Traffic signal light recognition method, device, electronic device and storage medium

By acquiring and processing the image data of the traffic light box, and using a multi-label classification network to determine the hierarchical attribute status of the traffic light, the problem of low accuracy of traffic light status recognition in the prior art is solved, and higher positioning accuracy and lower computing overhead are achieved.

CN117953464BActive Publication Date: 2025-05-09BEIJING PHIGENT TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311812302.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-05-09
Estimated Expiration
2043-12-26

AI Technical Summary

Technical Problem

In the prior art, the accuracy of traffic light status recognition is not high, and the recognition effect based on traditional image processing algorithms is poor and the robustness is poor. However, the recognition calculation overhead based on deep learning algorithms is large and the target size is too small, which reduces the detection accuracy.

Method used

By obtaining the location information of the target traffic light box in the image data of the current road, the multi-label classification network is used to determine the hierarchical attribute status of the traffic light box, and the multi-label classification results are post-processed to output the perceived results of the traffic light.

Benefits of technology

The positioning accuracy of traffic light status detection is improved, the positioning difficulty is reduced by increasing the detection size, and the complete semantic information is provided through multi-label classification tasks, reducing computational overhead and complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117953464B_ABST
    Figure CN117953464B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, electronic device and storage medium for identifying traffic lights, the method comprising: obtaining the location information of a target traffic light box in the image data of the current road; determining a multi-label classification result of the hierarchical attribute state of the target traffic light box according to the location information of the target traffic light box; post-processing the multi-label classification result, and outputting the perception result of the traffic lights on the target traffic light box. That is to say, in the present invention, the traffic light box is used as the target for detection, which increases the detection size of the target, makes it easier to detect the target, reduces the difficulty of target detection and positioning, and describes the semantic states of multiple lights on the traffic light box through a multi-label classification task of hierarchical attributes, provides complete semantic information, reduces the overhead and complexity caused by exhaustive enumeration, and thus improves the positioning accuracy of traffic light state detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method, device, electronic device and computer-readable storage medium for identifying a traffic light. Background Art

[0002] With the rapid development of the automobile industry and the improvement of people's living standards, cars have become people's main means of transportation, and people's requirements for cars have gradually shifted from safety and stability to intelligence. Assisted driving and future unmanned driving have become a trend. In assisted driving and future unmanned driving, real-time traffic conditions are very important, and real-time traffic conditions are largely regulated by traffic lights. Traffic lights are lights that direct traffic operations, generally consisting of red lights, green lights, and yellow lights. Red lights mean no passage, green lights mean passage is allowed, and yellow lights mean warnings. Traffic lights are divided into: motor vehicle lights, non-motor vehicle lights, pedestrian crossing lights, direction indicator lights (arrow lights), countdowns, lane lights, flashing warning lights, road and railway level crossing lights, etc.

[0003] In the related technology, traffic lights can be detected and identified based on traditional image processing algorithms, or by using deep learning algorithms. However, the solutions based on traditional image processing algorithms have few recognition types, poor results, and poor robustness. Complex algorithm logic needs to be designed, and it is difficult to use one algorithm to perceive all types of traffic lights. In the recognition based on deep learning algorithms, during the detection process of traffic information lights, although each traffic information light bulb is detected as a unit and the detected target position is classified, enumerating all categories requires more computing overhead, and the problem of small target size also greatly reduces the accuracy and effect of detection.

[0004] Therefore, how to improve the accuracy of traffic light status recognition is a technical problem that needs to be solved. Summary of the invention

[0005] The present invention provides a method, device, electronic device and computer-readable storage medium for identifying a traffic light, so as to at least solve the problem of low accuracy in identifying the state of a traffic light in the related art. The technical solution of the present invention is as follows:

[0006] According to a first aspect of an embodiment of the present invention, a method for identifying a traffic light is provided, comprising:

[0007] Obtain the location information of the target traffic light box in the image data of the current road;

[0008] Determining a multi-label classification result of a hierarchical attribute state of the target traffic light box according to the location information of the target traffic light box;

[0009] The multi-label classification result is post-processed to output a perception result of the traffic light on the target traffic light box.

[0010] Optionally, the step of obtaining the location information of the target traffic signal light box in the image data of the current road includes:

[0011] Get the image data of the current road;

[0012] The image data is input into a detection model trained in a target detection network for detection, and a plurality of bounding boxes including the location information of the target traffic light box, and the type and confidence level of each bounding box are output, wherein each bounding box is used to represent the location information of the target traffic light box, the type is used to confirm whether the bounding box contains the target traffic light box, and the confidence level represents the probability that the bounding box contains the target traffic light box.

[0013] Optionally, the method includes: pre-selecting a detection model in a training target detection network, including:

[0014] Obtain a training set, the training set including a plurality of images with traffic light boxes, and an image with a label annotated for each image, the label including: category and location information;

[0015] The training set is input into the detection model in the target detection network, and the target detection training is performed using the training script of YoloX. During the target training process, the target detection result of each image with a traffic light box is compared with the annotation label of the corresponding image, and the loss value is calculated according to the difference in the comparison result. Based on the loss value, the parameters of the detection model are updated through the back propagation mechanism. After multiple iterations, until the loss value reaches the minimum value or the detection model converges, a trained detection model is obtained.

[0016] Optionally, the multi-label classification result of determining the hierarchical attribute state of the target traffic light box according to the location information of the target traffic light box includes:

[0017] Based on the position information of the target traffic light box, cropping the corresponding target traffic light box image;

[0018] Inputting the target traffic light box image into a multi-label classification network for multi-label classification, and outputting a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box through each multi-label output channel;

[0019] Compare the vector of each attribute prediction value with the corresponding preset threshold value, and take the vector of each attribute prediction value greater than the preset threshold value as the classification result of the corresponding label;

[0020] The classification results of all labels are counted to obtain a multi-label classification result of the hierarchical attribute state of the traffic light on the target traffic light box, wherein the hierarchical attributes include: direction, state, shape, color and countdown value.

[0021] Optionally, the multi-label classification result of determining the hierarchical attribute state of the target traffic light box according to the location information of the target traffic light box further includes:

[0022] Performing super-resolution processing on the cropped target traffic light box image to obtain an ultra-high-resolution target traffic light box image;

[0023] The ultra-high resolution image of the target traffic light box is input into a multi-label classification network for multi-label classification, and a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box is output through each multi-label output channel.

[0024] Optionally, the post-processing of the multi-label classification result according to the attributes represented by each output channel and outputting the perception result of the traffic light on the target traffic light box includes:

[0025] Binarizing the multi-label classification result to obtain a binarized result;

[0026] The binarization result is used as a perception result of the traffic light on the target traffic light box, and the perception result includes: position information and category information.

[0027] Optionally, the binarization of the multi-label classification result according to the attributes represented by each output channel to obtain a binarization result includes:

[0028] Compare the multi-label classification result with the corresponding preset threshold according to the attribute represented by each output channel;

[0029] When the multi-label classification results all meet the corresponding preset thresholds, the multi-label classification results are fused in sequence according to the attribute channels, and the perception result of the traffic light on the target traffic light box is output.

[0030] Optionally, the binarization of the multi-label classification result according to the attributes represented by each output channel to obtain a binarization result further includes:

[0031] When there are multi-label classification results that do not meet the corresponding preset threshold, the classification results that do not meet the preset threshold are eliminated through a time domain filter;

[0032] The eliminated multi-label classification results are fused in sequence according to the attribute channels, and the perception results of the traffic light on the target traffic light box are output.

[0033] Optionally, before outputting the perception result of the traffic light on the target traffic light box, the method further includes:

[0034] Compare the perception result with the perception result of the previous time and the perception result of the previous time respectively;

[0035] If the perception results are all the same, then the step of outputting the perception result of the traffic light on the target traffic light box is performed; or

[0036] If the perception result is the same as the perception result at the previous time and different from the perception result at the time before that, then the step of outputting the perception result of the traffic light on the target traffic light box is performed; or

[0037] If the perception result is different from the perception result at the previous time and is the same as the perception result at the previous time, the perception result at the previous time is output.

[0038] According to a second aspect of an embodiment of the present invention, there is provided a traffic light recognition device, comprising:

[0039] An acquisition module is used to acquire the location information of the target traffic signal light box in the image data of the current road;

[0040] A determination module, used to determine a multi-label classification result of a hierarchical attribute state of the target traffic light box according to the location information of the target traffic light box;

[0041] The processing module is used to post-process the multi-label classification result and output the perception result of the traffic light on the target traffic light box.

[0042] Optionally, the acquisition module includes:

[0043] A data acquisition module, used to acquire image data of the current road;

[0044] A detection module is used to input the image data into a detection model trained in a target detection network for detection, and output a plurality of bounding boxes including the location information of a target traffic light box, and the type and confidence level of each bounding box, wherein each bounding box is used to represent the location information of the target traffic light box, the type is used to confirm whether the bounding box contains the target traffic light box, and the confidence level is the probability of including the target traffic light box in the bounding box.

[0045] Optionally, the device further includes: a training module, which is used to pre-select a detection model in the training target detection network, specifically including:

[0046] A training set acquisition module, used to acquire a training set, the training set including a plurality of images with traffic light boxes, and an image with a label annotated for each image, the label including: category and location information;

[0047] A training module is used to input the training set into the detection model in the target detection network, and use the YoloX training script to perform target detection training. During the target training process, the target detection result of each image with a traffic light box is compared with the annotation label of the corresponding image, and the loss value is calculated according to the difference in the comparison result. Based on the loss value, the parameters of the detection model are updated through a back propagation mechanism. After multiple iterations, until the loss value reaches a minimum value or the detection model converges, a trained detection model is obtained.

[0048] Optionally, the determining module includes:

[0049] A cropping module, used for cropping the corresponding target traffic light light box image based on the position information of the target traffic light light box;

[0050] A multi-label classification module, used for inputting the target traffic light box image into a multi-label classification network for multi-label classification, and outputting a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box through each multi-label output channel;

[0051] A first comparison module is used to compare the vector of each attribute prediction value with the corresponding preset threshold value, and take the vector of each attribute prediction value greater than the preset threshold value as the classification result of the corresponding label;

[0052] The statistics module is used to count the classification results of all labels to obtain the multi-label classification results of the hierarchical attribute status of the traffic light on the target traffic light box, wherein the hierarchical attributes include: direction, state, shape, color and countdown value.

[0053] Optionally, the determining module further includes:

[0054] A resolution processing module, used for performing super-resolution processing on the target traffic light box image cropped by the cropping module to obtain an ultra-high-resolution image of the target traffic light box;

[0055] The multi-label classification module is also used to input the ultra-high resolution image of the target traffic light box into the multi-label classification network for multi-label classification, and output a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box through each multi-label output channel.

[0056] Optionally, the processing module includes:

[0057] A binarization processing module is used to perform binarization processing on the multi-label classification result according to the attribute represented by each output channel to obtain a binarization processing result;

[0058] The perception determination module is used to use the binarization processing result as the perception result of the traffic light on the target traffic light box, and the perception result includes: position information and category information.

[0059] Optionally, the binarization processing module includes:

[0060] A second comparison module is used to compare the multi-label classification result with the corresponding preset threshold according to the attribute represented by each output channel;

[0061] The first fusion module is used to fuse the multi-label classification results in sequence according to the attribute channels when the multi-label classification results all meet the corresponding preset thresholds, and output the perception result of the traffic light on the target traffic light box.

[0062] Optionally, the binarization processing module further includes:

[0063] A filtering module, configured to, when any of the multi-label classification results do not meet the corresponding preset threshold, remove the classification results that do not meet the preset threshold through a time domain filter;

[0064] The second fusion module is used to fuse the eliminated multi-label classification results in sequence according to the attribute channels, and output the perception result of the traffic light on the target traffic light box.

[0065] Optionally, the device further comprises:

[0066] a third comparison module, configured to compare the perception result with the perception result at the previous time before the current time and the perception result at the previous time before the current time respectively before the processing module outputs the perception result of the traffic light on the target traffic light box;

[0067] A first output module, configured to output a perception result of the traffic light on the target traffic light box when the perception results compared by the third comparison module are all the same; and / or

[0068] a second output module, configured to output the perception result of the traffic light on the target traffic light box when the perception result compared by the third comparison module is the same as the perception result at the previous time and different from the perception result at the previous time; and / or

[0069] The third output module is used to output the perception result of the previous time when the perception result compared by the third comparison module is different from the perception result of the previous time and is the same as the perception result of the previous time.

[0070] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, including:

[0071] processor;

[0072] a memory for storing instructions executable by the processor;

[0073] The processor is configured to execute the instructions to implement the traffic light recognition method as described above.

[0074] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the traffic light recognition method as described above.

[0075] According to a fifth aspect of an embodiment of the present invention, there is provided a computer program product, comprising a computer program or instructions, which, when executed by a processor of an electronic device, implements the traffic light recognition method as described above.

[0076] The technical solution provided by the embodiments of the present invention brings at least the following beneficial effects:

[0077] In an embodiment of the present invention, the location information of the target traffic light box in the image data of the current road is obtained; based on the location information of the target traffic light box, the multi-label classification result of the hierarchical attribute state of the target traffic light box is determined; the multi-label classification result is post-processed to output the perception result of the traffic lights on the target traffic light box. In other words, in an embodiment of the present invention, the traffic light box is used as the target for detection, which increases the detection size of the target, makes it easier to detect the target, reduces the difficulty of detection and positioning, and describes the semantic states of the multiple lights on the traffic light box through the multi-label classification task of hierarchical attributes, giving complete semantic information, thereby improving the positioning accuracy of traffic light state detection.

[0078] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention, and do not constitute an improper limitation of the present invention. In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0080] Figure 1 The present invention provides a flow chart of a method for identifying a traffic light.

[0081] Figure 2 A schematic diagram of a tree structure map provided in an embodiment of the present invention.

[0082] Figure 3 A schematic diagram of a set of multi-label classification results provided by an embodiment of the present invention.

[0083] Figure 4 A schematic diagram of a traffic signal light box legend and its corresponding attribute states provided in an embodiment of the present invention.

[0084] Figure 5 This is another flow chart of a method for identifying a traffic light provided by an embodiment of the present invention.

[0085] Figure 6 The invention is a block diagram of a traffic light recognition device provided by an embodiment of the present invention.

[0086] Figure 7It is a block diagram of a determination module provided by an embodiment of the present invention.

[0087] Fig. 7A It is another block diagram of a determination module provided by an embodiment of the present invention.

[0088] Figure 8 It is a block diagram of an electronic device provided by an embodiment of the present invention.

[0089] Fig. 9 The invention is a block diagram of a traffic light recognition device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0090] In order to enable ordinary persons in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings.

[0091] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0092] Figure 1 is a flow chart of a method for identifying a traffic light provided by an embodiment of the present invention, such as Figure 1 As shown, the traffic light recognition method includes the following steps:

[0093] Step 101: Acquire the location information of the target traffic signal light box in the image data of the current road;

[0094] Step 102: determining a multi-label classification result of a hierarchical attribute state of the target traffic light box according to the location information of the target traffic light box;

[0095] Step 103: Post-process the multi-label classification result and output the perception result of the traffic light on the target traffic light box.

[0096] The traffic light recognition method described in the present invention can be applied to the vehicle side, the cloud side, the automatic lane change assistance system in intelligent driving, the automatic driving control system, etc., without limitation here. The vehicle-side implementation equipment can be an on-board terminal, a vehicle control platform, an industrial computer and other electronic equipment. The cloud side can be an independent server or a server cluster, or a server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, intermediate services, domain name services, security services, content distribution networks, or a big data and artificial intelligence platform, etc., without limitation here.

[0097] Combine the following Figure 1 , the specific implementation steps of a traffic signal light recognition method provided by an embodiment of the present invention are described in detail.

[0098] In step 101, the location information of the target traffic light box in the image data of the current road is obtained.

[0099] In this step, first, image data of the current road is obtained;

[0100] Specifically, the image data of the current road may be collected by hardware such as a camera on a vehicle or a car-mounted camera. The image data may be in RGB format or other formats, which is not limited in this embodiment.

[0101] Secondly, the target detection network is used to detect the location information of the target traffic light box from the image data. That is, the image data is input into the detection model trained in the target detection network for detection, and multiple bounding boxes including the location information of the target traffic light box are output, as well as the type and confidence of each bounding box, wherein each bounding box is used to represent the location information of the target traffic light box, the type is used to confirm whether the bounding box is the target traffic light box, and the confidence is the probability of the bounding box including the target traffic light box.

[0102] The detection model in the target detection network is a pre-trained detection model, which can use a traditional image algorithm or a deep learning algorithm to detect the image data of the current road, thereby detecting the location information of the target traffic light box. Of course, in this embodiment, other algorithms can also be used to detect the location information of the target traffic light box, which is not limited in this embodiment.

[0103] For example, the one stage and the two stage in the target detection network in deep learning are used to detect the location information of the target traffic light box from the image data.

[0104] Among them, the two-stage is to use the corresponding Region Proposal algorithm (which can be a traditional algorithm or a neural network, etc.) to generate target candidate regions from the image picture of the current road input, and then send all the target candidate regions to the classifier for classification, so as to obtain the location information of the target traffic light box.

[0105] One-stage, firstly divide the input image of the current road into NxN image regions (image patches), then each image patch can have M anchor boxes of fixed size or anchor boxes of different lengths and widths, output the location and classification label of the anchor box, and thus obtain the location information of the target traffic light box. The specific aspect ratio can be obtained by algorithms such as kmeans.

[0106] In this step, the traffic light is detected according to the light box instead of the light bulb. By increasing the detection size, the difficulty of detection and positioning is reduced and the detection and positioning accuracy is improved.

[0107] In step 102, a multi-label classification result of the hierarchical attribute state of the target traffic light box is determined according to the position information of the target traffic light box.

[0108] In this step, first, based on the location information of the target traffic light box, the corresponding target traffic light box image is cropped; then, the target traffic light box image is input into a multi-label classification network for multi-label classification, and a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box is output through each output channel of the multi-label, and each attribute prediction value vector is compared with the corresponding preset threshold value, and each attribute prediction value vector greater than the preset threshold value is used as the classification result of the corresponding label, and the classification results of all labels are counted to obtain the multi-label classification result of the hierarchical attribute state of each traffic light in the target traffic light box, wherein the hierarchical attributes include: direction, state, shape, color and countdown value. Wherein, the multi-label classification network includes a ResNet structure, and the ResNet structure includes: a convolutional layer, a residual block, a global average pooling layer, and a fully connected layer with multiple output nodes.

[0109] In this embodiment, the ResNet classification network is used as an example. Based on the ResNet classification network, the improvement is to change the output supervision from the one-hot vector [0,0,1,0,…,0,,0] to a multi-label [0,0,1,1,…,1,,0] vector representation; the loss supervision is changed from softmax multi-classification to the activation function (sigmoid) of each channel, which is used to distinguish each channel, that is, to distinguish whether the probability of the corresponding position of the vector of each channel is closer to 0 or 1.

[0110] That is to say, the multi-label classification network in this embodiment is a trained multi-label classification network, and the training process includes:

[0111] 1) Prepare a dataset, which includes: image data of traffic lights, and labels for each traffic light image that annotate multiple attributes, such as orientation, state, shape, color, and countdown value, to ensure that the labels contain various possible attribute combinations.

[0112] 2) Network structure: The multi-label classification network in this embodiment uses the classic ResNet structure as an example. The ResNet includes a convolutional layer, a residual block, and a global average pooling layer. Before the last fully connected layer, the number of output channels is adjusted to be equal to the number of labels required in the multi-label classification task.

[0113] 3) Output layer design: replace the last fully connected layer with a fully connected layer with multiple output nodes. Each node corresponds to a label, and the output represents the confidence of the label. The confidence uses the Sigmoid activation function to map the output value to the range of [0,1].

[0114] 4) Loss function. Since it is a multi-label classification task, the loss function is switched from the traditional multi-category cross entropy (softmax) to binary cross entropy. For each label, binary cross entropy loss is used to process the prediction of each output node independently.

[0115] 5) Training, using the prepared data set with multi-level attribute labels to train the multi-label classification network. During the training process, the multi-label classification network will learn to extract features of various attributes from the image, and perform convolutional layers, residual blocks, global average pooling layers and other processing on various attribute features. Finally, the number of output channels of the fully connected layer is adjusted to be equal to the number of labels required in the multi-label classification task, and a vector of multiple attribute prediction values ​​is output. The vector of each attribute prediction value is compared with the corresponding preset threshold, and the vector of each attribute prediction value greater than the preset threshold is used as the classification result of the corresponding label. The classification results of all labels are counted to obtain the multi-label classification result of the hierarchical attribute state of each traffic light in the target traffic light box, wherein the hierarchical attributes include: direction, state, shape, color and countdown value.

[0116] After the training is completed, the traffic light image to be tested is input into the multi-label classification network for multi-label classification. The multi-label classification network outputs a vector containing multiple attribute prediction values ​​of the traffic light on the target traffic light box through each multi-label output channel. These prediction values ​​can be converted into final attribute judgments through threshold judgment or other rule judgments. Finally, result analysis: Analyze the output of the multi-label classification network to obtain the prediction result of the multi-label classification result of the hierarchical attribute state of the traffic light to be tested, so as to obtain a description of the multi-level attributes of the traffic light to be tested, such as direction, state, shape, color and countdown value.

[0117] It should be noted that in this step, all target images on the target traffic light box can also be cropped according to the detection frame, and all target images can be input into the image classification network for multi-label classification; among them, the multi-label classification and recognition method can use vgg, resnet, mobilenet, RepVGG, CSPNet and NAS and other methods according to specific scenarios and requirements; of course, multi-label classification can also be performed according to the hierarchical attributes of each target image through a multi-label classification model to obtain a multi-label classification result of the hierarchical attribute state of the target traffic light box.

[0118] Among them, the multi-label (multi-state) classification result is a tree diagram, as shown in the following figure: Figure 2 As shown, Figure 2 A schematic diagram of a tree structure map provided in an embodiment of the present invention.

[0119] Depend on Figure 2As shown in the tree structure, each traffic light in the traffic light box is described according to hierarchical attributes such as direction, state, shape, color, countdown value, etc. The hierarchical attributes can be decoupled from each other, which greatly increases the degree of freedom and scalability, reduces the difficulty of exhaustive enumeration, and facilitates the attribute combination and utilization of subsequent processing algorithms.

[0120] like Figure 2 As shown, including:

[0121] Direction may include: left side (cannot see), left side (can see), forward, reverse, right side (can see), right side (cannot see), etc.

[0122] The status may include: whether an extinguished state exists, whether it exists or not.

[0123] Shapes, for example, the positive state may include a circle, an arrow up, an arrow down, an arrow left, an arrow right, ..., a pedestrian, a non-motor vehicle, a number (countdown time).

[0124] Colors can include red, yellow, and green. For example, the color of a circular traffic light changes to red, yellow, and green. For example, the color of a number changes to red, yellow, and green.

[0125] Numerical value, ones...ones, tens...tens, etc.

[0126] For example, the current attribute status of the traffic signal light box is: forward + there is an extinguished bulb + a lit circular green light + a countdown green value of 09, etc.

[0127] In step 103, the multi-label classification result is post-processed to output the perception result of the traffic light on the target traffic light box.

[0128] In this step, the multi-label classification result is first binarized to obtain a binarized result; and the binarized result is used as a perception result of the traffic light on the target traffic light box, and the perception result includes: location information and category information.

[0129] That is to say, in the above steps, the multi-label classification result output by the sigmoid of the multi-label classification network is a floating point number ranging from 0 to 1, and the floating point number needs to be binarized to a fixed point number through a preset threshold.

[0130] Specifically, the binarization processing is performed on the multi-label classification result to obtain a binarization processing result, including:

[0131] The multi-label classification results are compared with the corresponding preset thresholds respectively; when the multi-label classification results all meet the corresponding preset thresholds, the multi-label classification results are sequentially fused according to the attribute channels, and the perception results of the traffic lights on the target traffic light box are output; or, when there are multi-label classification results that do not meet the corresponding preset thresholds, the classification results that do not meet the preset thresholds are eliminated through a time domain filter; the eliminated multi-label classification results are sequentially fused according to the attribute channels, and the perception results of the traffic lights on the target traffic light box are output, and the perception results include: location information and category information.

[0132] In this embodiment, the subsequent processing of parsing multi-label results can refer to Figure 2 As shown, when the first few positions output by the multi-label classification network control the orientation, among the outputs of the first few control orientation positions, the position with the largest value is selected and recorded as the orientation of the traffic light box.

[0133] That is to say, if there is a light bulb that is extinguished, the output value of this position is judged to be 1 when it is greater than the preset threshold, indicating that there is a light bulb that is extinguished in the target traffic light box; and 0 when it is less than the preset threshold, indicating that there is no light bulb that is extinguished.

[0134] The next three positions form a group to judge the color and shape of the light bulb in the target traffic light box. There may be only red light, or there may be both red and green lights.

[0135] For example, there are 4 groups in a group, such as [round, green, yellow, red]

[0136] When the circular position output is 1, it means that a circular light is on. Then the one with the highest probability is selected from the last three positions to determine the color of the circular light, completing the combination and describing the color and shape attributes of the light bulb. Other colors and shapes are analyzed in the same way.

[0137] It should be noted that if the shape is a position of a number and there is a value, after determining the shape, it is necessary to determine the value in the subsequent 22 channels. 0-11 determines what the tens digit is, which is still unknown, and 12-22 determines what the digit is, which is still unknown. Then, they are combined to form a complete attribute description.

[0138] like Figure 3 As shown in FIG. 1 , a schematic diagram of a set of multi-label classification results provided by an embodiment of the present invention is shown. Figure 3 As shown, it indicates that the current attribute status of the traffic light box is: forward + there is an extinguished bulb + a lit circular green light + a countdown green value of 09.

[0139] Specifically, the schematic diagram of the traffic signal light box legend and its corresponding attribute state (ie, hierarchical attribute state) provided in this embodiment is as follows: Figure 4 As shown, Figure 4 This is just an example to illustrate, and is not limited to this in actual application.

[0140] like Figure 4 As shown, the property status of the first traffic light box is:

[0141] Forward + non-motor vehicle (green) + pedestrian (green) + number (green, tens digit 2, ones digit 8).

[0142] The attribute status of the second traffic signal light box is: forward + U-turn (red light) + off.

[0143] The attribute status of the third traffic light box is: forward + straight (red light) + left turn (green light) + right turn (green light).

[0144] The attribute status of the fourth traffic signal light box is: light can be seen on the right + turn left (red light) + number (red, tens digit 0, ones digit 7) + off.

[0145] In an embodiment of the present invention, the location information of the target traffic light box in the image data of the current road is obtained; based on the location information of the target traffic light box, the multi-label classification result of the hierarchical attribute state of the target traffic light box is determined; the multi-label classification result is post-processed to output the perception result of the traffic lights on the target traffic light box. In other words, in the present invention, the traffic light box is used as the target for detection, which increases the detection size of the target, makes it easier to detect the target, reduces the difficulty of target detection and positioning, and describes the semantic states of multiple lights on the traffic light box through the multi-label classification task of hierarchical attributes, provides complete semantic information, reduces the overhead and complexity caused by exhaustive enumeration, and thus improves the positioning accuracy of traffic light state detection.

[0146] Furthermore, the traffic light box is detected based on the target detection network (such as convolutional neural network); the target image is cropped according to the detection frame, and the image is input into the multi-label classification network for multi-label classification; and the multi-frame results are fused and filtered in time sequence to ensure the stability of perception, and finally the perception results of the traffic lights in the traffic light box are obtained, including location information and category information. In other words, by combining the tree logic of the knowledge graph in the form of multiple labels, each traffic light box can be fully described and combined, which greatly expands the upper limit of the perception ability of the solution, without exhausting all combinations, and reducing the computational overhead.

[0147] In the embodiment of the present invention, the positioning (detection), identification (classification) and post-processing of traffic lights adopt a decoupled design. The decoupling of detection, classification and post-processing can be debugged, optimized and iterated separately, forming a building block assembly, which is conducive to each part being responsible for its own work and can conveniently replace and expand the functions of each part; the embodiment of the present invention can be enhanced at the data end to facilitate the recognition of scenes such as weather, time, and lighting, thereby improving the robustness of the solution. In the later stage, the perception categories can be expanded according to different countries and cities to cover more usage scenarios at any time.

[0148] The use of a multi-label classification network and a tree-like attribute state design of a knowledge graph in an embodiment of the present invention can easily cover all attribute state combinations of traffic lights without enumerating all categories, thereby reducing the amount of calculation and the difficulty of network output; the attributes can also be decoupled from each other, which can be more conveniently provided to the post-processing algorithm, giving the post-processing operation more room to play.

[0149] In another embodiment, based on the above embodiment, the multi-label classification result of determining the hierarchical attribute state of the target traffic light box according to the location information of the target traffic light box may also include: performing super-resolution processing on the cropped target traffic light box image to obtain an ultra-high resolution image of the target traffic light box; inputting the ultra-high resolution image of the target traffic light box into a multi-label classification network for multi-label classification, specifically as follows: Figure 5 As shown, Figure 5 Another flow chart of a method for identifying a traffic light provided by an embodiment of the present invention, the method comprising:

[0150] Step 501: Acquire the location information of the target traffic signal light box in the image data of the current road.

[0151] Step 502: based on the position information of the target traffic light box, cropping out the corresponding target traffic light box image.

[0152] Step 503: performing super-resolution processing on the cropped target traffic light box image to obtain an ultra-high-resolution target traffic light box image.

[0153] In this step, super-resolution processing can improve the resolution of the original image through hardware or software methods, that is, a low-resolution image can be processed by super-resolution technology to obtain a large-size high-resolution image. The process is super-resolution reconstruction, or the size of the traffic light box in the current road image obtained is relatively small, and the image needs to be super-resolution processed to obtain a larger-size and ultra-high-resolution traffic light box image.

[0154] Step 504: Input the ultra-high resolution target traffic light box image into a multi-label classification network for multi-label classification.

[0155] Step 505: Output a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box through each output channel of the multiple labels.

[0156] Step 506: Compare the vector of each attribute prediction value with the corresponding preset threshold value, and take the vector of each attribute prediction value greater than the preset threshold value as the classification result of the corresponding label.

[0157] Step 507: Count the classification results of all labels to obtain a multi-label classification result of the hierarchical attribute state of the traffic light on the target traffic light box, wherein the hierarchical attributes include: direction, state, shape, color and countdown value.

[0158] Step 508: Post-process the multi-label classification result and output the perception result of the traffic light on the target traffic light box.

[0159] In this embodiment, the implementation process of step 501, step 502, and step 504 to step 507 is detailed in the implementation process of the corresponding steps in the above embodiment, and will not be repeated here.

[0160] In an embodiment of the present invention, between detection and classification, super-resolution processing is added to the detected image to improve the resolution of the image, that is, the detection result is cropped to obtain a cropped target traffic light box image, and then the cropped target traffic light box image is super-resolution processed to obtain an ultra-high resolution target traffic light box image. In other words, if the cropped target traffic light box image is relatively small, it is necessary to perform super-resolution processing on the image, that is, the size of the distant target (i.e., the traffic light box) reflected on the image pixels is too small (too few pixels), it is necessary to perform super-resolution processing first to obtain an image with increased size and higher resolution. The target traffic light box after super-resolution processing is directly input into the multi-label classification network for multi-label classification, which reduces the need to add additional processing operations and details to the target image when the size of the input target image is small in the multi-label classification network, thereby improving the perception effect.

[0161] Optionally, in another embodiment, based on the above embodiment, before outputting the perception result of the traffic light on the target traffic light box, the method may further include:

[0162] The perception result is compared with the perception result of the previous time and the perception result of the previous time before the current time respectively; wherein, in this embodiment, the previous time and the previous time both refer to the time of the perception result of the traffic light, and the previous time refers to the time of the previous perception result adjacent to the previous time.

[0163] If the perception results are all the same, then the step of outputting the perception result of the traffic light on the target traffic light box is performed; or

[0164] If the perception result is the same as the perception result at the previous time and different from the perception result at the time before that, then the step of outputting the perception result of the traffic light on the target traffic light box is performed; or

[0165] If the perception result is different from the perception result at the previous time and is the same as the perception result at the previous time, the perception result at the previous time is output.

[0166] Optionally, in an embodiment of the present invention, time-series multi-frame fusion post-processing can be used. Since the camera can continuously capture images of the current road, this embodiment can use the perception results of multiple frames to track each traffic light box, and then use a time domain filter to eliminate sudden recognition results (false recognition), thereby ensuring the stability of the output and improving the positioning accuracy of traffic light status detection.

[0167] In the embodiment of the present invention, the reduction of false detection by time series fusion through filters usually involves smoothing and filtering the time series to eliminate short-term noise or false detection. In the corresponding scenario, a simple filter technique, such as moving average or exponentially weighted moving average, can be used to smooth the time series to reduce false detection.

[0168] Taking moving average as an example, given a window size, the filter calculates the average value of the data in the window and uses this average value as the output, which can effectively smooth the signal and make it less sensitive to short-term noise or sudden changes.

[0169] Suppose there is a time series of traffic lights, where 1 represents a green light and 0 represents a red light. Consider a case where the moving average window size is 3:

[0170] Red Green Red Red:

[0171] Initially, the average value in the window is [red, green, red] and the output is red.

[0172] Then the window moves one position to the right, becoming [green, red, red], and the output is green.

[0173] Next, the window moves right again, becoming [red, red, red], and the output is red.

[0174] In this case, the short green light signal is smoothed out and the output is still red.

[0175] Red Red Green Green Green:

[0176] Initially, the average value in the window is [red, red, green] and the output is red.

[0177] The window moves right and becomes [red, green, green], and the output is red.

[0178] Shift right again, it becomes [green, green, green], and the output is green.

[0179] In this case, since the duration of the green light is longer, the green light signal can be maintained in the output for a period of time.

[0180] The above is just a simple example. In actual applications, different filters and parameters can be selected according to specific needs. Other filter technologies, such as exponential weighted moving average and median filtering, can also be used for time series fusion to reduce false detection.

[0181] It should be noted that, for the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that this implementation disclosure is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the present invention.

[0182] See also Figure 5 , is a block diagram of a traffic signal light recognition device provided by an embodiment of the present invention. Figure 5 The device includes: an acquisition module 501, a determination module 502 and a processing module 503, wherein:

[0183] An acquisition module 501 is used to acquire the location information of a target traffic signal light box in the image data of the current road;

[0184] A determination module 502, configured to determine a multi-label classification result of a hierarchical attribute state of the target traffic light box according to the location information of the target traffic light box;

[0185] The processing module 503 is used to post-process the multi-label classification result and output the perception result of the traffic light on the target traffic light box.

[0186] Optionally, in another embodiment, based on the above embodiment, the acquisition module includes: a data acquisition module and a detection module, wherein:

[0187] A data acquisition module, used to acquire image data of the current road;

[0188] A detection module is used to input the image data into a detection model trained in a target detection network for detection, and output a plurality of bounding boxes including the location information of a target traffic light box, and the type and confidence level of each bounding box, wherein each bounding box is used to represent the location information of the target traffic light box, the type is used to confirm whether the bounding box contains the target traffic light box, and the confidence level is the probability of including the target traffic light box in the bounding box.

[0189] Optionally, in another embodiment, based on the above embodiment, the device further includes: a training module for preselecting a detection model in a training target detection network, specifically including:

[0190] A training set acquisition module, used to acquire a training set, the training set including a plurality of images with traffic light boxes, and an image with a label annotated for each image, the label including: category and location information;

[0191] A training module is used to input the training set into the detection model in the target detection network, and use the YoloX training script to perform target detection training. During the target training process, the target detection result of each image with a traffic light box is compared with the annotation label of the corresponding image, and the loss value is calculated according to the difference in the comparison result. Based on the loss value, the parameters of the detection model are updated through a back propagation mechanism. After multiple iterations, until the loss value reaches a minimum value or the detection model converges, a trained detection model is obtained.

[0192] Optionally, in another embodiment, based on the above embodiment, the determination module 502 includes: a clipping module 701, a multi-label classification module 702, a first comparison module 703 and a statistical module 704, and its structural block diagram is as follows: Figure 7 As shown,

[0193] A cropping module 701, configured to crop a corresponding target traffic light light box image based on the position information of the target traffic light light box;

[0194] A multi-label classification module 702 is used to input the target traffic light box image into a multi-label classification network for multi-label classification, and output a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box through each multi-label output channel;

[0195] A first comparison module 703 is used to compare the vector of each attribute prediction value with the corresponding preset threshold value, and take the vector of each attribute prediction value greater than the preset threshold value as the classification result of the corresponding label;

[0196] The statistics module 704 is used to count the classification results of all labels to obtain the multi-label classification results of the hierarchical attribute state of each traffic light in the target traffic light box, wherein the hierarchical attributes include: direction, state, shape, color and countdown value.

[0197] Optionally, in another embodiment, based on the above embodiment, the determination module 502 further includes: a resolution processing module 705, whose structural block diagram is as follows: Fig. 7A As shown, Fig. 7A Another block diagram of a determination module provided by an embodiment of the present invention, wherein:

[0198] The resolution processing module 705 is used to perform super-resolution processing on the target traffic light box image cropped by the cropping module 701 to obtain an ultra-high-resolution image of the target traffic light box;

[0199] The multi-label classification module 702 is further used to input the ultra-high resolution image of the target traffic light box into the multi-label classification network for multi-label classification, and output a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box through each multi-label output channel.

[0200] Optionally, in another embodiment, based on the above embodiment, the processing module includes: a binarization processing module and a perception determination module, wherein:

[0201] A binarization processing module is used to perform binarization processing on the multi-label classification result according to the attribute represented by each output channel to obtain a binarization processing result;

[0202] The perception determination module is used to use the binarization processing result as the perception result of the traffic light on the target traffic light box, and the perception result includes: position information and category information.

[0203] Optionally, in another embodiment, based on the above embodiment, the binarization processing module includes:

[0204] A second comparison module is used to compare the multi-label classification result with the corresponding preset threshold according to the attribute represented by each output channel;

[0205] The first fusion module is used to fuse the multi-label classification results in sequence according to the attribute channels when the multi-label classification results all meet the corresponding preset thresholds, and output the perception result of the traffic light on the target traffic light box.

[0206] Optionally, in another embodiment, based on the above embodiment, the binarization processing module further includes:

[0207] A filtering module, configured to, when any of the multi-label classification results do not meet the corresponding preset threshold, remove the classification results that do not meet the preset threshold through a time domain filter;

[0208] The second fusion module is used to fuse the eliminated multi-label classification results in sequence according to the attribute channels, and output the perception result of the traffic light on the target traffic light box.

[0209] Optionally, in another embodiment, based on the above embodiment, the device further includes:

[0210] a third comparison module, configured to compare the perception result with the perception result at the previous time before the current time and the perception result at the previous time before the current time respectively before the processing module outputs the perception result of the traffic light on the target traffic light box;

[0211] A first output module, configured to output a perception result of the traffic light on the target traffic light box when the perception results compared by the third comparison module are all the same; and / or

[0212] a second output module, configured to output the perception result of the traffic light on the target traffic light box when the perception result compared by the third comparison module is the same as the perception result at the previous time and different from the perception result at the previous time; and / or

[0213] The third output module is used to output the perception result of the previous time when the perception result compared by the third comparison module is different from the perception result of the previous time and is the same as the perception result of the previous time.

[0214] Optionally, an embodiment of the present invention further provides an electronic device, including:

[0215] processor;

[0216] a memory for storing instructions executable by the processor;

[0217] The processor is configured to execute the instructions to implement the traffic light recognition method as described above.

[0218] Optionally, an embodiment of the present invention further provides a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the traffic light recognition method as described above.

[0219] Optionally, an embodiment of the present invention further provides a computer program product, including a computer program or instructions, which implement the above-mentioned traffic light recognition method when executed by a processor of an electronic device.

[0220] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0221] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0222] See also Figure 8 , is a frame of an electronic device 800 provided by an embodiment of the present invention Figure 8 As shown in the figure, it includes a processor 801, a communication interface 802, a memory 803 and a communication bus 804, wherein the processor 801, the communication interface 802, and the memory 803 communicate with each other through the communication bus 804;

[0223] A memory 803, used to store the processor executable instructions;

[0224] The processor 801 is used to implement the automatic lane change control method as described above when executing the executable instructions in the memory 803.

[0225] The communication bus in this embodiment may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0226] The communication interface is used for communication between the above electronic device and other devices.

[0227] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0228] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0229] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.

[0230] In an embodiment, a computer-readable storage medium is also provided, and when the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device 800 can perform the above-mentioned traffic light recognition method. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0231] In an embodiment, a computer program product is further provided, including a computer program or instructions. When the computer program or instructions are executed by the processor 801 of the electronic device 800, the electronic device 800 executes the traffic light recognition method shown above.

[0232] Fig. 9 is a block diagram of a device 900 for identifying a traffic light provided by an embodiment of the present invention. For example, the device 900 may be provided as a server. Fig. 9 , the apparatus 900 includes a processing component 922, which further includes one or more processors, and a memory resource represented by a memory 932 for storing instructions, such as an application, that can be executed by the processing component 922. The application stored in the memory 932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 922 is configured to execute the instructions to perform the above method.

[0233] The device 900 may also include a power supply component 926 configured to perform power management of the device 900, a wired or wireless network interface 950 configured to connect the device 900 to a network, and an input / output (I / O) interface 958. The device 900 may operate based on an operating system stored in the memory 932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.

[0234] In the embodiments of the present invention, the term "module" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processor or memory) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module.

[0235] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art that are not disclosed by the present invention. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present invention are indicated by the following claims.

[0236] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A method for identifying a traffic light, characterized in that: include: Obtain the location information of the target traffic light box in the image data of the current road; Based on the position information of the target traffic light box, cropping a corresponding target traffic light box image; Inputting the target traffic light light box image into a multi-label classification network for multi-label classification, and obtaining a multi-label classification result of the hierarchical attribute state of each traffic light in the target traffic light light box; The hierarchical attributes are decoupled from each other, and the hierarchical attributes include: orientation, state, shape, color and countdown value; Binarizing the multi-label classification result to obtain a binarized result; The binarization processing result is output as the perception result of the traffic light on the target traffic light box.

2. The method for identifying a traffic light according to claim 1, characterized in that: The step of obtaining the location information of the target traffic signal light box in the image data of the current road includes: Get the image data of the current road; The image data is input into a detection model trained in a target detection network for detection, and a plurality of bounding boxes including the location information of the target traffic light box, and the type and confidence level of each bounding box are output, wherein each bounding box is used to represent the location information of the target traffic light box, the type is used to confirm whether the bounding box contains the target traffic light box, and the confidence level represents the probability that the bounding box contains the target traffic light box.

3. The method for identifying a traffic light according to claim 2, characterized in that: The method comprises: pre-selecting a detection model in a training target detection network, including: Obtain a training set, the training set including a plurality of images with traffic light boxes, and an image with a label annotated for each image, the label including: category and location information; The training set is input into the detection model in the target detection network, and the target detection training is performed using the training script of YoloX. During the target training process, the target detection result of each image with a traffic light box is compared with the annotation label of the corresponding image, and the loss value is calculated according to the difference in the comparison result. Based on the loss value, the parameters of the detection model are updated through the back propagation mechanism. After multiple iterations, until the loss value reaches the minimum value or the detection model converges, a trained detection model is obtained.

4. The method for identifying a traffic light according to claim 1, characterized in that: The step of inputting the target traffic light box image into a multi-label classification network for multi-label classification to obtain a multi-label classification result of the hierarchical attribute state of each traffic light in the target traffic light box includes: Inputting the target traffic light box image into a multi-label classification network for multi-label classification, and outputting a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box through each multi-label output channel; Compare the vector of each attribute prediction value with the corresponding preset threshold value, and take the vector of each attribute prediction value greater than the preset threshold value as the classification result of the corresponding label; The classification results of all labels are counted to obtain a multi-label classification result of the hierarchical attribute state of the traffic light on the target traffic light box.

5. The method for identifying a traffic light according to claim 4, characterized in that: The step of inputting the target traffic light box image into a multi-label classification network for multi-label classification to obtain a multi-label classification result of the hierarchical attribute state of each traffic light in the target traffic light box further includes: Performing super-resolution processing on the cropped target traffic light box image to obtain an ultra-high-resolution target traffic light box image; The ultra-high resolution image of the target traffic light box is input into a multi-label classification network for multi-label classification, and a vector of multiple attribute prediction values ​​of the traffic light on the target traffic light box is output through each multi-label output channel.

6. The method for identifying a traffic light according to claim 1, characterized in that: The perception result includes: location information and category information.

7. The method for identifying a traffic light according to claim 6, characterized in that: The binarization processing of the multi-label classification result to obtain the binarization processing result includes: Compare the multi-label classification result with the corresponding preset threshold according to the attribute represented by each output channel; When the multi-label classification results all meet the corresponding preset thresholds, the multi-label classification results are fused in sequence according to the attribute channels, and the perception result of the traffic light on the target traffic light box is output.

8. The method for identifying a traffic light according to claim 7, characterized in that: The binarization of the multi-label classification result according to the attribute represented by each output channel to obtain a binarization result also includes: When there are multi-label classification results that do not meet the corresponding preset threshold, the classification results that do not meet the preset threshold are eliminated through a time domain filter; The eliminated multi-label classification results are fused in sequence according to the attribute channels, and the perception results of the traffic light on the target traffic light box are output.

9. The method for identifying a traffic light according to any one of claims 1 to 8, characterized in that: Before outputting the perception result of the traffic light on the target traffic light box, the method further includes: Compare the perception result with the perception result of the previous time before the current time and the perception result of the previous time before that respectively; If the perception results are all the same, then the step of outputting the perception result of the traffic light on the target traffic light box is performed; or If the perception result is the same as the perception result at the previous time and different from the perception result at the time before that, then the step of outputting the perception result of the traffic light on the target traffic light box is performed; or If the perception result is different from the perception result at the previous time and is the same as the perception result at the previous time, the perception result at the previous time is output.

10. A traffic light recognition device, characterized in that: include: An acquisition module is used to acquire the location information of the target traffic light box in the image data of the current road; A cropping module, used for cropping a corresponding target traffic light light box image based on the position information of the target traffic light light box; A determination module, used for inputting the target traffic light light box image into a multi-label classification network for multi-label classification, and obtaining a multi-label classification result of the hierarchical attribute state of each traffic light in the target traffic light light box; The hierarchical attributes are decoupled from each other, and the hierarchical attributes include: orientation, state, shape, color and countdown value; The processing module is used to perform binarization processing on the multi-label classification result to obtain a binarization processing result; and output the binarization processing result as a perception result of the traffic light on the target traffic light box.

11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the traffic light recognition method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method for identifying a traffic light as claimed in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Multi-label learning method for tactile attribute recognition of target object

    CN112668607A

  • Traffic information identification and intelligent driving method and device, equipment and storage medium

    CN112991791A

  • Power grid infrastructure construction risk management and control method based on image recognition and knowledge graph

    CN114694098A

  • Traffic signal lamp classification method and device and electronic equipment

    CN115761699A