Traffic light classification method, classification device, electronic device and storage medium

By clustering and adjusting the detection network model of traffic light images, the problem of low classification accuracy of traffic lights in the prior art is solved, and higher detection and classification accuracy of traffic lights are achieved.

CN114155392BActive Publication Date: 2025-05-13SF TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010926714.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-07
Publication Date
2025-05-13
Estimated Expiration
2040-09-07

AI Technical Summary

Technical Problem

In the prior art, the classification accuracy of traffic lights is not high, mainly because traffic lights account for a small proportion in pictures and the detection field of vision of the detection network model is limited, resulting in missed detection.

Method used

By acquiring the first traffic light image, performing traffic light detection to obtain multiple object detection frames, clustering to obtain the second traffic light image, and then traffic light detection and classification of the second traffic light image is performed. The object detection network model that adjusts the size of the convolution kernel and the number of anchor boxes is used, and the non-maximum suppression algorithm and DIoU distance measurement are combined to accurately detect and classify traffic lights.

Benefits of technology

The classification accuracy of traffic lights is improved, and the detection of clustered images is performed, which reduces the difficulty of small target detection and enhances the detection confidence of traffic lights.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155392B_ABST
    Figure CN114155392B_ABST
Patent Text Reader

Abstract

The present application provides a traffic light classification method, classification device, electronic device and storage medium, the traffic light classification method comprising: obtaining a first traffic light image; performing traffic light detection on the first traffic light image to obtain multiple first target detection frames; clustering the multiple first target detection frames to obtain a second traffic light image; performing traffic light detection on the second traffic light image to obtain multiple second target detection frames; and classifying traffic lights based on the multiple second target detection frames to obtain traffic light classification results. The traffic lights of the present application are easier to detect, and the accuracy of traffic light classification can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and specifically to a traffic light classification method, classification device, electronic device and storage medium. Background Art

[0002] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0003] In the road feature detection and recognition project, we need to detect all traffic lights that appear on the road image. However, since traffic lights are generally long strips and account for a small proportion of the entire image, it is difficult to detect all traffic lights in the same scene using existing mature algorithms. In the existing technology, traffic lights are detected by only one detector. Since the detection network model has a limited detection field of view, and most traffic lights are generally very small in the field of view of the image, the direct detection of a detection network model is prone to false detection and missed detection, and the traffic light classification accuracy is not high.

[0004] That is, the classification accuracy of traffic lights in the prior art is not high. Summary of the invention

[0005] The present application aims to provide a traffic light classification method, a classification device, an electronic device and a storage medium, aiming to solve the problem of low accuracy in traffic light classification in the prior art.

[0006] In one aspect, the present application provides a method for classifying traffic lights, the method comprising:

[0007] Acquire a first traffic light image;

[0008] Performing traffic light detection on the first traffic light image to obtain a plurality of first target detection frames;

[0009] Clustering the multiple first target detection frames to obtain a second traffic light image;

[0010] Performing traffic light detection on the second traffic light image to obtain a plurality of second object detection frames;

[0011] Traffic light classification is performed based on the multiple second target detection frames to obtain a traffic light classification result.

[0012] The performing traffic light detection on the second traffic light image to obtain a plurality of second target detection frames includes:

[0013] Obtain a preset target detection network model, wherein the size of the convolution kernel of the convolution layer of the target detection network model is a*a, where a is an integer greater than 0;

[0014] Adjust the size of the convolution kernel of at least one convolution layer in the target detection network model to b1*h1, to obtain the second traffic light detection model, wherein b1 is not equal to h1, and h1 and b1 are both integers greater than 0;

[0015] Traffic light detection is performed on the second traffic light image based on the second traffic light detection model to obtain multiple second target detection frames.

[0016] The performing traffic light classification based on the multiple second target detection frames to obtain a traffic light classification result includes:

[0017] The size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b2*h2 to obtain a third traffic light detection model, wherein when b1 is less than h1, b2 is set to be greater than h2, and when b1 is greater than h1, b2 is set to be less than h2, and both h2 and b2 are integers greater than 0;

[0018] Performing traffic light detection on the second traffic light image based on the third traffic light detection model to obtain a plurality of third target detection frames;

[0019] Traffic light classification is performed based on the multiple second target detection frames and the multiple third target detection frames to obtain a traffic light classification result.

[0020] The performing traffic light classification based on the multiple second target detection frames and the multiple third target detection frames to obtain a traffic light classification result includes:

[0021] Performing traffic light detection on the second traffic light image based on the target detection network model to obtain a plurality of fourth target detection frames;

[0022] Traffic light classification is performed based on the multiple second target detection frames, the multiple third target detection frames, and the multiple fourth target detection frames to obtain a traffic light classification result.

[0023] The step of adjusting the size of the convolution kernel of at least one convolutional layer in the target detection network model to b1*h1 to obtain the second traffic light detection model includes:

[0024] The size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b1*h1, and the number of anchor frames in the target detection network model is reduced to a first preset value to obtain the second traffic light detection model.

[0025] The step of clustering the plurality of first target detection frames to obtain a second traffic light image includes:

[0026] Obtaining prediction confidences of the multiple first target detection boxes;

[0027] Filtering a preset number of first target detection frames from the multiple first target detection frames as fifth target detection frames based on the prediction confidences of the multiple first target detection frames;

[0028] The fifth target detection frame is clustered to obtain a second traffic light image.

[0029] The performing traffic light classification based on the multiple second target detection frames to obtain a traffic light classification result includes:

[0030] Removing redundant frames in the multiple second target detection frames based on a non-maximum suppression algorithm to obtain a target prediction frame, wherein the non-maximum suppression algorithm uses DIoU as a distance metric;

[0031] Traffic light classification is performed based on the target predicted bounding box to obtain a traffic light classification result.

[0032] The performing traffic light detection on the second traffic light image based on the second traffic light detection model to obtain multiple second target detection frames includes:

[0033] enlarging the second traffic light image according to a preset ratio to obtain an enlarged second traffic light image;

[0034] The enlarged second traffic light image is input into the second traffic light detection model to perform traffic light detection, and a plurality of second target detection frames are obtained.

[0035] The traffic light classification is performed based on the target prediction border to obtain the traffic light classification result, including:

[0036] Get the captured original traffic light image;

[0037] The original traffic lights are labeled based on the number of lamps on the traffic lights, the patterns displayed on the traffic lights, and the colors presented by the traffic lights to obtain a traffic light training set;

[0038] Using the traffic light training set to train a traffic light classification model to obtain a trained traffic light classification model;

[0039] The target prediction bounding box is input into the trained traffic light classification model to obtain a traffic light classification result.

[0040] In one aspect, the present application provides a traffic light classification device, the traffic light classification device comprising:

[0041] An acquisition unit, configured to acquire a first traffic light image;

[0042] A first detection unit, configured to perform traffic light detection on the first traffic light image to obtain a plurality of first target detection frames;

[0043] A clustering unit, configured to cluster the plurality of first target detection frames to obtain a second traffic light image;

[0044] A second detection unit, configured to perform traffic light detection on the second traffic light image to obtain a plurality of second target detection frames;

[0045] A classification unit is used to classify traffic lights based on the multiple second target detection frames to obtain a traffic light classification result.

[0046] The second detection unit is further used to obtain a preset target detection network model, wherein the size of the convolution kernel of the convolution layer of the target detection network model is a*a, where a is an integer greater than 0;

[0047] Adjust the size of the convolution kernel of at least one convolution layer in the target detection network model to b1*h1, to obtain the second traffic light detection model, wherein b1 is not equal to h1, and h1 and b1 are both integers greater than 0;

[0048] Traffic light detection is performed on the second traffic light image based on the second traffic light detection model to obtain multiple second target detection frames.

[0049] The classification unit is further used to adjust the size of the convolution kernel of at least one convolution layer in the target detection network model to b2*h2, so as to obtain a third traffic light detection model, wherein when b1 is less than h1, b2 is set to be greater than h2, and when b1 is greater than h1, b2 is set to be less than h2, and both h2 and b2 are integers greater than 0;

[0050] Performing traffic light detection on the second traffic light image based on the third traffic light detection model to obtain a plurality of third target detection frames;

[0051] Traffic light classification is performed based on the multiple second target detection frames and the multiple third target detection frames to obtain a traffic light classification result.

[0052] The classification unit is further used to perform traffic light detection on the second traffic light image based on the target detection network model to obtain a plurality of fourth target detection frames;

[0053] Traffic light classification is performed based on the multiple second target detection frames, the multiple third target detection frames, and the multiple fourth target detection frames to obtain a traffic light classification result.

[0054] Among them, the second detection unit is also used to adjust the size of the convolution kernel of at least one convolutional layer in the target detection network model to b1*h1, reduce the number of anchor frames in the target detection network model to a first preset value, and obtain the second traffic light detection model.

[0055] The clustering unit is further used to obtain prediction confidences of the multiple first target detection boxes;

[0056] Filtering a preset number of first target detection frames from the multiple first target detection frames as fifth target detection frames based on the prediction confidences of the multiple first target detection frames;

[0057] The fifth target detection frame is clustered to obtain a second traffic light image.

[0058] The classification unit is further used to remove redundant frames in the multiple second target detection frames based on a non-maximum suppression algorithm to obtain a target prediction frame, wherein the non-maximum suppression algorithm uses DIoU as a distance metric;

[0059] Traffic light classification is performed based on the target predicted bounding box to obtain a traffic light classification result.

[0060] The second detection unit is further configured to amplify the second traffic light image according to a preset ratio to obtain an amplified second traffic light image;

[0061] The enlarged second traffic light image is input into the second traffic light detection model to perform traffic light detection, and a plurality of second target detection frames are obtained.

[0062] Wherein, the classification unit is also used to obtain the captured original traffic light image;

[0063] The original traffic lights are labeled based on the number of lamps on the traffic lights, the patterns displayed on the traffic lights, and the colors presented by the traffic lights to obtain a traffic light training set;

[0064] Using the traffic light training set to train a traffic light classification model to obtain a trained traffic light classification model;

[0065] The target prediction bounding box is input into the trained traffic light classification model to obtain a traffic light classification result.

[0066] In one aspect, the present application further provides an electronic device, the electronic device comprising:

[0067] one or more processors;

[0068] Memory; and

[0069] One or more applications, wherein the one or more applications are stored in the memory and are configured to be executed by the processor to implement the traffic light classification method described in any one of the first aspects.

[0070] In one aspect, the present application further provides an electronic device, the electronic device comprising:

[0071] one or more processors;

[0072] Memory; and

[0073] One or more applications, wherein the one or more applications are stored in the memory and are configured to be executed by the processor to implement the traffic light classification method described in any one of the first aspects.

[0074] On the one hand, the present application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is loaded by a processor to execute the steps in the traffic light classification method described in any one of the first aspects.

[0075] The present application provides a method for classifying traffic lights. Traffic light detection is first performed on a first traffic light image to obtain multiple first target detection frames, thereby preliminarily determining the position of the traffic light. The multiple first target detection frames are then clustered to obtain a second traffic light image. The proportion of traffic lights on the second traffic light image is greater than the proportion of traffic lights in the first traffic light image. Therefore, traffic light detection on the second traffic light image is no longer small target detection, and traffic lights are more easily detected, thereby improving the accuracy of traffic light classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0077] Figure 1 A schematic diagram of a scene of a traffic light classification system provided in an embodiment of the present application;

[0078] Figure 2This is a flow chart of an embodiment of a method for classifying traffic lights provided in an embodiment of the present application;

[0079] Figure 3 is a flow chart of another embodiment of the method for classifying traffic lights provided in an embodiment of the present application;

[0080] Figure 4 It is a schematic structural diagram of an embodiment of a traffic light classification device provided in an embodiment of the present application;

[0081] Figure 5 It is a schematic diagram of the structure of an embodiment of an electronic device provided in the embodiments of the present application. DETAILED DESCRIPTION

[0082] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0083] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the accompanying drawings, which are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0084] In this application, the word "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described in this application as "exemplary" is not necessarily to be construed as being preferred or advantageous over other embodiments. The following description is given to enable any technician in the field to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that a person of ordinary skill in the art can recognize that the present application can be implemented without using these specific details. In other instances, well-known structures and processes will not be elaborated in detail to avoid obscuring the description of the present application with unnecessary details. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest scope consistent with the principles and features disclosed in the present application.

[0085] It should be noted that since the method of the embodiment of the present application is executed in an electronic device, the processing objects of each electronic device exist in the form of data or information. For example, time is actually time information. It can be understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data for processing by the electronic device. The details will not be elaborated here.

[0086] The embodiments of the present application provide a traffic light classification method, a classification device, an electronic device, and a storage medium, which are described in detail below.

[0087] See also Figure 1 , Figure 1 A schematic diagram of a traffic light classification system provided in an embodiment of the present application is provided. The traffic light classification system may include an electronic device 100, in which a traffic light classification device is integrated. Figure 1 Electronic devices in.

[0088] In the embodiment of the present application, the electronic device 100 may be an independent server, or a server network or server cluster composed of servers. For example, the electronic device 100 described in the embodiment of the present application includes but is not limited to a computer, a network host, a single network server, a plurality of network server sets or a cloud server composed of a plurality of servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.

[0089] Those skilled in the art will understand that Figure 1 The application environment shown in the figure is only one application scenario of the present application solution and does not constitute a limitation on the application scenario of the present application solution. Other application environments may also include Figure 1 More or fewer electronic devices as shown in, e.g. Figure 1Only one electronic device is shown. It can be understood that the traffic light classification system may also include one or more other services, which are not specifically limited here.

[0090] In addition, if Figure 1 As shown, the traffic light classification system may further include a memory 200 for storing data, such as training data.

[0091] It should be noted that Figure 1 The scenario diagram of the traffic light classification system shown is merely an example. The traffic light classification system and scenario described in the embodiment of the present application are intended to more clearly illustrate the technical solution of the embodiment of the present application, and do not constitute a limitation on the technical solution provided in the embodiment of the present application. A person of ordinary skill in the art can appreciate that with the evolution of the traffic light classification system and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is equally applicable to similar technical problems.

[0092] First, a traffic light classification method is provided in an embodiment of the present application. The execution subject of the traffic light classification method is a traffic light classification device. The traffic light classification device is applied to an electronic device. The traffic light classification method includes:

[0093] Acquire a first traffic light image;

[0094] Performing traffic light detection on the first traffic light image to obtain a plurality of first target detection frames;

[0095] Clustering the multiple first target detection frames to obtain a second traffic light image;

[0096] Performing traffic light detection on the second traffic light image to obtain a plurality of second object detection frames;

[0097] Traffic light classification is performed based on the multiple second target detection frames to obtain a traffic light classification result.

[0098] See also Figure 2 , Figure 2 FIG. 1 is a flow chart of an embodiment of a method for classifying traffic lights provided in an embodiment of the present application. Figure 2 As shown, the classification method of the traffic light includes:

[0099] S201: Acquire a first traffic light image.

[0100] In the embodiment of the present application, there are two types of traffic lights. The ones for motor vehicles are called motor vehicle lights, which usually refer to signal lights composed of red, yellow and green lights used to direct traffic. When the green light is on, vehicles are allowed to pass. When the yellow light is flashing, vehicles that have crossed the stop line can continue to pass; vehicles that have not passed should slow down to the stop line and wait. When the red light is on, vehicles are prohibited from passing. The ones for pedestrians are called crosswalk lights, which usually refer to signal lights composed of red and green lights used to direct traffic. Red light means stop and green light means go. The first traffic light image can be a picture of the traffic light taken by a camera, a driving recorder, etc. There may be one or more traffic lights in the first traffic light image.

[0101] S202: Perform traffic light detection on the first traffic light image to obtain a plurality of first target detection frames.

[0102] In an embodiment of the present application, the first traffic light image is input into the first traffic light detection model for traffic light detection, and multiple first target detection frames are obtained. Specifically, the first traffic light detection model is a traffic light detection model trained with a traffic light training set. Among them, the first traffic light detection model can be SSD (Single Shot MultiBox Detector) or YOLOV3-tiny. SSD is a paper in ICCV in 2016 and is the main target detection algorithm so far. The main advantages of the algorithm: faster than Faster-Rcnn and higher accuracy than Yolo. While taking into account speed, the accuracy is also very high. In order to improve the accuracy, the results are predicted under different feature maps, and the feature pyramid prediction method is used. The END-TO-END training method is adopted, and the classification results are very accurate even for pictures with relatively small resolution. YOLO (You Only Look Once) is an object recognition and positioning algorithm based on deep neural networks. Its biggest feature is that it runs very fast and can be used in real-time systems. Now YOLO has developed to version v3. YOLOV3-tiny removes some feature layers based on YOLOV3 and only retains two independent prediction branches.

[0103] According to the design of YOLO, the input image is divided into a 7*7 grid, and the 7*7 in the output tensor corresponds to the 7*7 grid of the input image. Or we regard the 7*7*30 tensor as 7*7=49 30-dimensional vectors, that is, each grid in the input image corresponds to a 30-dimensional vector output. The first traffic light image is input into the first traffic light detection model for traffic light detection. The first traffic light detection model predicts a fixed number of first target detection frames and the prediction confidence of the first target detection frame for each grid in the first traffic light image, so multiple first target detection frames can be obtained.

[0104] S203: Cluster the multiple first target detection frames to obtain a second traffic light image.

[0105] In an embodiment of the present application, a k-means clustering algorithm is used to cluster multiple first target detection frames to obtain a second traffic light image. The k-means clustering algorithm is an iterative clustering analysis algorithm, and its steps are: pre-dividing the data into K groups, then randomly selecting K objects as initial cluster centers, and then calculating the distance between each object and each seed cluster center, and assigning each object to the cluster center closest to it. The cluster centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the cluster center of the cluster will be recalculated based on the existing objects in the cluster. This process will be repeated until a certain termination condition is met. The termination condition can be that no (or a minimum number) of objects are reassigned to different clusters, no (or a minimum number) of cluster centers change again, and the sum of squared errors is locally minimized.

[0106] Specifically, multiple first target detection frames are clustered to obtain cluster detection frames, and images in the cluster detection frames are extracted from the first traffic light image to obtain the second traffic light image.

[0107] Generally speaking, the proportion of traffic lights in the first traffic light image is small, and it is a small target detection, which is not easy to be detected. Therefore, the first traffic light image is subjected to preliminary traffic light detection, and the first target detection frame obtained is likely to be the location of the traffic light. The second traffic light image is obtained by clustering the first target detection frame. The proportion of traffic lights on the second traffic light image is greater than the proportion of traffic lights in the first traffic light image. Therefore, traffic light detection on the second traffic light image is no longer a small target detection, and it is easier to be detected.

[0108] In a specific embodiment, clustering a plurality of first object detection frames to obtain a second traffic light image includes:

[0109] (1) Obtain the prediction confidence of multiple first target detection boxes.

[0110] The first traffic light detection model performs traffic light detection on the first traffic light image and outputs multiple first target detection frames and prediction confidences of the multiple first target detection frames. The prediction confidences of the multiple first target detection frames are obtained from the output results of the first traffic light detection model. For example, target detection frame A, prediction confidence 0.7; target detection frame B, prediction confidence 0.6; target detection frame C, prediction confidence 0.5.

[0111] (2) Based on the prediction confidences of the multiple first target detection frames, a preset number of first target detection frames are selected from the multiple first target detection frames as the fifth target detection frames.

[0112] The preset number may be 50, 100, 200, etc., and is set according to specific circumstances. The preset number is less than the number of the multiple first target detection frames. Specifically, the multiple first target detection frames are sorted from high to low according to the prediction confidence, and the preset number of first target detection frames with the top sorting positions are screened out to obtain the preset number of fifth target detection frames.

[0113] For example, target detection frame A has a prediction confidence of 0.7; target detection frame B has a prediction confidence of 0.6; and target detection frame C has a prediction confidence of 0.5. The preset number is 2, and the fifth target detection frame of the preset number is target detection frame A and target detection frame B.

[0114] (3) Cluster the fifth target detection frame to obtain a second traffic light image.

[0115] Specifically, a preset number of fifth target detection frames are clustered using a k-means clustering algorithm to obtain a second traffic light image. A preset number of fifth target detection frames are obtained by clustering after a portion of the first target detection frames are selected to obtain a second traffic light image.

[0116] Preferably, a preset number of fifth target detection frames are clustered using a k-means clustering algorithm to obtain a square second traffic light image. The side length of the second traffic light image is a variable parameter, for example, set to 320. Specifically, a preset number of fifth target detection frames are clustered to obtain a square cluster detection frame, and the image in the square cluster detection frame is extracted from the first traffic light image to obtain a square second traffic light image.

[0117] S204: Perform traffic light detection on the second traffic light image to obtain multiple second target detection frames.

[0118] It should be noted that the second traffic light image obtained by clustering may be one or more. When there are more than one second traffic light images, traffic light detection is performed on the more than one second traffic light images respectively to obtain more than one second target detection frames.

[0119] In the embodiment of the present application, traffic light detection is performed on the second traffic light image to obtain multiple second target detection frames, including:

[0120] (1) Obtain the preset target detection network model.

[0121] The preset target detection network model may be YOLOV1, YOLOV2, YOLOV3, etc., and this application is described with the preset target detection network model being YOLOV3. The size of the convolution kernel of the convolution layer of the target detection network model is a*a, where a is an integer greater than 0. The first traffic light detection model may be a target detection network model, or the first traffic light detection model may be a detection model having fewer detection layers than the target detection network model.

[0122] (2) The size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b1*h1 to obtain a second traffic light detection model.

[0123] In a specific embodiment, the size of the convolution kernel of at least one convolution layer in the target detection network model is adjusted to b1*h1 to obtain a second traffic light detection model, wherein b1 is not equal to h1, and h1 and b1 are both integers greater than 0. The convolution kernel size of the convolution layer in YOLOV3 is generally 3*3. For example, the size of the convolution kernel of at least one convolution layer in the target detection network model is adjusted to 1*3 or 3*1. That is, the second traffic light detection model uses a rectangular convolution kernel, which has a better detection effect on detecting narrow and long traffic lights.

[0124] As the number and scale of the output feature maps change, the size of the anchor box also needs to be adjusted accordingly. YOLOV2 has begun to use the k-means clustering algorithm to obtain the size of the anchor box. YOLOV3 continues this method, setting 3 anchor boxes for each downsampling scale, and clustering a total of 9 sizes of anchor boxes. In the COCO dataset, these 9 anchor boxes are: (10x13), (16x30), (33x23), (30x61), (62x45), (59x119), (116x90), (156x198), (373x326).

[0125] In another specific embodiment, the size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b1*h1, and the number of anchor boxes in the target detection network model is reduced to a first preset value to obtain a second traffic light detection model. The first preset value is set according to the specific situation. For example, the 9 anchor boxes in YOLOV3 are reduced to 3, and the aspect ratios are: 1:1, 1:2, 1:3.

[0126] Further, the number of detection layers in the target detection network model is reduced to a second preset value to obtain a second traffic light detection model. The second preset value is set according to the specific situation. The number of detection layers in YOLOV3 is generally 3. For example, the number of detection layers in the target detection network model is reduced to 2.

[0127] (3) Perform traffic light detection on the second traffic light image based on the second traffic light detection model to obtain multiple second target detection frames.

[0128] In an embodiment of the present application, traffic light detection is performed on the second traffic light image based on the second traffic light detection model to obtain multiple second target detection frames, including: enlarging the second traffic light image at a preset ratio to obtain an enlarged second traffic light image; inputting the enlarged second traffic light image into the second traffic light detection model to perform traffic light detection to obtain multiple second target detection frames.

[0129] It should be noted that after obtaining the second traffic light detection model, the second traffic light detection model can be stored, and when traffic light detection is performed on the second traffic light image based on the second traffic light detection model to obtain multiple second target detection frames, the stored second traffic light detection model can be directly read, and traffic light detection is performed on the second traffic light image based on the second traffic light detection model to obtain multiple second target detection frames.

[0130] In the embodiment of the present application, the preset ratio is set according to the specific situation, such as 1:2; 1:1.5, etc. Preferably, the preset ratio is 1:1.5. Taking into account the balance between computational efficiency and accuracy, the second traffic light image is magnified 1.5 times and then placed in the second-order detection model to detect the accurate target. This method saves computing costs (no need to run the entire image, only need to detect in the small image) by enlarging the target area generated first, while also improving the detection effect of small targets.

[0131] S205 , classify traffic lights based on the multiple second target detection frames to obtain traffic light classification results.

[0132] In the embodiment of the present application, traffic light classification is performed based on multiple second target detection frames to obtain traffic light classification results, including:

[0133] (1) Based on the non-maximum suppression algorithm, redundant frames in multiple second target detection frames are removed to obtain the target prediction frame, where the non-maximum suppression algorithm uses DIoU as the distance metric.

[0134] In the embodiment of the present application, the core idea of ​​the non-maximum suppression algorithm (NMS) is to select the one with the highest score as the output, remove the ones overlapping with the output, and repeat this process until all the candidates are processed. The non-maximum suppression algorithm is widely used in target detection algorithms. Its purpose is to eliminate redundant boxes and find the best object detection position.

[0135] The conventional non-maximum suppression algorithm uses IoU as the distance metric. IoU is what we call the intersection over union ratio. It is the most commonly used indicator in target detection. It can reflect the detection effect of the predicted detection box and the real detection box.

[0136] In addition, since the target is narrow, long and small, the IOU algorithm widely used to pair the target box and the real box is easily invalid. To solve this problem, we modify the IOU in the non-maximum suppression algorithm to DIOU as follows:

[0137]

[0138] where b, b gt Represent the center points of the second target detection frame and the real frame respectively, p represents the Euclidean distance between the two center points, and c represents the diagonal length of the closure of the two frames. This method can effectively avoid the mismatch between the target frame and the target that may occur in IOU.

[0139] (2) Traffic lights are classified based on the target prediction bounding box to obtain the traffic light classification result.

[0140] In a specific embodiment, the target prediction bounding box is input into the trained traffic light classification model to obtain the traffic light classification result. The traffic light classification model can be an Efficientnet classification network model. The basic network architecture of Efficientnet (Rethinking Model Scaling for Convolutional Neural Networks) is designed by using neural architecture search. A simple and efficient composite coefficient is used to weigh the network depth, width and input image resolution.

[0141] Specifically, before classifying traffic lights based on the target predicted bounding box to obtain the traffic light classification result, the method also includes: obtaining the original captured traffic light image, annotating the original traffic light based on the number of lamps on the traffic light, the pattern displayed on the traffic light, and the color presented by the traffic light, to obtain a traffic light training set, and using the traffic light training set to train a traffic light classification model to obtain a trained traffic light classification model.

[0142] Due to the differences in traffic light systems in different places, it is a luxury to use a classification model to divide traffic lights into lane lights, sidewalks and other subcategories without a standard. We have summarized a three-level granularity classification specification through a large amount of data analysis in multiple cities, trying to solve this problem and have achieved initial success. This specification actually relies on three conditions to determine the subcategories of traffic lights, namely the number of lamps on the traffic light, the pattern displayed on the traffic light, and the color of the traffic light. Through this classification, we divide traffic lights into the following major categories, namely round lights, straight lights, left turn lights, right turn lights, U-turn lights, sidewalk lights, bicycle lights, and lane lights. Effectively reduce the difficulty of labeling and improve the actual labeling efficiency.

[0143] When a traffic light is on, generally only part of the area is lit, and the other areas are invalid areas. In addition, specific classification requires observing the pattern displayed in the lit area, which is obviously quite different from the object classification under normal circumstances.

[0144] Therefore, in order to improve the training effect of the traffic light classification model, in a specific embodiment, the original traffic light is labeled based on the number of lamps on the traffic light, the pattern displayed on the traffic light, and the color presented by the traffic light, including:

[0145] The original traffic light image is preprocessed to obtain a preprocessed traffic light image, and the preprocessed traffic light image is labeled based on the number of lamps on the traffic light, the pattern displayed on the traffic light, and the color presented by the traffic light.

[0146] There are two ways to preprocess the original traffic light image: (1) centering the rectangular original traffic light image and complementing it to obtain a square traffic light image; (2) cropping the rectangular original traffic light image into a square traffic light image. This greatly improves the efficiency of data augmentation.

[0147] See also Figure 3 , Figure 3 FIG. 2 is another flow chart of a method for classifying traffic lights provided in an embodiment of the present application. Figure 3 As shown, the classification method of the traffic light includes:

[0148] S301: Acquire a first traffic light image.

[0149] In the embodiment of the present application, S301 is the same as S201, and the implementation of S302 may refer to the specific implementation process of S201, which will not be repeated here.

[0150] S302: Perform traffic light detection on the first traffic light image to obtain multiple first target detection frames.

[0151] In the embodiment of the present application, S302 is the same as S202. The specific implementation of S302 can refer to the specific implementation process of S202, which will not be repeated here.

[0152] S303: Cluster the multiple first target detection frames to obtain a second traffic light image.

[0153] In the embodiment of the present application, S303 is the same as S203. The specific implementation of S303 may refer to the specific implementation process of S203, which will not be repeated here.

[0154] S304: Perform traffic light detection on the second traffic light image to obtain multiple second target detection frames.

[0155] In the embodiment of the present application, S304 is the same as S204. The implementation of S304 may refer to the specific implementation process of S204, which will not be repeated here.

[0156] It should be noted that, in S304, the size of the convolution kernel of at least one convolution layer in the second traffic light detection model is b1*h1, and b1 is smaller than h1. For example, the size of the convolution kernel in the second convolution layer and the third convolution layer in the second traffic light detection model is b1*h1. Preferably, b1*h1 is 1*3. In other embodiments, b1 may be larger than h1.

[0157] S305. Adjust the size of the convolution kernel of at least one convolutional layer in the target detection network model to b2*h2 to obtain a third traffic light detection model.

[0158] When b1 is smaller than h1, b2 is set to be larger than h2; when b1 is larger than h1, b2 is set to be smaller than h2; both h2 and b2 are integers larger than 0.

[0159] For example, the size of the convolution kernel in the second convolution layer and the third convolution layer in the second traffic light detection model is b2*h2. Preferably, b2*h2 is 3*1.

[0160] Further, the size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b2*h2, and the number of anchor boxes in the target detection network model is reduced to a first preset value, thereby obtaining a third traffic light detection model. The first preset value is set according to the specific situation. For example, the 9 anchor boxes in YOLOV3 are reduced to 3, and the aspect ratios are: 1:1, 1:2, and 1:3 respectively.

[0161] In other embodiments, the size of the convolution kernel of at least one convolution layer of the second traffic light detection model is b1*h1, b1 is greater than h1, and the size of the convolution kernel of at least one convolution layer of the third traffic light detection model is b2*h2, b2 is less than h2.

[0162] S306: Perform traffic light detection on the second traffic light image based on the third traffic light detection model to obtain multiple third target detection frames.

[0163] Specifically, the second traffic light image is enlarged according to a preset ratio to obtain an enlarged second traffic light image; the enlarged second traffic light image is input into a third traffic light detection model to perform traffic light detection to obtain multiple third target detection frames.

[0164] S307: Perform traffic light detection on the second traffic light image based on the target detection network model to obtain multiple fourth target detection frames.

[0165] Specifically, the second traffic light image is enlarged according to a preset ratio to obtain an enlarged second traffic light image; the enlarged second traffic light image is input into the target detection network model to perform traffic light detection to obtain multiple fourth target detection frames.

[0166] Further, the number of anchor frames of the target detection network model is reduced to 3, and a fourth traffic light detection model is obtained. The second traffic light image is enlarged according to a preset ratio to obtain an enlarged second traffic light image; the enlarged second traffic light image is input into the fourth traffic light detection model for traffic light detection, and multiple fourth target detection frames are obtained.

[0167] S308. Classify traffic lights based on the multiple second target detection frames, the multiple third target detection frames, and the multiple fourth target detection frames to obtain a traffic light classification result.

[0168] Specifically, based on the non-maximum suppression algorithm, redundant frames in multiple third target detection frames and multiple fourth target detection frames are removed to obtain target prediction frames, wherein the non-maximum suppression algorithm uses DIoU as a distance metric. Traffic light classification is performed based on the target prediction frame to obtain a traffic light classification result.

[0169] In other embodiments, S304, S306, and S307 may be performed in different orders or simultaneously, and this application does not limit this.

[0170] In order to better implement the traffic light classification method in the embodiment of the present application, based on the traffic light classification method, the embodiment of the present application also provides a traffic light classification device, such as Figure 4 As shown, Figure 4 : is a schematic diagram of the structure of an embodiment of a traffic light classification device provided in an embodiment of the present application, and the traffic light classification device includes:

[0171] An acquisition unit 401 is used to acquire a first traffic light image;

[0172] A first detection unit 402, configured to perform traffic light detection on the first traffic light image to obtain a plurality of first target detection frames;

[0173] A clustering unit 403, configured to cluster the plurality of first target detection frames to obtain a second traffic light image;

[0174] A second detection unit 404 is used to perform traffic light detection on the second traffic light image to obtain a plurality of second object detection frames;

[0175] The classification unit 405 is used to classify the traffic lights based on the multiple second target detection frames to obtain a traffic light classification result.

[0176] The second detection unit 404 is further used to obtain a preset target detection network model;

[0177] The size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b1*h1 to obtain a second traffic light detection model, wherein b1 is not equal to h1, and h1 and b1 are both integers greater than 0.

[0178] The classification unit 405 is further used to adjust the size of the convolution kernel of at least one convolution layer in the target detection network model to b2*h2, so as to obtain a third traffic light detection model, wherein when b1 is less than h1, b2 is set to be greater than h2, and when b1 is greater than h1, b2 is set to be less than h2, and both h2 and b2 are integers greater than 0;

[0179] Performing traffic light detection on the second traffic light image based on the third traffic light detection model to obtain multiple third target detection frames;

[0180] Traffic light classification is performed based on the multiple second target detection frames and the multiple third target detection frames to obtain a traffic light classification result.

[0181] The classification unit 405 is further configured to perform traffic light detection on the second traffic light image based on the target detection network model to obtain a plurality of fourth target detection frames;

[0182] Traffic light classification is performed based on the multiple second target detection frames, the multiple third target detection frames, and the multiple fourth target detection frames to obtain a traffic light classification result.

[0183] Among them, the second detection unit 404 is also used to adjust the size of the convolution kernel of at least one convolutional layer in the target detection network model to b1*h1, reduce the number of anchor frames in the target detection network model to a first preset value, and obtain a second traffic light detection model.

[0184] The clustering unit 403 is further used to obtain prediction confidences of multiple first target detection boxes;

[0185] Filtering a preset number of first target detection frames from the multiple first target detection frames based on the prediction confidences of the multiple first target detection frames as fifth target detection frames;

[0186] The fifth target detection frame is clustered to obtain a second traffic light image.

[0187] The classification unit 405 is further used to remove redundant frames in the plurality of second target detection frames based on a non-maximum suppression algorithm to obtain a target prediction frame, wherein the non-maximum suppression algorithm uses DIoU as a distance metric;

[0188] Traffic lights are classified based on the target prediction bounding box to obtain the traffic light classification result.

[0189] The second detection unit 404 is further used to amplify the second traffic light image according to a preset ratio to obtain an amplified second traffic light image;

[0190] The enlarged second traffic light image is input into the second traffic light detection model to perform traffic light detection, and a plurality of second target detection frames are obtained.

[0191] Wherein, the classification unit 405 is also used to obtain the captured original traffic light image;

[0192] The original traffic lights are labeled based on the number of lamps on the traffic lights, the patterns displayed on the traffic lights, and the colors presented by the traffic lights to obtain a traffic light training set;

[0193] The traffic light classification model is trained using the traffic light training set to obtain a trained traffic light classification model;

[0194] The target prediction bounding box is input into the trained traffic light classification model to obtain the traffic light classification result.

[0195] The present application also provides an electronic device that integrates any of the traffic light classification devices provided in the present application. Figure 5 As shown, it shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application, specifically:

[0196] The electronic device may include components such as a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, a power supply 603, and an input unit 604. Those skilled in the art will appreciate that the electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. Among them:

[0197] The processor 601 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 602, and calling data stored in the memory 602, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 601.

[0198] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.

[0199] The electronic device also includes a power supply 603 for supplying power to each component. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, so as to manage charging, discharging, power consumption and other functions through the power management system. The power supply 603 can also include one or more DC or AC power supplies, recharging systems, power failure detection circuits, power converters or inverters, power status indicators and other arbitrary components.

[0200] The electronic device may further include an input unit 604, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0201] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 601 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 will run the application programs stored in the memory 602, thereby realizing various functions, as follows:

[0202] Acquire a first traffic light image;

[0203] Performing traffic light detection on the first traffic light image to obtain a plurality of first target detection frames;

[0204] Clustering the multiple first target detection frames to obtain a second traffic light image;

[0205] Performing traffic light detection on the second traffic light image to obtain a plurality of second object detection frames;

[0206] Traffic light classification is performed based on the multiple second target detection frames to obtain a traffic light classification result.

[0207] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0208] To this end, the embodiment of the present application provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the traffic light classification methods provided in the embodiment of the present application. For example, the computer program loaded by the processor can execute the following steps:

[0209] Acquire a first traffic light image;

[0210] Performing traffic light detection on the first traffic light image to obtain a plurality of first target detection frames;

[0211] Clustering the multiple first target detection frames to obtain a second traffic light image;

[0212] Performing traffic light detection on the second traffic light image to obtain a plurality of second object detection frames;

[0213] Traffic light classification is performed based on the multiple second target detection frames to obtain a traffic light classification result.

[0214] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of other embodiments above, and will not be repeated here.

[0215] In specific implementation, the above units or structures can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units or structures can refer to the previous method embodiments, which will not be repeated here.

[0216] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0217] The above is a detailed introduction to a traffic light classification method, classification device, electronic device and storage medium provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A traffic light classification method, characterized in that: The traffic light classification method includes: Acquire a first traffic light image; Performing traffic light detection on the first traffic light image to obtain a plurality of first target detection frames; Clustering the multiple first target detection frames to obtain a second traffic light image; Performing traffic light detection on the second traffic light image to obtain a plurality of second object detection frames; Performing traffic light classification based on the multiple second target detection frames to obtain a traffic light classification result; The performing traffic light detection on the second traffic light image to obtain a plurality of second target detection frames includes: Obtain a preset target detection network model, wherein the size of the convolution kernel of the convolution layer of the target detection network model is a*a, where a is an integer greater than 0; Adjust the size of the convolution kernel of at least one convolution layer in the target detection network model to b1*h1, to obtain a second traffic light detection model, wherein b1 is not equal to h1, and h1 and b1 are both integers greater than 0; Performing traffic light detection on the second traffic light image based on the second traffic light detection model to obtain a plurality of second target detection frames; The performing traffic light classification based on the multiple second target detection frames to obtain a traffic light classification result includes: The size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b2*h2 to obtain a third traffic light detection model, wherein when b1 is less than h1, b2 is set to be greater than h2, and when b1 is greater than h1, b2 is set to be less than h2, and both h2 and b2 are integers greater than 0; Performing traffic light detection on the second traffic light image based on the third traffic light detection model to obtain a plurality of third target detection frames; Traffic light classification is performed based on the multiple second target detection frames and the multiple third target detection frames to obtain a traffic light classification result.

2. The method for classifying traffic lights according to claim 1, characterized in that: The performing traffic light classification based on the multiple second target detection frames and the multiple third target detection frames to obtain a traffic light classification result includes: Performing traffic light detection on the second traffic light image based on the target detection network model to obtain a plurality of fourth target detection frames; Traffic light classification is performed based on the multiple second target detection frames, the multiple third target detection frames, and the multiple fourth target detection frames to obtain a traffic light classification result.

3. The traffic light classification method according to any one of claims 1 to 2, characterized in that: The step of adjusting the size of the convolution kernel of at least one convolutional layer in the target detection network model to b1*h1 to obtain the second traffic light detection model includes: The size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b1*h1, and the number of anchor frames in the target detection network model is reduced to a first preset value to obtain the second traffic light detection model.

4. The method for classifying traffic lights according to claim 1, characterized in that: The clustering of the plurality of first target detection frames to obtain a second traffic light image includes: Obtaining prediction confidences of the multiple first target detection boxes; Filtering a preset number of first target detection frames from the multiple first target detection frames as fifth target detection frames based on the prediction confidences of the multiple first target detection frames; The fifth target detection frame is clustered to obtain a second traffic light image.

5. The traffic light classification method according to claim 1, characterized in that: The performing traffic light classification based on the multiple second target detection frames to obtain a traffic light classification result includes: Removing redundant frames in the multiple second target detection frames based on a non-maximum suppression algorithm to obtain a target prediction frame, wherein the non-maximum suppression algorithm uses DIoU as a distance metric; Traffic light classification is performed based on the target predicted bounding box to obtain a traffic light classification result.

6. The method for classifying traffic lights according to claim 3, characterized in that: The performing traffic light detection on the second traffic light image based on the second traffic light detection model to obtain a plurality of second target detection frames includes: enlarging the second traffic light image according to a preset ratio to obtain an enlarged second traffic light image; The enlarged second traffic light image is input into the second traffic light detection model to perform traffic light detection, and a plurality of second target detection frames are obtained.

7. The traffic light classification method according to claim 5, characterized in that: The traffic light classification is performed based on the target prediction border to obtain the traffic light classification result, including: Get the captured original traffic light image; The original traffic lights are labeled based on the number of lamps on the traffic lights, the patterns displayed on the traffic lights, and the colors presented by the traffic lights to obtain a traffic light training set; Using the traffic light training set to train a traffic light classification model to obtain a trained traffic light classification model; The target prediction bounding box is input into the trained traffic light classification model to obtain a traffic light classification result.

8. A traffic light classification device, characterized in that: The traffic light classification device comprises: An acquisition unit, configured to acquire a first traffic light image; A first detection unit, configured to perform traffic light detection on the first traffic light image to obtain a plurality of first target detection frames; A clustering unit, configured to cluster the plurality of first target detection frames to obtain a second traffic light image; A second detection unit, configured to perform traffic light detection on the second traffic light image to obtain a plurality of second target detection frames; a classification unit, configured to classify the traffic lights based on the plurality of second target detection frames to obtain a traffic light classification result; The performing traffic light detection on the second traffic light image to obtain a plurality of second target detection frames includes: Obtain a preset target detection network model, wherein the size of the convolution kernel of the convolution layer of the target detection network model is a*a, where a is an integer greater than 0; Adjust the size of the convolution kernel of at least one convolution layer in the target detection network model to b1*h1, to obtain a second traffic light detection model, wherein b1 is not equal to h1, and h1 and b1 are both integers greater than 0; Performing traffic light detection on the second traffic light image based on the second traffic light detection model to obtain a plurality of second target detection frames; The performing traffic light classification based on the multiple second target detection frames to obtain a traffic light classification result includes: The size of the convolution kernel of at least one convolutional layer in the target detection network model is adjusted to b2*h2 to obtain a third traffic light detection model, wherein when b1 is less than h1, b2 is set to be greater than h2, and when b1 is greater than h1, b2 is set to be less than h2, and both h2 and b2 are integers greater than 0; Performing traffic light detection on the second traffic light image based on the third traffic light detection model to obtain a plurality of third target detection frames; Traffic light classification is performed based on the multiple second target detection frames and the multiple third target detection frames to obtain a traffic light classification result.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and are configured to be executed by the processor to implement the traffic light classification method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the traffic light classification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and device of target detection based on multiple models

    CN108009558A

  • Method and device for identifying traffic light signal, readable medium and electronic equipment

    CN108830199A

  • Semantic traffic signal lamp detection method based on multi-scale attention mechanism network model

    CN110532961A

  • Traffic signal identification method and device, vehicle navigation equipment and unmanned vehicle

    CN110688992A