Signal lamp identification and grouping method and device, electronic equipment and storage medium

By using a recurrent neural network model, especially a bidirectional gated recurrent unit model, combined with corner position heatmaps and embedding vectors, and employing three loss functions, the problem of poor traffic light recognition and matching performance was solved, enabling accurate recognition and grouping of traffic lights in autonomous vehicles.

CN116259041BActive Publication Date: 2026-04-21ZHIDAO NETWORK TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHIDAO NETWORK TECH (BEIJING) CO LTD
Filing Date
2023-03-28
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies have poor traffic light recognition and matching performance, especially in autonomous vehicles where traffic light targets are small, difficult to match, time-consuming, and with unsatisfactory results.

Method used

By employing a recurrent neural network model, particularly a bidirectional gated recurrent unit model, and by identifying the corner location heatmap and embedding vector in traffic light images, combined with three loss functions, the four corner points of the same traffic light and the traffic light regions in different frame images are determined.

Benefits of technology

It achieves accurate identification and grouping of traffic lights, improves identification efficiency, and enables efficient identification of traffic lights in consecutive frame images in autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116259041B_ABST
    Figure CN116259041B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, electronic device, and storage medium for traffic light identification and grouping. The method includes: acquiring multiple consecutive frames of traffic light images to be identified, using them as input to a recurrent neural network model to obtain a positional heatmap and embedding vector for each corner point of each traffic light to be identified in each frame of the image; determining the four corner points of each traffic light to be identified in each frame of the image; and determining the traffic light regions belonging to the same traffic light in the multiple consecutive frames of the image based on the four corner points and embedding vectors of each traffic light to be identified. The embodiments of this application employ four different loss functions to identify corner points, determining the four corner points belonging to the same traffic light in the same frame and the traffic light regions belonging to the same traffic light in different frames. This method can accurately identify traffic lights while also grouping traffic lights in consecutive frames, resulting in good identification performance and high efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, electronic device and storage medium for identifying and grouping traffic lights. Background Technology

[0002] With the development of autonomous driving technology, autonomous vehicles are becoming increasingly automated. For safety and navigation reasons, autonomous vehicles require higher accuracy in recognizing traffic lights.

[0003] In related technical solutions, neural networks are generally used to identify traffic lights in a single road image, and then the Hungarian algorithm or feature point extraction method is used to match the signals in multiple images. Since traffic lights are small targets and difficult to match, this method is time-consuming and the recognition and matching effect is poor. Alternatively, images of consecutive frames are used to identify traffic lights, and then algorithms such as optimal matching lights are used to match the traffic lights. However, this method is not ideal for matching and grouping traffic lights and needs improvement. Summary of the Invention

[0004] To address or partially address the problems existing in related technologies, this application provides a method, apparatus, electronic device, and storage medium for identifying and grouping traffic lights, which can accurately identify traffic lights in consecutive frame images and match and group the same traffic light in different images.

[0005] The first aspect of this application provides a method for identifying and grouping traffic lights, including:

[0006] Acquire multiple consecutive frames of traffic light images to be identified, wherein each traffic light image includes at least one traffic light to be identified;

[0007] The multiple consecutive frames of traffic light images to be identified are used as input to a recurrent neural network model. The corner points of each traffic light to be identified in each frame of the traffic light image are identified to obtain the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image.

[0008] The four corner points of each traffic light to be identified in each frame of the traffic light image are determined based on the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image;

[0009] Based on the four corner points of each traffic light to be identified and the embedding vector, the traffic light regions belonging to the same traffic light in the multiple consecutive frames of traffic light images to be identified are determined.

[0010] In one possible implementation of this application, the recurrent neural network model is a bidirectional gated recurrent unit (GROUP) model. The step of using the multiple consecutive frames of traffic light images to be identified as input to the recurrent neural network model, and identifying the corner points of each traffic light to be identified in each frame of the traffic light image to obtain the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image, includes:

[0011] The multiple consecutive frames of traffic light images to be identified are input into the bidirectional gate loop unit model;

[0012] The corner points of each traffic light to be identified in each traffic light image to be identified are identified, and a heat map of the corner point position is generated based on the position of each corner point of each traffic light to be identified. The heat map of the corner point position consists of 4 images, and each heat map of the corner point position records the corner point position of all traffic lights to be identified in one direction in the multiple consecutive frames of traffic light images to be identified.

[0013] Based on the location heatmap of the corner points, an embedding vector for each corner point of each traffic light to be identified is generated.

[0014] As one possible implementation of this application, in this implementation, determining the four corner points of each traffic light to be identified in each frame of the traffic light image based on the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image includes:

[0015] The first vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image is calculated using a first loss function, wherein the first loss function is:

[0016]

[0017] Where L_pull1 is the value of the first loss function, e k Let e ​​be the embedding vector of the k-th corner point. center It is the average value of the embedding vectors of the four corner points of the same traffic light to be identified;

[0018] The second loss function is used to calculate the second vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image, where the second loss function is:

[0019]

[0020] Where L_push is the value of the second loss function, N is the number of traffic lights to be identified in each frame of the traffic light image, and e centeri e is the average of the embedding vectors of the four corner points of the i-th traffic light to be identified. centerjLet Δ be the average of the embedding vectors of the four corner points of the j-th traffic light to be identified, where Δ is a constant 1;

[0021] Based on the first loss function and the second loss function, the four corner points belonging to the same traffic light to be identified in each frame of traffic light image are determined. Among them, the first vector distance of the four corner points of the same traffic light to be identified in each frame of traffic light image is the closest, while the second vector distance of the four corner points of different traffic lights to be identified in each frame of traffic light image is not the closest.

[0022] As one possible implementation of this application, in this implementation, determining the signal light region belonging to the same signal light in the multiple consecutive frames of signal light images to be identified based on the four corner points of each signal light to be identified and the embedding vector includes:

[0023] The third vector distance between each traffic light to be identified in multiple consecutive frames of traffic light images is calculated using a third loss function, wherein the third loss function is:

[0024]

[0025] Where L_pull2 is the value of the third loss function, and X and Y represent that there are X traffic lights to be identified in each of the Y consecutive frames of traffic light images to be identified. Let be the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the y-th traffic light image. This is the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the first traffic light image;

[0026] Based on the third loss function, the signal light regions belonging to the same signal light in the multiple consecutive frames of signal light images to be identified are determined, wherein the third vector distance between the corner points belonging to the same signal light in different signal light images is the smallest.

[0027] In one possible implementation of this application, generating the embedding vector of each corner point of each traffic light to be identified based on the positional heatmap of the corner points includes:

[0028] Obtain the pixel values ​​of each corner point of each traffic light to be identified in the position heatmap of the corner points;

[0029] An embedding vector for each corner point is generated based on the pixel values.

[0030] A second aspect of this application provides a device for identifying and grouping traffic lights, comprising:

[0031] An image acquisition module is used to acquire multiple consecutive frames of traffic light images to be identified, wherein each traffic light image includes at least one traffic light to be identified;

[0032] The model recognition module is used to take the multiple consecutive frames of traffic light images to be recognized as input to the recurrent neural network model, and to recognize the corner points of each traffic light to be recognized in each frame of traffic light images, so as to obtain the position heatmap and embedding vector of each corner point of each traffic light to be recognized in each frame of traffic light images.

[0033] The traffic light recognition module is used to determine the four corner points of each traffic light to be recognized in each frame of traffic light image based on the position heatmap and embedding vector of each corner point of each traffic light to be recognized in each frame of traffic light image;

[0034] The traffic light grouping module is used to determine the traffic light regions belonging to the same traffic light in the multiple consecutive frames of traffic light images to be identified, based on the four corner points of each traffic light to be identified and the embedding vector.

[0035] In one possible implementation of this application, the traffic light recognition module is used for:

[0036] The first vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image is calculated using a first loss function, wherein the first loss function is:

[0037]

[0038] Where L_pull1 is the value of the first loss function, e k Let e ​​be the embedding vector of the k-th corner point. center It is the average value of the embedding vectors of the four corner points of the same traffic light to be identified;

[0039] The second loss function is used to calculate the second vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image, where the second loss function is:

[0040]

[0041] Where L_push is the value of the second loss function, N is the number of traffic lights to be identified in each frame of the traffic light image, and e centeri e is the average of the embedding vectors of the four corner points of the i-th traffic light to be identified. centerj Let Δ be the average of the embedding vectors of the four corner points of the j-th traffic light to be identified, where Δ is a constant 1;

[0042] Based on the first loss function and the second loss function, the four corner points belonging to the same traffic light to be identified in each frame of traffic light image are determined. Among them, the first vector distance of the four corner points of the same traffic light to be identified in each frame of traffic light image is the closest, while the second vector distance of the four corner points of different traffic lights to be identified in each frame of traffic light image is not the closest.

[0043] In one possible implementation of this application, the traffic light grouping module is used for:

[0044] The third vector distance between each traffic light to be identified in multiple consecutive frames of traffic light images is calculated using a third loss function, wherein the third loss function is:

[0045]

[0046] Where L_pull2 is the value of the third loss function, and X and Y represent that there are X traffic lights to be identified in each of the Y consecutive frames of traffic light images to be identified. Let be the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the y-th traffic light image. This is the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the first traffic light image;

[0047] Based on the third loss function, the signal light regions belonging to the same signal light in the multiple consecutive frames of signal light images to be identified are determined, wherein the third vector distance between the corner points belonging to the same signal light in different signal light images is the smallest.

[0048] A third aspect of this application provides an electronic device, comprising:

[0049] Processor; and

[0050] A memory that stores executable code, which, when executed by the processor, causes the processor to perform the method described above.

[0051] A fourth aspect of this application provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method described above.

[0052] This application embodiment acquires multiple consecutive frames of traffic light images to be identified, uses these images as input to a recurrent neural network model, and identifies the corner points of each traffic light in each frame. This yields a positional heatmap and embedding vector for each corner point of each traffic light in each frame. Then, based on the positional heatmap and embedding vector of each corner point, three different loss functions are used to identify the corner points, determining the four corner points belonging to the same traffic light in the same frame and the traffic light regions belonging to the same traffic light in different frames. This approach can accurately identify traffic lights and also group traffic lights in consecutive frames, resulting in good recognition performance and high recognition efficiency.

[0053] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0054] The above and other objects, features and advantages of this application will become more apparent from the following description of exemplary embodiments of this application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components.

[0055] Figure 1 This is a flowchart illustrating a method for identifying and grouping traffic lights according to an embodiment of this application;

[0056] Figure 2 This is a flowchart illustrating a method for generating a corner location heatmap according to an embodiment of this application;

[0057] Figure 3 This is a flowchart illustrating a method for generating corner embedding vectors according to an embodiment of this application;

[0058] Figure 4 This is a schematic diagram of the structure of a traffic light identification and grouping device shown in an embodiment of this application;

[0059] Figure 5 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application.

[0060] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. Detailed Implementation

[0061] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.

[0062] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0063] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0064] With the development of autonomous driving technology, the level of automation in autonomous vehicles is increasing. For safety and navigation reasons, autonomous vehicles require higher accuracy in traffic light recognition. Current technical solutions typically employ neural networks to identify traffic lights in a single road image, then use the Hungarian algorithm or feature point extraction methods to match signals from multiple images. However, because traffic lights are small targets and difficult to match, this process is time-consuming and yields poor recognition and matching results. Alternatively, consecutive frames of images can be used to identify traffic lights, followed by algorithms such as optimal matching lights. However, this method also has unsatisfactory grouping results for traffic lights and needs improvement.

[0065] To address the aforementioned issues, this application provides a method for identifying and grouping traffic lights, which can accurately identify traffic lights and group the same traffic lights in consecutive frame images.

[0066] Figure 1 This is a flowchart illustrating a method for identifying and grouping traffic lights according to an embodiment of this application.

[0067] See Figure 1 The method for identifying and grouping traffic lights provided in this application includes:

[0068] Step S101: Obtain multiple consecutive frame images of traffic lights to be identified, wherein each traffic light image includes at least one traffic light to be identified.

[0069] In this embodiment, when identifying and grouping traffic lights, multiple images can be identified simultaneously. These multiple consecutive frame traffic light images to be identified are consecutive frame images from a video of traffic lights on the road surface captured by an autonomous vehicle while driving on the target road. To ensure the effectiveness of traffic light identification and grouping, these multiple consecutive frame traffic light images are taken at different times for the same group of traffic lights. Optionally, the multiple consecutive frame traffic light images may contain one or more traffic lights to be identified; this application does not impose any restrictions on this.

[0070] Step S102: The multiple consecutive frames of traffic light images to be identified are used as input to the recurrent neural network model. The corner points of each traffic light to be identified in each frame of the traffic light image are identified to obtain the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image.

[0071] In this embodiment of the application, after acquiring multiple consecutive frames of traffic light images to be identified, they are used as input to a recurrent neural network. The recurrent neural network can be an RNN (Recurrent Neural Network). The characteristic of this recurrent neural network is that the number of outputs is the same as the number of inputs. By using the RNN to identify the corner points of each traffic light to be identified in each frame of traffic light images, the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of traffic light images can be obtained.

[0072] As one possible implementation of this application, such as Figure 2 As shown, the recurrent neural network model is a bidirectional gated recurrent unit model. The process involves using the multiple consecutive frames of traffic light images to be identified as input to the recurrent neural network model, identifying the corner points of each traffic light in each frame, and obtaining the position heatmap and embedding vector of each corner point of each traffic light in each frame, including:

[0073] Step S201: Input the multiple consecutive frame traffic light images to be identified into the bidirectional gate loop unit model.

[0074] In this embodiment of the application, multiple consecutive frame traffic light images to be identified contain one or more identical traffic lights to be identified. These multiple consecutive frame traffic light images to be identified are input into a bidirectional GRU (Gate Recurrent Unit) for identification.

[0075] Step S202: Identify the corner points of each traffic light to be identified in each traffic light image to be identified, and generate a corner point position heatmap based on the position of each corner point of each traffic light to be identified. The corner point position heatmap consists of 4 images, and each corner point position heatmap records the corner point position of all traffic lights to be identified in one direction in the multiple consecutive frames of traffic light images to be identified.

[0076] In this embodiment of the application, for a traffic light image to be identified, it may contain multiple traffic lights to be identified. For each traffic light, there should be four corner points. Generally, the traffic light is a rectangular frame, and the four vertices of the rectangular frame are the four corner points of the traffic light. When using a bidirectional GRU model to identify each traffic light in each frame of the traffic light image to be identified, the four corner points of each traffic light can be identified. Therefore, when outputting the position heatmap of the four corner points of the traffic light, four position heatmaps can be used to display the four corner points of the traffic light respectively. Optionally, each of the four location heatmaps is used to record the corner point heatmap of one direction. Specifically, each traffic light has four corner points: the upper left corner, the upper right corner, the lower left corner, and the lower right corner. The four location heatmaps are used to record the corner point position of each traffic light in one direction. For example, in a frame of traffic light image to be identified, there are three traffic lights to be identified. When identifying the corner points of the traffic lights to be identified in the image, the four location heatmaps obtained are as follows: the first location heatmap records the position of the upper left corner of the three traffic lights to be identified; the second location heatmap records the position of the upper right corner of the three traffic lights to be identified; the third location heatmap records the position of the lower left corner of the three traffic lights to be identified; and the fourth location heatmap records the position of the lower right corner of the three traffic lights to be identified.

[0077] Step S203: Based on the positional heatmap of the corner points, generate the embedding vector of each corner point of each traffic light to be identified.

[0078] In this embodiment, when the bidirectional GRU model identifies the traffic light image to be identified, it generates corresponding embedding vectors for each corner point of each traffic light to be identified. Specifically, as shown below... Figure 3 As shown, the step of generating the embedding vector for each corner point of each traffic light to be identified based on the positional heatmap of the corner points includes:

[0079] Step S301: Obtain the pixel values ​​of each corner point of each traffic light to be identified in the position heat map of the corner points;

[0080] Step S302: Generate embedding vectors for each corner point based on the pixel values.

[0081] In this embodiment, after determining the positional heatmap of each corner point of each traffic light to be identified, the pixel value of each corner point is obtained. Then, based on the matrix containing the pixel value, an embedding vector is generated. Specifically, for each corner point, the pixels within a preset range of its positional heatmap can be defined as the matrix square range. Then, the pixels within this matrix range are vectorized to obtain the embedding vector of the corner point. In this embodiment, the embedding vector of each corner point of each traffic light to be identified can be obtained using the above method.

[0082] Step S103: Determine the four corner points of each traffic light to be identified in each frame of the traffic light image based on the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image.

[0083] In this embodiment of the application, after determining the position heatmap and embedding vector of each traffic light to be identified in each frame of traffic light image, it is necessary to identify each traffic light to be identified in each frame of traffic light image, specifically including:

[0084] The first vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image is calculated using a first loss function, wherein the first loss function is:

[0085]

[0086] Where L_pull1 is the value of the first loss function, e k Let e ​​be the embedding vector of the k-th corner point. center It is the average value of the embedding vectors of the four corner points of the same traffic light to be identified;

[0087] The second loss function is used to calculate the second vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image, where the second loss function is:

[0088]

[0089] Where L_push is the value of the second loss function, N is the number of traffic lights to be identified in each frame of the traffic light image, and e centeri e is the average of the embedding vectors of the four corner points of the i-th traffic light to be identified. centerj Let Δ be the average of the embedding vectors of the four corner points of the j-th traffic light to be identified, where Δ is a constant 1;

[0090] Based on the first loss function and the second loss function, the four corner points belonging to the same traffic light to be identified in each frame of traffic light image are determined. Among them, the first vector distance of the four corner points of the same traffic light to be identified in each frame of traffic light image is the closest, while the second vector distance of the four corner points of different traffic lights to be identified in each frame of traffic light image is not the closest.

[0091] In this embodiment, when identifying different traffic lights in the same frame of a traffic light image, based on the embedding vectors of each corner point of each traffic light, the vector distance between the embedding vectors of corner points belonging to the same traffic light should be as close as possible, and the vector distance between the embedding vectors of corner points belonging to different traffic lights should be as far as possible. Therefore, a first loss function and a second loss function are set to calculate the first vector distance and the second vector distance, respectively. The first vector distance is used to represent the vector distance between different corner points belonging to the same traffic light, and the second vector distance is used to represent the vector distance between corner points belonging to different traffic lights. In this embodiment, by using the first loss function and the second loss function, the accuracy of traffic light identification in a single frame of a traffic light image can be guaranteed when the bidirectional GRU model is used to identify the image, especially when there are multiple traffic lights to be identified in a single frame of a traffic light image.

[0092] Step S104: Based on the four corner points of each traffic light to be identified and the embedding vector, determine the traffic light region belonging to the same traffic light in the multiple consecutive frames of traffic light images to be identified.

[0093] In this embodiment of the application, after identifying each traffic light in each frame of the traffic light image to be identified, it is necessary to group the same traffic light in different frames of traffic light images to be identified, that is, it is necessary to identify the same traffic light in consecutive frames of images collected by the acquisition vehicle. Specifically, this includes:

[0094] The third vector distance between each traffic light to be identified in multiple consecutive frames of traffic light images is calculated using a third loss function, wherein the third loss function is:

[0095]

[0096] Where L_pull2 is the value of the third loss function, and X and Y represent that there are X traffic lights to be identified in each of the Y consecutive frames of traffic light images to be identified. Let be the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the y-th traffic light image. This is the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the first traffic light image;

[0097] Based on the third loss function, the signal light regions belonging to the same signal light in the multiple consecutive frames of signal light images to be identified are determined, wherein the third vector distance between the corner points belonging to the same signal light in different signal light images is the smallest.

[0098] In this embodiment of the application, for the same traffic light in different frame images, the vector distance between the embedding vectors of its corner points in different frame images should be as small as possible. Therefore, a third loss function is introduced to calculate the vector distance between the embedding vectors of the corner points of the traffic light in different frame images. Optionally, the first frame image in multiple consecutive frame images of traffic lights to be identified is used as a reference, and the third vector distance between the same corner point between two consecutive frame images is calculated respectively. The corner point with the smallest third vector distance is taken as the corner point of the same traffic light in the consecutive frames, thereby determining the region belonging to the same traffic light to be identified in different consecutive frame images.

[0099] This application embodiment acquires multiple consecutive frames of traffic light images to be identified, uses these images as input to a recurrent neural network model, and identifies the corner points of each traffic light in each frame. This yields a positional heatmap and embedding vector for each corner point of each traffic light in each frame. Then, based on the positional heatmap and embedding vector of each corner point, three different loss functions are used to identify the corner points, determining the four corner points belonging to the same traffic light in the same frame and the traffic light regions belonging to the same traffic light in different frames. This approach can accurately identify traffic lights and also group traffic lights in consecutive frames, resulting in good recognition performance and high recognition efficiency.

[0100] Corresponding to the aforementioned application function implementation method embodiments, this application also provides a lane line recognition error feedback device, electronic device, and corresponding embodiments.

[0101] Figure 4 This is a schematic diagram of the structure of a traffic light identification and grouping device shown in an embodiment of this application.

[0102] See Figure 4 The traffic light identification and grouping device 40 provided in this application embodiment includes an image acquisition module 410, a model recognition module 420, a traffic light identification module 430, and a traffic light grouping module 440, wherein:

[0103] Image acquisition module 410 is used to acquire multiple consecutive frames of traffic light images to be identified, wherein the traffic light images include at least one traffic light to be identified;

[0104] The model recognition module 420 is used to take the multiple consecutive frames of traffic light images to be recognized as input to the recurrent neural network model, and to recognize the corner points of each traffic light to be recognized in each frame of traffic light images, so as to obtain the position heatmap and embedding vector of each corner point of each traffic light to be recognized in each frame of traffic light images.

[0105] Traffic light recognition module 430 is used to determine the four corner points of each traffic light to be recognized in each frame of traffic light image based on the position heatmap and embedding vector of each corner point of each traffic light to be recognized in each frame of traffic light image;

[0106] The traffic light grouping module 440 is used to determine the traffic light region belonging to the same traffic light in the multiple consecutive frames of traffic light images to be identified, based on the four corner points of each traffic light to be identified and the embedding vector.

[0107] As one possible implementation of this application, the recurrent neural network model is a bidirectional gated recurrent unit model. The step of using the multiple consecutive frames of traffic light images to be identified as input to the recurrent neural network model, identifying the corner points of each traffic light to be identified in each frame of the traffic light image, and obtaining the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image, includes:

[0108] The multiple consecutive frames of traffic light images to be identified are input into the bidirectional gate loop unit model;

[0109] The corner points of each traffic light to be identified in each traffic light image to be identified are identified, and a heat map of the corner point position is generated based on the position of each corner point of each traffic light to be identified. The heat map of the corner point position consists of 4 images, and each heat map of the corner point position records the corner point position of all traffic lights to be identified in one direction in the multiple consecutive frames of traffic light images to be identified.

[0110] Based on the location heatmap of the corner points, an embedding vector for each corner point of each traffic light to be identified is generated.

[0111] As one possible implementation of this application, the traffic light recognition module is used for:

[0112] The first vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image is calculated using a first loss function, wherein the first loss function is:

[0113]

[0114] Where L_pull1 is the value of the first loss function, e k Let e ​​be the embedding vector of the k-th corner point. centerIt is the average value of the embedding vectors of the four corner points of the same traffic light to be identified;

[0115] The second loss function is used to calculate the second vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image, where the second loss function is:

[0116]

[0117] Where L_push is the value of the second loss function, N is the number of traffic lights to be identified in each frame of the traffic light image, and e centeri e is the average of the embedding vectors of the four corner points of the i-th traffic light to be identified. centerj Let Δ be the average of the embedding vectors of the four corner points of the j-th traffic light to be identified, where Δ is a constant 1;

[0118] Based on the first loss function and the second loss function, the four corner points belonging to the same traffic light to be identified in each frame of traffic light image are determined. Among them, the first vector distance of the four corner points of the same traffic light to be identified in each frame of traffic light image is the closest, while the second vector distance of the four corner points of different traffic lights to be identified in each frame of traffic light image is not the closest.

[0119] As one possible implementation of this application, the traffic light grouping module is used for:

[0120] The third vector distance between each traffic light to be identified in multiple consecutive frames of traffic light images is calculated using a third loss function, wherein the third loss function is:

[0121]

[0122] Where L_pull2 is the value of the third loss function, and X and Y represent that there are X traffic lights to be identified in each of the Y consecutive frames of traffic light images to be identified. Let be the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the y-th traffic light image. This is the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the first traffic light image;

[0123] Based on the third loss function, the signal light regions belonging to the same signal light in the multiple consecutive frames of signal light images to be identified are determined, wherein the third vector distance between the corner points belonging to the same signal light in different signal light images is the smallest.

[0124] As one possible implementation of this application, generating the embedding vector of each corner point of each traffic light to be identified based on the positional heatmap of the corner points includes:

[0125] Obtain the pixel values ​​of each corner point of each traffic light to be identified in the position heatmap of the corner points;

[0126] An embedding vector for each corner point is generated based on the pixel values.

[0127] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated further here.

[0128] This application embodiment acquires multiple consecutive frames of traffic light images to be identified, uses these images as input to a recurrent neural network model, and identifies the corner points of each traffic light in each frame. This yields a positional heatmap and embedding vector for each corner point of each traffic light in each frame. Then, based on the positional heatmap and embedding vector of each corner point, three different loss functions are used to identify the corner points, determining the four corner points belonging to the same traffic light in the same frame and the traffic light regions belonging to the same traffic light in different frames. This approach can accurately identify traffic lights and also group traffic lights in consecutive frames, resulting in good recognition performance and high recognition efficiency.

[0129] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device 500 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0130] The electronic device includes a memory and a processor, wherein the processor may be referred to as processing device 501 as described below, and the memory may include at least one of read-only memory (ROM) 502, random access memory (RAM) 503, and storage device 508 as described below, as follows:

[0131] like Figure 5As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0132] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0133] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.

[0134] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0135] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0136] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire multiple consecutive frames of traffic light images to be identified, wherein each traffic light image includes at least one traffic light to be identified; use the multiple consecutive frames of traffic light images to be identified as input to a recurrent neural network model, identify the corner points of each traffic light to be identified in each frame of traffic light images, and obtain a positional heatmap and embedding vector for each corner point of each traffic light to be identified in each frame of traffic light images; determine the four corner points of each traffic light to be identified in each frame of traffic light images based on the positional heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of traffic light images; and determine the traffic light region belonging to the same traffic light to be identified in the multiple consecutive frames of traffic light images to be identified based on the four corner points of each traffic light to be identified and the embedding vector.

[0137] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0139] The modules or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules or units do not necessarily limit the specific unit itself.

[0140] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0141] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0142] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0143] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0144] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for identifying and grouping traffic lights, characterized in that, include: Acquire multiple consecutive frames of traffic light images to be identified, wherein each traffic light image includes at least one traffic light to be identified; The multiple consecutive frames of traffic light images to be identified are used as input to a recurrent neural network model. The corner points of each traffic light to be identified in each frame of the traffic light image are identified to obtain the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image. The four corner points of each traffic light to be identified in each frame of the traffic light image are determined based on the position heatmap and embedding vector of each corner point of each traffic light to be identified in each frame of the traffic light image; It includes: The first vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image is calculated using a first loss function, wherein the first loss function is: in, The value of the first loss function, Let be the embedding vector of the k-th corner point. It is the average value of the embedding vectors of the four corner points of the same traffic light to be identified; The second loss function is used to calculate the second vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image, where the second loss function is: in, The value of the second loss function, N The number of traffic lights to be identified in each frame of the traffic light image. For the first i The average value of the embedding vectors of the four corner points of each traffic light to be identified. For the first j The average value of the embedding vectors of the four corner points of each traffic light to be identified. It is a constant of 1; Based on the first loss function and the second loss function, the four corner points belonging to the same traffic light to be identified in each frame of traffic light image are determined. Among them, the first vector distance of the four corner points of the same traffic light to be identified in each frame of traffic light image is the closest, and the second vector distance of the four corner points of different traffic lights to be identified in each frame of traffic light image is not the closest. Based on the four corner points of each traffic light to be identified and the embedding vector, the traffic light regions belonging to the same traffic light in the multiple consecutive frames of traffic light images to be identified are determined.

2. The method for identifying and grouping traffic lights according to claim 1, characterized in that, The recurrent neural network model is a bidirectional gated recurrent unit model. The multiple consecutive frames of traffic light images to be identified are used as input to the recurrent neural network model. The corner points of each traffic light to be identified in each frame of the traffic light image are identified to obtain a position heatmap and embedding vector for each corner point of each traffic light to be identified in each frame of the traffic light image, including: The multiple consecutive frames of traffic light images to be identified are input into the bidirectional gate loop unit model; The corner points of each traffic light to be identified in each traffic light image to be identified are identified, and a heat map of the corner point position is generated based on the position of each corner point of each traffic light to be identified. The heat map of the corner point position consists of 4 images, and each heat map of the corner point position records the corner point position of all traffic lights to be identified in one direction in the multiple consecutive frames of traffic light images to be identified. Based on the location heatmap of the corner points, an embedding vector for each corner point of each traffic light to be identified is generated.

3. The method for identifying and grouping traffic lights according to claim 1, characterized in that, The step of determining the signal light region belonging to the same signal light in the multiple consecutive frames of signal light images to be identified, based on the four corner points of each signal light to be identified and the embedding vector, includes: A third loss function is used to calculate the third vector distance between each traffic light to be identified in multiple consecutive frames of traffic light images. The third loss function is: in, Let X and Y represent the values ​​of the third loss function, where X represents the number of traffic lights to be identified in each of the Y consecutive frames of traffic light images to be identified. Let be the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the y-th traffic light image. This is the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the first traffic light image; Based on the third loss function, the signal light regions belonging to the same signal light in the multiple consecutive frames of signal light images to be identified are determined, wherein the third vector distance between the corner points belonging to the same signal light in different signal light images is the smallest.

4. The method for identifying and grouping traffic lights according to claim 2, characterized in that, The step of generating embedding vectors for each corner point of each traffic light to be identified based on the location heatmap of the corner points includes: Obtain the pixel values ​​of each corner point of each traffic light to be identified in the position heatmap of the corner points; An embedding vector for each corner point is generated based on the pixel values.

5. A device for identifying and grouping traffic lights, characterized in that, include: An image acquisition module is used to acquire multiple consecutive frames of traffic light images to be identified, wherein each traffic light image includes at least one traffic light to be identified; The model recognition module is used to take the multiple consecutive frames of traffic light images to be recognized as input to the recurrent neural network model, and to recognize the corner points of each traffic light to be recognized in each frame of traffic light images, so as to obtain the position heatmap and embedding vector of each corner point of each traffic light to be recognized in each frame of traffic light images. The traffic light recognition module is used to determine the four corner points of each traffic light to be recognized in each frame of traffic light image based on the position heatmap and embedding vector of each corner point of each traffic light to be recognized in each frame of traffic light image; It includes: The first vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image is calculated using a first loss function, wherein the first loss function is: in, The value of the first loss function, Let be the embedding vector of the k-th corner point. It is the average value of the embedding vectors of the four corner points of the same traffic light to be identified; The second loss function is used to calculate the second vector distance between the embedding vectors of each traffic light corner point in each frame of the traffic light image, where the second loss function is: in, The value of the second loss function, N The number of traffic lights to be identified in each frame of the traffic light image. For the first i The average value of the embedding vectors of the four corner points of each traffic light to be identified. For the first j The average value of the embedding vectors of the four corner points of each traffic light to be identified. It is a constant of 1; Based on the first loss function and the second loss function, the four corner points belonging to the same traffic light to be identified in each frame of traffic light image are determined. Among them, the first vector distance of the four corner points of the same traffic light to be identified in each frame of traffic light image is the closest, and the second vector distance of the four corner points of different traffic lights to be identified in each frame of traffic light image is not the closest. The traffic light grouping module is used to determine the traffic light regions belonging to the same traffic light in the multiple consecutive frames of traffic light images to be identified, based on the four corner points of each traffic light to be identified and the embedding vector.

6. The device for identifying and grouping traffic lights according to claim 5, characterized in that, The traffic light grouping module is used for: A third loss function is used to calculate the third vector distance between each traffic light to be identified in multiple consecutive frames of traffic light images. The third loss function is: in, Let X and Y represent the values ​​of the third loss function, where X represents the number of traffic lights to be identified in each of the Y consecutive frames of traffic light images to be identified. Let be the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the y-th traffic light image. This is the average of the embedding vectors of the four corner points of the x-th traffic light to be identified in the first traffic light image; Based on the third loss function, the signal light regions belonging to the same signal light in the multiple consecutive frames of signal light images to be identified are determined, wherein the third vector distance between the corner points belonging to the same signal light in different signal light images is the smallest.

7. An electronic device, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-4.

8. A computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Method and device for detecting face image

    CN110276277A

  • Signal lamp state correction method and device and computer readable storage medium

    CN112633137A