Recognition Method, Device and Electronic Equipment for Traffic Lights

By using multiple cameras to acquire images on autonomous vehicles and extracting cross-fusion features in combination with deep learning models, the problem of low accuracy in traffic light recognition is solved, and more efficient traffic light status recognition is achieved.

CN115273021BActive Publication Date: 2025-07-25GUANGZHOU WERIDE TECH LTD CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210557652.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-19
Publication Date
2025-07-25
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

In autonomous driving systems, it is difficult for the prior art to effectively use images captured by multiple cameras to accurately identify the traffic lights' traffic state, especially in assisted driving systems with low automation, where the cross-image obstacle tracking task is complex and the recognition accuracy is poor.

Method used

By setting up multiple on-board cameras on the vehicle to obtain target images, the pre-trained traffic light recognition model extracts the cross-fusion characteristics of traffic lights, and combining the multi-head attention mechanism and cross-attention mechanism, the interaction relationship between the time dimension and the spatial dimension between each traffic light is determined, and the traffic state is determined.

Benefits of technology

It improves the accuracy of traffic light recognition, simplifies the problem of cross-camera obstacle tracking, reduces information loss, and improves recognition efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273021B_ABST
    Figure CN115273021B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus and electronic device for identifying traffic lights, which acquire a target area of an image containing a target traffic light in a plurality of target images; input the target area into a pre-trained traffic light recognition model, extract cross-fusion features of the target traffic light from the target area through the traffic light recognition model, determine traffic characteristics of a preset traffic direction based on the cross-fusion features, and determine the traffic state indicated by the traffic light based on the traffic characteristics. In the process of identifying traffic lights through a plurality of vehicle-mounted images captured by multiple cameras, this method extracts cross-fusion features that can represent the interaction relationship between time dimension and / or space dimension of each traffic light, further determines the influence of the traffic light on the traffic state of each traffic direction, thereby determining the traffic state indicated by the traffic light and improving the accuracy of the recognition result of the traffic light.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and in particular, to a method, device, and electronic device for identifying traffic lights. Background Art

[0002] In order to capture a relatively complete image of a traffic light, multiple image sensors are usually distributed on a vehicle body. During the process of identifying a traffic light, it is necessary to determine the traffic state indicated by the traffic light based on multiple images of the target traffic light captured by multiple image sensors.

[0003] In related technologies, usually, through the identification of traffic lights in a semantic map, or mapping each traffic light in each image to the four directions of vehicle travel to form a tracking identifier, and voting on the tracking results during the process of traffic light identification to achieve the identification of the target traffic light. However, semantic maps cannot be used in vehicles with a low degree of automation, and the method based on rule voting has certain limitations, resulting in a relatively poor accuracy of the traffic light identification result. Summary of the Invention

[0004] In view of this, an object of the present invention is to provide a method, device, and electronic device for identifying traffic lights, so as to improve the accuracy of the identification result during the process of identifying traffic lights through multiple vehicle-mounted images captured by multiple cameras.

[0005] In a first aspect, an embodiment of the present invention provides a method for identifying traffic lights, the method including: obtaining a target area in multiple target images; the multiple target images are captured by multiple vehicle-mounted cameras arranged on the same vehicle for the same traffic scene; the target area includes at least a partial image of a target traffic light; inputting the target area into a pre-trained traffic light identification model, extracting, by the traffic light identification model, cross-fusion features of the target traffic light from the target area, determining, based on the cross-fusion features, traffic features of a preset traffic direction, and determining, based on the traffic features, the traffic state indicated by the traffic light; wherein the cross-fusion features indicate the interaction relationship in the time dimension and / or space dimension between the traffic lights in the target traffic light; the traffic features indicate the influence of the target traffic light on the traffic state of the traffic direction.

[0006] Optionally, in the first implementation manner of the first aspect of the present invention, the step of obtaining the target area in the multiple target images includes: for each target image, processing the target image through a pre-trained obstacle recognition model to obtain a recognition result; the recognition result includes the category of the recognized obstacle and the position information of the obstacle in the target image; based on the position information of the obstacle with the category of traffic signal in the target image, intercept the target area including the obstacle with the category of traffic signal from the target image.

[0007] Optionally, in the second implementation manner of the first aspect of the present invention, the above traffic signal recognition model includes a feature extraction network and a feature fusion network; the feature fusion network is established based on the multi-head attention mechanism; the step of processing the target area through the pre-trained traffic signal recognition model to obtain the cross-fusion feature of the target traffic signal includes: performing feature extraction processing on the target area through the feature extraction network to obtain the image feature in the target area; the image feature includes the attribute features of each traffic signal in the target traffic signal in the time dimension and / or the space dimension; performing feature fusion processing on the image feature through the feature fusion network based on the multi-head attention mechanism to obtain the cross-fusion feature of the target traffic signal.

[0008] Optionally, in the third implementation manner of the first aspect of the present invention, the above feature extraction network includes a residual neural network; the step of performing feature extraction processing on the target area through the feature extraction network to obtain the image feature in the target area includes: mapping the target area to a pre-set high-dimensional space through the residual neural network to obtain a mapping result; performing feature extraction processing on the mapping result to obtain the image feature in the target area.

[0009] Optionally, in the fourth implementation manner of the first aspect of the present invention, there are multiple above target areas, and the cross-fusion feature corresponds to the target area; the traffic signal recognition model includes a feature screening network; the feature screening network is established based on the cross-attention mechanism; the preset traffic directions include: straight, left turn, right turn, and U-turn; the step of determining the traffic feature of the preset traffic direction based on the cross-fusion feature includes: performing screening processing on the cross-fusion features of multiple target areas through the feature screening network based on the cross-attention mechanism to obtain a screening result; the screening result includes the cross-fusion features related to each traffic direction; for each traffic direction, determining the cross-fusion feature related to the traffic direction in the screening result as the traffic feature of the traffic direction.

[0010] Optionally, in the fifth implementation manner of the first aspect of the present invention, the above traffic signal recognition model includes a decoding network; the step of determining the traffic state indicated by the traffic signal based on the traffic feature includes: performing decoding processing on the traffic feature through the decoding network to obtain the traffic state indicated by the traffic signal.

[0011] Optionally, in the sixth implementation manner of the first aspect of the present invention, the above traffic signal recognition model is trained as follows: Determine training data from a preset sample set; the training data includes a target area containing multiple sample images and the traffic signal indication of the current traffic state in the target area; the multiple sample images are taken by multiple vehicle-mounted cameras arranged on the same vehicle for the same traffic scene; the target area includes at least part of the traffic signal images; input the target areas of the multiple sample images into the initial model to obtain the processing result output by the initial model; determine the loss value of the initial model based on the processing result and the current traffic state indicated by the traffic signal; update the model parameters of the initial model based on the loss value; continue to execute the step of determining training data from the preset sample data until the loss value converges, and determine the initial model with the converged loss value as the traffic signal recognition model.

[0012] In a second aspect, an embodiment of the present invention provides a traffic signal recognition device, which includes: a target area recognition module, configured to obtain a target area in multiple target images; the multiple target images are taken by multiple vehicle-mounted cameras arranged on the same vehicle for the same traffic scene; the target area includes at least part of the target traffic signal images; a traffic state determination module, configured to input the target area into a pre-trained traffic signal recognition model, extract the cross-fusion features of the target traffic signal from the target area through the traffic signal recognition model, determine the traffic features of the preset traffic direction based on the cross-fusion features, and determine the traffic state indicated by the traffic signal based on the traffic features; wherein, the cross-fusion features indicate the interaction relationship between each traffic signal in the target traffic signal in the time dimension and / or space dimension; the traffic features indicate the influence of the target traffic signal on the traffic state of the traffic direction.

[0013] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above traffic signal recognition method.

[0014] In a fourth aspect, an embodiment of the present invention provides a machine-readable storage medium, which stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the above traffic signal recognition method.

[0015] The embodiments of the present invention bring the following beneficial effects:

[0016] An embodiment of the present invention provides a method, apparatus, and electronic device for identifying traffic lights, which includes obtaining a target area of an image containing a target traffic light in a plurality of target images; inputting the target area into a pre-trained traffic light recognition model, extracting cross-fusion features of the target traffic light from the target area through the traffic light recognition model, determining traffic characteristics of a preset traffic direction based on the cross-fusion features, and determining the traffic state indicated by the traffic light based on the traffic characteristics. In the process of identifying traffic lights through multiple vehicle-mounted images captured by multiple cameras, this method extracts cross-fusion features that can represent the interaction relationship between each traffic light in terms of time dimension and / or space dimension, further determines the impact of traffic lights on the traffic states of each traffic direction, and thus determines the traffic state indicated by the traffic light, improving the accuracy of the recognition result of the traffic light.

[0017] Other features and advantages of the present invention will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the specific structures particularly pointed out in the specification, claims, and drawings.

[0018] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a flowchart of a method for identifying a traffic light provided by an embodiment of the present invention;

[0021] Figure 2 It is a flowchart of another method for identifying a traffic light provided by an embodiment of the present invention;

[0022] Figure 3 It is a schematic structural diagram of a traffic light recognition model provided by an embodiment of the present invention;

[0023] Figure 4 It is a schematic structural diagram of an apparatus for identifying a traffic light provided by an embodiment of the present invention;

[0024] Figure 5A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] The recognition of traffic lights is an important task in the fields of autonomous driving, assisted driving, and driving safety. Since there are usually multiple image sensors on the vehicle body to collect images of traffic lights. During the recognition process of traffic lights, it is necessary to determine the indicated traffic state based on multiple images of traffic lights collected by multiple image sensors and related to the time sequence. Common processing solutions are generally divided into two stages. The first stage is the detection stage, which is mainly responsible for detecting the position and corresponding attributes of traffic lights, such as color and shape. The second stage is the decision-making stage, which combines the detection results of multiple frames to give the decision result in terms of time sequence, determines the traffic state indicated by the traffic lights, and then further plans the driving trajectory through the path planning module based on the traffic state indicated by the traffic lights.

[0027] There are the following two common solutions in the decision-making stage:

[0028] (1) Fuse the time sequence information of each traffic light, and then give the state of the final time sequence. Among them, the acquisition of each time sequence information strongly depends on the time sequence tracking of the box. Since the shooting angles between multiple cameras usually vary greatly, the task of obstacle tracking across images is a relatively difficult problem, and there are also certain bottlenecks in the final accuracy. Although in an autonomous driving system with a level of autonomous driving of L4, this problem can be avoided by relying on the unique identifier (id) of the semantic map, in an assisted driving system with a level of autonomous driving of L2 or L3 where the semantic map cannot be relied on, the problem of the recognition accuracy of traffic lights caused by obstacle tracking across images cannot be avoided.

[0029] (2) Avoid the difficulties in cross-camera obstacle tracking through a preprocessing method. Specifically: Map all traffic lights detected in each frame to the four driving directions of the vehicle, namely left turn, straight ahead, right turn, and U-turn (also known as "uturn"). In this way, each frame has a natural tracking ID. Then, based on the voted results, use a temporal model to process the features of each traffic light. This rule-based voting has certain limitations and poor scalability. Of course, it is also possible to put the voting process into the model for learning, but inevitably, it will increase the difficulty of model learning and is difficult to handle problems such as data balance during the training process, ultimately resulting in poor model accuracy.

[0030] In addition, in the above solution, there will be a certain degree of information loss during the information transmission process (at the decision-making stage), that is, it is difficult to obtain the information of the original sensor, and even if it is obtained, it will increase a lot of repeated calculations for image feature extraction, resulting in low recognition efficiency of traffic lights.

[0031] Based on this, an identification method, device, and electronic device for traffic lights provided by an embodiment of the present invention can be applied to the identification process of traffic lights in various scenarios.

[0032] For the convenience of understanding this embodiment, first, a detailed introduction to an identification method for traffic lights disclosed in an embodiment of the present invention is as follows Figure 1 As shown, the method includes the following steps:

[0033] Step 102, obtain the target area in multiple target images; the multiple target images are taken by multiple on-vehicle cameras arranged on the same vehicle for the same traffic scene; the target area includes at least part of the image of the target traffic light.

[0034] The above target images are usually images collected by multiple on-vehicle cameras on the same vehicle. These target images have a certain shooting order, so that there is a certain temporal relationship in the display states of the traffic lights in the target images. For example, when the vehicle drives towards the traffic intersection from far to near, eight images including the target traffic lights at the traffic intersection are successively collected by camera A, camera B, camera C, and camera D. These eight images can be used as target images for identifying the target traffic lights at the traffic intersection.

[0035] The above-mentioned target area is the image area in the target image that includes the target traffic signal. First, the traffic signal in the target image can be recognized. For example, a traffic signal detection model established based on neural network or deep learning theory can be used to detect the traffic signal in the target image, or various traffic signal detection and positioning technologies such as image detection technology based on color and shape can be used to determine parameters such as the contour and position of the traffic signal in the target image. Then, based on the parameters such as the contour and position of the traffic signal in the target image, the target area is further determined. The target area can be the smallest rectangular area containing the traffic signal, or can be appropriately expanded to obtain a more complete image of the traffic signal.

[0036] Step 104, input the target area into the pre-trained traffic signal recognition model. The traffic signal recognition model extracts the cross-fusion features of the target traffic signal from the target area, determines the traffic features of the preset traffic direction based on the cross-fusion features, and determines the traffic state indicated by the traffic signal based on the traffic features. Among them, the cross-fusion features indicate the interaction relationship in the time dimension and / or space dimension among the traffic signals in the target traffic signal; the traffic features indicate the influence of the target traffic signal on the traffic state of the traffic direction.

[0037] The above-mentioned traffic signal recognition model can be established based on neural network and deep learning theory. Before the traffic signal recognition model is put into use, it needs to be trained with a large amount of training data. These training data are usually the target areas in multiple groups of target images containing traffic signals whose indicated traffic states have been determined. Through these training data, the traffic signal recognition model learns how to extract the cross-fusion features of the target traffic signal from the target areas in multiple groups of target images, determines the traffic features of each traffic direction based on the cross-fusion features of the traffic signal, and then determines the traffic state indicated by the traffic signal based on the traffic features.

[0038] To implement the above functions, the traffic signal recognition model usually needs to include a feature extraction network, a feature fusion network, a classification network, and an output network connected in sequence. The feature extraction network can be used to extract the basic features of the traffic signal from the target area. These basic features can include features in the time dimension and features in the space dimension, such as the sequence of the lighting state and extinguishing state of each traffic signal of the target signal light, and the position of each traffic signal in the target area.

[0039] The correlation relationships between the above basic features need to be obtained through cross-fusion processing by a feature fusion network, generating cross-fusion features that can represent the interaction relationships in the time dimension and / or space dimension among the traffic lights in the target traffic light. The traffic lights can represent the following interaction relationships: the first traffic light is above the second traffic light, and the second traffic light lights up after the first traffic light goes out, etc. Among them, the feature fusion network can be established based on a multi-head attention mechanism or a cross-attention mechanism, and learn how to determine the interaction relationships among the traffic lights during the training process.

[0040] The above classification network can also be called a feature screening network, which is used to determine the passing features related to each passing direction from the cross-fusion features, and can also be considered as classifying the cross-fusion features into each passing direction. The above passing directions usually include straight, left turn, right turn, and U-turn in the direction where the vehicle is located, and can also include straight, left turn, right turn, and U-turn in the vertical direction of the direction where the vehicle is located, etc. After determining the passing features related to each passing direction, the output network can process the passing features related to each passing direction, such as decoding processing or mapping processing, etc., to obtain the passing states indicated by the traffic lights. For example, a left turn is prohibited, a right turn and a straight pass are normal passes, a U-turn is prohibited, and so on.

[0041] An embodiment of the present invention provides a method for identifying traffic lights, which includes obtaining a target area of an image containing a target traffic light in a plurality of target images; inputting the target area into a pre-trained traffic light recognition model, extracting cross-fusion features of the target traffic light from the target area by the traffic light recognition model, determining passing features of a preset passing direction based on the cross-fusion features, and determining the passing states indicated by the traffic lights based on the passing features. In the process of identifying traffic lights through a plurality of vehicle-mounted images captured by multiple cameras, this method extracts cross-fusion features in the images that can represent the interaction relationships in the time dimension and / or space dimension among the traffic lights, further determines the influence of the traffic lights on the passing states of each passing direction, thereby determining the passing states indicated by the traffic lights, and improving the accuracy of the traffic light recognition results.

[0042] An embodiment of the present invention also provides another method for identifying traffic lights, and this method is implemented based on the method Figure 1 shown. This method mainly describes the specific process of obtaining the target area in a plurality of target images, and the specific process of processing the target area by the traffic light recognition model to further determine the passing states indicated by the traffic lights. As Figure 2 shown, this method includes the following steps:

[0043] Step S202: For each target image, process the target image through a pre-trained obstacle recognition model to obtain a recognition result. The recognition result includes the category of the recognized obstacle and the position information of the obstacle in the target image.

[0044] The above obstacle recognition model can be established based on traditional machine learning algorithms, such as logistic regression and support vector machines, or based on deep learning algorithms, such as deep convolutional neural networks. During the model training process, data collected by an in-vehicle camera during road driving can be used as samples first, then obstacles are labeled in the camera images, and then the parameters and outputs of the model are trained through the labeled sample data.

[0045] Step S204: Based on the position information of the obstacle with the category of traffic signal in the target image, intercept the target area containing the obstacle with the category of traffic signal from the target image.

[0046] The obstacle with the category of traffic signal in the recognition result can be intercepted (also known as "cropping") from the original image captured by the in-vehicle camera to obtain the target area. Usually, at least one target area will be obtained in a target image. Correspondingly, the number of target areas is multiple.

[0047] The traffic signal recognition model for processing multiple target areas usually includes a feature extraction network, a feature fusion network, a feature screening network, and a decoding network connected in sequence, as Figure 3 shown.

[0048] The traffic signal recognition model can be trained in the following way:

[0049] (1) Determine the training data from a preset sample set. The training data includes the target area containing multiple sample images and the traffic signal indication of the current traffic state in the target area. The multiple sample images are captured by multiple in-vehicle cameras arranged on the same vehicle for the same traffic scene. The target area includes at least part of the traffic signal image.

[0050] (2) Input the target areas of the multiple sample images into the initial model to obtain the processing result output by the initial model. This processing result usually includes the traffic signal indication of the current traffic state in each passing direction obtained by the initial model.

[0051] (3) Based on the processing result and the current traffic state indicated by the traffic signal, determine the loss value of the initial model. The loss function can be selected according to requirements to calculate the loss value of the initial model based on the processing result and the current traffic state indicated by the traffic signal.

[0052] (4) Update the model parameters of the initial model based on the loss value; continue to execute the step of determining the training data from the preset sample data until the loss value converges, and determine the initial model with the converged loss value as the traffic signal recognition model.

[0053] Step S206, perform feature extraction processing on the target area through the feature extraction network to obtain the image features in the target area; the image features include the attribute features of each traffic signal in the target traffic signal in the time dimension and / or space dimension.

[0054] In order to extract more abundant attribute features, the feature extraction network can be a residual network. Specifically, the target area can be mapped to a pre-set high-dimensional space through a residual neural network to obtain a mapping result, and then feature extraction processing is performed on the mapping result to obtain the image features in the target area. The attribute features of the traffic signal in the time dimension and / or space dimension can include the position of the traffic signal in the image, the display content, and its state change features in chronological order in the target area.

[0055] Step S208, perform feature fusion processing on the image features through the feature fusion network based on the multi-head attention mechanism to obtain the cross-fusion features of the target traffic signal. The above feature fusion network is established based on the multi-head attention mechanism (multihead attention) to fuse the image features of the traffic signals with interaction relationships in the target traffic signal.

[0056] Step S210, perform screening processing on the cross-fusion features of multiple target areas through the feature screening network based on the cross-attention mechanism to obtain a screening result; the screening result includes the cross-fusion features related to each traffic direction.

[0057] The above target areas include multiple ones, and the cross-fusion features correspond to the target areas; the traffic signal recognition model includes a feature screening network; the feature screening network is established based on the cross-attention mechanism. The preset traffic directions can include: straight, left turn, right turn, and U-turn.

[0058] Step S212, for each traffic direction, determine the cross-fusion features related to the traffic direction in the screening result as the traffic features of the traffic direction.

[0059] Step S214, perform decoding processing on the traffic features through the decoding network to obtain the traffic states indicated by the traffic signals.

[0060] The above method can improve the accuracy of the result obtained by recognizing traffic signals through multiple vehicle-mounted images captured by multiple cameras.

[0061] An embodiment of the present invention also provides another method for recognizing traffic lights. This method is implemented on the basis of the method shown in Figure 1 This method provides an end-to-end traffic light recognition solution. The specific process is as follows: First, the upstream module for obstacle detection based on in-vehicle images gives the basic position information of traffic lights, and then the region of interest is intercepted (also called "cropped") according to the basic position information and used as a batch of original features to be input into an end-to-end network (traffic light decision maker) for inference. Finally, the states of four directions (left turn, straight, right turn, U-turn (also called uturn)) are obtained, and finally the state is output to the downstream path planning module.

[0062] This method is mainly implemented through the following steps:

[0063] S1: Use deep learning methods to detect and recognize obstacles in in-vehicle camera images.

[0064] Use deep learning methods to detect and recognize traffic lights in in-vehicle camera images, and detect the position and category information of obstacles.

[0065] S2: According to the detected traffic light position information in S1, crop the region of interest from the in-vehicle image, and send this batch of images into an end-to-end network (traffic light decision maker, equivalent to the above-mentioned "traffic light recognition model") for inference.

[0066] Take out the obstacles with the category of traffic lights in the detection results of S1 as an input similar to the region of interest, crop out the original in-vehicle camera images, and then feed this series of images into an end-to-end network for inference to obtain the final results of four directions (left turn, straight, right turn, uturn).

[0067] The main process of this part of the network is as follows:

[0068] S2-1: According to the position information in S1, perform cropping processing on the original images of multiple cameras to obtain a batch of images of regions of interest.

[0069] This module performs cropping processing on the original image based on the position information output in S1. When cropping, a certain scale can be expanded outward to retain certain environmental information. In the specific implementation scheme, parallel processing on the GPU (graphics processing unit) can be considered.

[0070] S2-2, Image feature extraction module: This module extracts features from the input batch of images based on a convolutional neural network. Common implementation schemes include common convolutional neural networks such as resnet and regnet.

[0071] S2-3, a multi-frame multi-camera fusion module based on the attention mechanism. This module is a fusion module based on the attention mechanism. For each piece of input traffic light image features output from S2-2, it can perceive the surrounding traffic light features (in both the time dimension and the space dimension) information. Through the transfer of cross information, richer features can be better extracted.

[0072] S2-4, a decoding module based on the attention mechanism: This module performs decoding operations on the features fused by S2-3. Using the attention mechanism, it performs an interrogation operation (equivalent to a screening operation) on the features of a period of time series, interrogating the features in four directions, that is, perceiving the influence of all traffic light information in all time dimensions and space dimensions on this driving direction, and performing decoding operations to obtain the final output result. The implementation is mainly based on cross-attention.

[0073] S2-5, a traffic light detection and classification module: This module is an auxiliary branch and does not output during actual vehicle operation. Its main function is to supervise the learning progress of the S2-2 module. Implementations can use common detection networks such as the YOLO series of networks.

[0074] The above method bypasses the difficult problem of cross-camera tracking with a deep learning-based method; uses an end-to-end solution to avoid the loss of original image information during the transmission process in a two-stage solution; uses a more elegant feature expression method and model design to improve the computational parallelism of image feature extraction, as well as the capacity and scalability of this solution. Compared with the original algorithm, the process is more concise, the model implementation and feature expression are more elegant, the scalability is strong, the engineering implementation and maintenance are easier, the cost of obtaining model data is greatly reduced, and it does not need to rely on any additional auxiliary modules such as semantic maps other than vision, and also avoids the difficult problem of cross-image tracking.

[0075] Corresponding to the above method embodiment, an embodiment of the present invention provides an identification device for traffic lights, as Figure 4 shown. The device includes:

[0076] A target area recognition module 400, configured to obtain the target area in multiple target images; the multiple target images are taken by multiple on-vehicle cameras arranged on the same vehicle for the same traffic scene; the target area includes at least part of the images of the target traffic lights;

[0077] The traffic state determination module 402 is configured to input the target area into a pre-trained traffic signal recognition model, extract the cross-fusion features of the target traffic signal from the target area through the traffic signal recognition model, determine the traffic features of the preset traffic direction based on the cross-fusion features, and determine the traffic state indicated by the traffic signal based on the traffic features; wherein, the cross-fusion features indicate the interaction relationship between the time dimension and / or the spatial dimension of each traffic signal in the target traffic signal; the traffic features indicate the influence of the target traffic signal on the traffic state of the traffic direction.

[0078] Further, the above-mentioned target area recognition module is further configured to: for each target image, process the target image through a pre-trained obstacle recognition model to obtain a recognition result; the recognition result includes the category of the recognized obstacle and the position information of the obstacle in the target image; based on the position information of the obstacle with the category of traffic signal in the target image, intercept the target area including the obstacle with the category of traffic signal from the target image.

[0079] Further, the above-mentioned traffic signal recognition model includes a feature extraction network and a feature fusion network; the feature fusion network is established based on the multi-head attention mechanism; the above-mentioned traffic state determination module is further configured to: perform feature extraction processing on the target area through the feature extraction network to obtain the image features in the target area; the image features include the attribute features of the time dimension and / or the spatial dimension of each traffic signal in the target traffic signal; perform feature fusion processing on the image features through the feature fusion network based on the multi-head attention mechanism to obtain the cross-fusion features of the target traffic signal.

[0080] Further, the above-mentioned feature extraction network includes a residual neural network; the above-mentioned traffic state determination module is further configured to: map the target area to a preset high-dimensional space through the residual neural network to obtain a mapping result; perform feature extraction processing on the mapping result to obtain the image features in the target area.

[0081] Further, there are multiple above-mentioned target areas, and the cross-fusion features correspond to the target areas; the traffic signal recognition model includes a feature screening network; the feature screening network is established based on the cross-attention mechanism; the preset traffic directions include: straight, left turn, right turn, and U-turn; the above-mentioned traffic state determination module is further configured to: perform screening processing on the cross-fusion features of multiple target areas through the feature screening network based on the cross-attention mechanism to obtain a screening result; the screening result includes the cross-fusion features related to each traffic direction; for each traffic direction, determine the cross-fusion features related to the traffic direction in the screening result as the traffic features of the traffic direction.

[0082] Further, the above traffic signal recognition model includes a decoding network; the above traffic state determination module is further configured to: perform decoding processing on the traffic features through the decoding network to obtain the traffic state indicated by the traffic signal.

[0083] Further, the above device further includes a model training module, configured to: determine training data from a preset sample set; the training data includes a target region containing multiple sample images and the traffic state indicated by the traffic signal in the target region; the multiple sample images are captured by multiple in-vehicle cameras arranged on the same vehicle for the same traffic scene; the target region includes images of at least part of the traffic signals; input the target regions of the multiple sample images into an initial model to obtain a processing result output by the initial model; determine a loss value of the initial model based on the processing result and the current traffic state indicated by the traffic signal; update the model parameters of the initial model based on the loss value; continue to execute the step of determining training data from the preset sample data until the loss value converges, and determine the initial model with the converged loss value as the traffic signal recognition model.

[0084] This embodiment further provides an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above traffic signal recognition method.

[0085] See Figure 5 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the above traffic signal recognition method.

[0086] Further, Figure 5 The electronic device shown further includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.

[0087] Among them, the memory 101 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), communication connection between the system network element and at least one other network element can be realized, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 5 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0088] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0089] This embodiment also provides a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the above-mentioned traffic signal recognition method.

[0090] A traffic signal recognition method, device, and electronic device provided by an embodiment of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For the specific implementation, reference can be made to the method embodiments, and details are not described herein again.

[0091] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.

[0092] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0093] If the above-mentioned functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0094] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0095] Finally, it should be noted that the above embodiments are only specific implementation manners of the present invention to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: any person skilled in the art can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements for some of the technical features within the technical scope disclosed by the present invention; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for identifying traffic signals, characterized in that, The method includes: Obtaining a target area in multiple target images; the multiple target images are captured by multiple on-vehicle cameras arranged on the same vehicle for the same traffic scene; the target area includes an image of at least part of a target traffic signal; Inputting the target area into a pre-trained traffic signal recognition model, extracting the cross-fusion features of the target traffic signal from the target area through the traffic signal recognition model, determining the traffic features of a preset traffic direction based on the cross-fusion features, and determining the traffic state indicated by the traffic signal based on the traffic features; Wherein, the cross-fusion features indicate the interaction relationship in the time dimension and / or space dimension between the traffic signals in the target traffic signal; the traffic features indicate the influence of the target traffic signal on the traffic state of the traffic direction; The traffic signal recognition model includes a feature extraction network and a feature fusion network; the feature fusion network is established based on a multi-head attention mechanism; The step of obtaining the cross-fusion features of the target traffic signal by processing the target area through a pre-trained traffic signal recognition model includes: Performing feature extraction processing on the target area through the feature extraction network to obtain image features in the target area; the image features include the attribute features in the time dimension and / or space dimension of each traffic signal in the target traffic signal; Performing feature fusion processing on the image features through the feature fusion network based on the multi-head attention mechanism to obtain the cross-fusion features of the target traffic signal; There are multiple target areas, and the cross-fusion features correspond to the target areas; the traffic signal recognition model includes a feature screening network; the feature screening network is established based on a cross-attention mechanism; the preset traffic directions include: straight, left turn, right turn, and U-turn; The step of determining the traffic features of a preset traffic direction based on the cross-fusion features includes: Performing screening processing on the cross-fusion features of multiple target areas through the feature screening network based on the cross-attention mechanism to obtain a screening result; the screening result includes the cross-fusion features related to each traffic direction; For each traffic direction, determining the cross-fusion features related to the traffic direction in the screening result as the traffic features of the traffic direction.

2. The method according to claim 1, wherein The step of obtaining the target area in multiple target images includes: For each target image, processing the target image through a pre-trained obstacle recognition model to obtain a recognition result; the recognition result includes the category of the recognized obstacle and the position information of the obstacle in the target image; Based on the position information of the obstacle with the category of traffic signal in the target image, intercepting a target area including the obstacle with the category of traffic signal from the target image.

3. The method according to claim 1, wherein The feature extraction network includes a residual neural network; The step of performing feature extraction processing on the target area through the feature extraction network to obtain image features in the target area includes: Mapping the target region to a pre-set high-dimensional space through a residual neural network to obtain a mapping result; Performing feature extraction processing on the mapping result to obtain image features in the target region.

4. The method according to claim 1, wherein The traffic signal recognition model includes a decoding network; The step of determining the traffic state indicated by the traffic signal based on the traffic feature includes: Performing decoding processing on the traffic feature through the decoding network to obtain the traffic state indicated by the traffic signal.

5. The method according to claim 1, wherein The traffic signal recognition model is trained in the following manner: Determining training data from a preset sample set; the training data includes a target region containing a plurality of sample images and the traffic state indicated by the traffic signal in the target region; The plurality of sample images are captured by a plurality of on-vehicle cameras arranged on the same vehicle for the same traffic scene; the target region includes images of at least part of the traffic signals; Inputting the target regions of the plurality of sample images into an initial model to obtain a processing result output by the initial model; Determining the loss value of the initial model based on the processing result and the current traffic state indicated by the traffic signal; Updating the model parameters of the initial model based on the loss value; continuing to execute the step of determining training data from the preset sample data until the loss value converges, and determining the initial model with the converged loss value as the traffic signal recognition model.

6. An identification device for traffic lights, characterized in that, The device includes: A target region recognition module, configured to obtain target regions in a plurality of target images; the plurality of target images are captured by a plurality of on-vehicle cameras arranged on the same vehicle for the same traffic scene; the target region includes images of at least part of the target traffic signals; A traffic state determination module, configured to input the target region into a pre-trained traffic signal recognition model, extract the cross-fusion features of the target traffic signal from the target region through the traffic signal recognition model, determine the traffic features of a preset traffic direction based on the cross-fusion features, and determine the traffic state indicated by the traffic signal based on the traffic features; Wherein, the cross-fusion feature indicates the interaction relationship of the time dimension and / or space dimension between each traffic signal in the target traffic signal; the traffic feature indicates the influence of the target traffic signal on the traffic state of the traffic direction; The traffic signal recognition model includes a feature extraction network and a feature fusion network; the feature fusion network is established based on a multi-head attention mechanism; The traffic state determination module is further configured to: Performing feature extraction processing on the target region through the feature extraction network to obtain image features in the target region; the image features include the attribute features of the time dimension and / or space dimension of each traffic signal in the target traffic signal; Performing feature fusion processing on the image features through the feature fusion network based on the multi-head attention mechanism to obtain the cross-fusion features of the target traffic signal; The target regions include multiple ones, and the cross-fusion features correspond to the target regions; the traffic signal recognition model includes a feature screening network; the feature screening network is established based on a cross-attention mechanism; the preset traffic directions include: straight, left turn, right turn, and U-turn; The traffic state determination module is further configured to: perform screening processing on the cross-fusion features of multiple target regions through the feature screening network based on the cross-attention mechanism to obtain a screening result; the screening result includes cross-fusion features related to each traffic direction; For each traffic direction, determine the cross-fusion features related to the traffic direction in the screening result as the traffic features of the traffic direction.

7. An electronic device, characterized in that, It includes a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the traffic signal recognition method according to any one of claims 1-5.

8. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the traffic signal recognition method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Traffic signal lamp state identification method, device and equipment and storage medium

    CN111310708A

  • Training method of traffic signal lamp identification model and identification method of traffic signal lamp

    CN114332704A