Signal lamp identification method and device, electronic equipment and storage medium
By fusing the recognition results from intersection monitoring cameras and vehicle-side sensors, the accuracy and robustness issues of traffic light recognition in autonomous driving have been resolved, enabling stable traffic light perception and recognition in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-29
AI Technical Summary
Existing autonomous driving technologies lack environmental adaptability and recognition accuracy in traffic light recognition in complex road scenarios. Furthermore, existing solutions are costly or rely on inadequate infrastructure, resulting in poor recognition reliability.
By utilizing existing intersection surveillance cameras and matching traffic light positions with historical frames, the system enables beyond-line-of-sight perception and secondary calibration of traffic lights. By combining vehicle-side sensor information with the recognition results, the accuracy and robustness of traffic light recognition are improved.
It effectively improves the accuracy and robustness of traffic light recognition, reduces sensitivity to environmental changes, reduces false detections and missed detections, and improves the safety and reliability of autonomous vehicles.
Smart Images

Figure CN122116319A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, particularly to intelligent transportation, computer vision and other fields, and can be used in application scenarios such as autonomous driving. Specifically, it relates to traffic light recognition methods, devices, electronic devices and storage media. Background Technology
[0002] In the field of autonomous driving, existing technologies mainly employ three approaches to accurately identify traffic lights in complex road scenarios: First, camera-based visual recognition schemes, which analyze images to detect information such as the color, shape, and direction of traffic lights. These are low-cost but susceptible to ambient light and occlusion. Second, radar-visual fusion schemes, combining lidar positioning with camera semantic recognition, improve recognition reliability at night or in strong light. Third, vehicle-to-infrastructure (V2I) schemes, which directly acquire traffic light status information through vehicle-to-infrastructure communication, enabling beyond-line-of-sight perception, but rely on roadside infrastructure. These approaches still have limitations in terms of environmental adaptability, recognition accuracy, and infrastructure dependence in complex road scenarios. Summary of the Invention
[0003] This disclosure provides a traffic light identification method, apparatus, electronic device, and storage medium.
[0004] According to a first aspect of this disclosure, a traffic light recognition method is provided, comprising: acquiring a current frame image containing traffic lights; inputting the current frame image into a pre-trained traffic light position detection model to obtain an initial position prediction result for at least one traffic light in the current frame image; matching the initial position prediction result with historical position prediction results of corresponding traffic lights in at least one historical frame, and correcting the initial position prediction result according to the matching result to obtain a traffic light position prediction result; and recognizing the traffic lights in the current frame image based on the traffic light position prediction result to obtain a traffic light recognition result.
[0005] According to a second aspect of this disclosure, a traffic light recognition method is provided, comprising: acquiring a first recognition result and a second recognition result for a target traffic light; the first recognition result being a roadside recognition result obtained by the method of the first aspect; the second recognition result being a vehicle-side recognition result obtained based on an on-board sensor; and fusing the first recognition result and the second recognition result to obtain a fused recognition result for the target traffic light.
[0006] According to a third aspect of this disclosure, a traffic light recognition device is provided, comprising: an image acquisition module for acquiring a current frame image containing traffic lights; a position retrieval module for inputting the current frame image into a pre-trained traffic light position detection model to obtain an initial position prediction result for at least one traffic light in the current frame image; a history correction module for matching the initial position prediction result with the historical position prediction results of corresponding traffic lights in at least one historical frame, and correcting the initial position prediction result according to the matching result to obtain a traffic light position prediction result; and a result generation module for recognizing the traffic lights in the current frame image based on the traffic light position prediction result to obtain a traffic light recognition result.
[0007] According to a fourth aspect of this disclosure, a traffic light recognition device is provided, comprising: a result acquisition module, configured to acquire a first recognition result and a second recognition result for a target traffic light; the first recognition result is a roadside recognition result obtained by the method of the first aspect; the second recognition result is a vehicle-side recognition result obtained based on an on-board sensor; and a result fusion module, configured to fuse the first recognition result and the second recognition result to obtain a fused recognition result for the target traffic light.
[0008] According to a fifth aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform any of the methods described in the embodiments of this disclosure.
[0009] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of this disclosure.
[0010] According to a seventh aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of this disclosure.
[0011] According to the eighth aspect of this disclosure, an autonomous vehicle is provided, including electronic devices as described in the fifth aspect.
[0012] The solution disclosed herein can reuse existing intersection surveillance cameras to achieve beyond-line-of-sight perception and secondary calibration, effectively improving the accuracy and robustness of vehicle traffic light recognition.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic flowchart of a traffic light recognition method applied to roadside or cloud-based systems according to an embodiment of this disclosure; Figure 2 This is a schematic flowchart of a traffic light recognition method applied to a vehicle according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram illustrating the application of the traffic light recognition method according to an embodiment of the present disclosure; Figure 4 This is another schematic flowchart of the traffic light recognition method according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the feature fusion process according to an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of a traffic light recognition device applied to the roadside or cloud-based system according to an embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of a traffic light recognition device applied to a vehicle according to an embodiment of the present disclosure; Figure 8 This is a schematic diagram of a scenario for a traffic light recognition method applied to a roadside or cloud-based system according to an embodiment of this disclosure; Figure 9 This is a schematic diagram of a scenario for a traffic light recognition method applied to a vehicle, according to an embodiment of this disclosure. Figure 10 This is a structural diagram of an electronic device used to implement the traffic light recognition method of the embodiments of this disclosure. Detailed Implementation
[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0016] Before introducing the technical solutions of the embodiments of this disclosure, the technical terms that may be used in this disclosure will be further explained: Traffic lights refer to road traffic lights, namely red, yellow, and green lights and their extended forms, which are set up at intersections, road junctions, ramps, etc., to direct the passage of vehicles and pedestrians. These include straight-ahead and turning arrow lights, lane indicator lights, pedestrian crossing lights, and traffic signal display devices with countdown numbers.
[0017] Among related technologies, pure vision solutions are highly sensitive to lighting conditions, exhibiting decreased recognition stability in environments such as strong light, backlight, rain, fog, and nighttime. Furthermore, due to their primary reliance on monocular depth estimation, their accuracy in estimating the distance and 3D position of traffic lights in long-distance scenes is insufficient, making them susceptible to occlusion. Additionally, bright targets such as vehicle lights, billboards, and screens are easily misidentified as traffic lights. The laser-vision fusion solution's reliance on high-cost sensors like LiDAR significantly increases the overall vehicle hardware cost. Moreover, the large data bandwidth of multiple sensors, coupled with complex time synchronization and fusion calculations, results in high system resource consumption and a heavy system weight. Vehicle-to-everything (V2X) solutions heavily rely on the large-scale deployment of roadside infrastructure and communication networks, leading to high construction and maintenance costs. Currently, they are mainly concentrated in demonstration areas and some smart highway pilot sections. Furthermore, different regions have not yet fully unified protocols, message formats, and security authentication, resulting in poor cross-regional compatibility. Simultaneously, most vehicles generally lack onboard units, limiting system coverage. Strict encryption and authentication mechanisms are also required to address security issues such as spoofed signals, malicious attacks, and privacy leaks.
[0018] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, this disclosure proposes a traffic light recognition method that can reuse existing intersection monitoring cameras, achieve beyond-line-of-sight perception and secondary calibration, and effectively improve the accuracy and robustness of vehicles in traffic light recognition.
[0019] This disclosure provides a traffic light recognition method applicable to roadside or cloud-based systems. Figure 1 This is a flowchart illustrating a traffic light recognition method applied to a roadside or cloud-based system according to embodiments of this disclosure. This traffic light recognition method can be applied to a traffic light recognition device. The traffic light recognition device is located in an electronic device. This electronic device includes, but is not limited to, fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or ordinary servers. Mobile devices include, but are not limited to, autonomous driving devices, which can be mobile phones, tablets, etc. In some possible implementations, the traffic light recognition method can also be implemented by a processor calling computer-readable instructions stored in memory. Figure 1 As shown, the traffic light recognition method includes: S101. Obtain the current frame image containing the traffic lights.
[0020] S102. Input the current frame image into the pre-trained traffic light position detection model to obtain the initial position prediction result of at least one traffic light in the current frame image.
[0021] S103. Match the initial position prediction result with the historical position prediction result of the corresponding traffic light in at least one historical frame, and correct the initial position prediction result according to the matching result to obtain the traffic light position prediction result.
[0022] S104. Based on the traffic light position prediction result, identify the traffic lights in the current frame image to obtain the traffic light identification result.
[0023] In this embodiment of the disclosure, an existing camera can be used in the target intersection or road scene to continuously capture video of the area to be detected. The camera outputs the original video stream according to a preset frame rate. After receiving the video stream, the video is parsed frame by frame according to the time sequence, and the corresponding single frame image is extracted from the video stream at the current processing time and the frame is used as the current frame image.
[0024] In this embodiment of the disclosure, the current frame image can be input into a pre-trained traffic light location detection model according to a predetermined input format. The traffic light location detection model can employ a convolutional neural network-based target detection structure. During the forward inference process, it performs feature extraction and multi-scale target search on the current frame image, automatically identifies potential traffic light candidate regions in the image, outputs the bounding box position parameters of each candidate target, and uses these detection boxes as the initial position prediction results for at least one traffic light in the current frame.
[0025] Here, the historical position prediction results refer to the traffic light position prediction results obtained after processing in at least one historical frame before the current frame. These results include information such as the image coordinate position, size, confidence level, and optional target identifier of the corresponding traffic light at different historical time points, which are used to characterize the movement trajectory or position change of the traffic light in the time series.
[0026] In this embodiment, a cache queue containing multiple historical frames can be pre-maintained. The queue records the traffic light position prediction results and their timestamps or frame numbers corresponding to each historical frame. Further, based on chronological order, the results of one or more historical frames adjacent to the current frame can be selected from the historical cache. Using the spatial proximity, shape similarity, and possible target identification information of candidate targets as criteria, the initial detection box of the current frame is matched one-to-one or one-to-many with the corresponding traffic light positions in the historical frames. After matching, based on the continuity of the historical trajectory and motion constraints, abnormal detection boxes deviating from the historical trajectory in the current frame can be filtered or their positions corrected. Finally, the corrected traffic light position prediction result for the current frame is output.
[0027] In this embodiment, based on the traffic light position prediction result, the corresponding region in the current frame image is cropped according to the coordinates of each predicted box to obtain multiple local traffic light sub-images. Each sub-image is then input into a pre-trained traffic light state recognition model. This model can be a classification network for traffic light color and shape, or a multi-task network that simultaneously recognizes multiple states. Further, the recognition model extracts and discriminates features from the local images, outputting the category label and corresponding probability value of each traffic light's current state. It then organizes the state category, confidence level, and associated location and time information of each traffic light into structured data, which serves as the traffic light recognition result for the current frame.
[0028] The technical solution of this disclosure realizes a complete traffic light recognition link on the roadside, from video frame acquisition, target position detection, temporal position correction to state recognition. It uses historical position prediction results to match and correct the initial detection results of the current frame, effectively introducing a continuity constraint in the time dimension. This not only alleviates the detection jitter and false detection problems in complex scenarios such as changes in lighting, occlusion, and long-distance imaging, but also significantly improves the stability and accuracy of traffic light position estimation. Furthermore, it can provide a more accurate target area for state recognition, thereby improving the reliability and anti-interference capability of traffic light recognition. This results in an overall improvement in the accuracy and robustness of traffic light recognition in complex traffic environments, providing safer and more reliable traffic light perception results for autonomous vehicles or vehicle-road cooperative systems.
[0029] In some embodiments, acquiring the current frame image containing the traffic lights includes: acquiring a continuous video stream or image sequence; and selecting a target frame from the video stream or image sequence as the current frame image according to a preset sampling interval.
[0030] In this embodiment, existing camera devices can be selected from the target intersection or road segment, and connected to the roadside processing unit or cloud processing unit via wired or wireless network. The camera continuously captures images of the monitored area according to a preset frame rate and resolution, continuously outputting video signals. After receiving the video signal, the roadside processing unit or cloud processing unit performs necessary encoding and decoding processing, parsing the compressed video data into frame-by-frame image data, and buffering or storing them in chronological order to form a video stream or image sequence.
[0031] In this embodiment, a sampling interval parameter can be preset. This parameter can be a fixed number of frame intervals or a fixed time interval, and the frame number or timestamp of each frame is recorded while receiving or reading continuous image data. Furthermore, as image data in the video stream or image sequence continuously arrives, it can be determined whether the preset sampling interval threshold has been reached based on the difference between the frame number of the current frame and the frame number of the previously selected target frame, or based on the difference between the current time and the previously selected time. When the sampling condition is met, the corresponding image frame is marked as the target frame and extracted from the continuous data stream as the current frame image. In particular, the sampling interval can also be dynamically adjusted in conjunction with factors such as intersection signal timing and vehicle flow.
[0032] In this way, representative keyframes can be selectively extracted from a large amount of raw image data while ensuring the continuity of the time series. This avoids the excessive computational overhead of detecting and recognizing traffic lights in every frame, and ensures that the selected target frames are sufficiently dense in time to avoid missing critical events such as changes in traffic light status. By pre-setting or adaptively adjusting the sampling interval, the recognition frequency and system real-time performance can be flexibly controlled according to the timing characteristics of traffic lights, vehicle speed, and system processing capabilities in the actual scenario. This results in a better overall effect between resource consumption, response speed, and recognition stability in the entire traffic light recognition process.
[0033] In some embodiments, the traffic light location detection model is obtained by: acquiring an image training set containing traffic lights; the image training set containing traffic light location annotation information; and using the image training set to train the model to be trained to obtain the traffic light location detection model.
[0034] Here, the image training set refers to a set of labeled image data used to train the traffic light location detection model. Each image in the set contains the visible area of a road traffic light and is accompanied by precise location labeling information for each traffic light target.
[0035] In this embodiment, a large number of raw images containing road traffic lights can be acquired first. These raw images can be derived from historical video frames extracted from roadside cameras, photos captured from real vehicles or simulated scenes, or publicly available datasets. Specifically, the acquisition process should cover as many typical operating conditions as possible, such as daytime, nighttime, backlighting, rain and snow, different urban roads, different installation heights, and different traffic light styles, to improve data representativeness. Subsequently, these raw images are manually or semi-automatically annotated using annotation tools. The positions of all visible traffic lights are bounded in each image, and their bounding box coordinates are recorded. Finally, these image files and their corresponding annotation information are organized and stored in a unified format to form an image training set for training purposes.
[0036] In this embodiment, the model structure to be trained can be constructed first, such as a single-stage detection network or a two-stage detection network based on a convolutional neural network, and the model parameters can be initialized. Subsequently, the labeled images in the image training set can be divided into a training set and a validation set according to a certain ratio. During the training phase, each image in the training set and its corresponding location label information are used as a set of samples input into the model. The model's prediction result for the traffic light position in the image is calculated through forward propagation. A loss function is constructed based on the deviation between the predicted position and the actual labeled position. Combining multiple losses such as target classification and bounding box regression, the model parameters are backpropagated and gradients are updated. After multiple rounds of iterative training, when the model's performance on the validation set reaches the preset requirements, its parameters are fixed, and this model is used as the final traffic light position detection model for predicting the traffic light position in the current frame image of a road scene during actual operation.
[0037] Thus, high-quality and diverse supervised data are introduced during the model training phase, enabling the model to learn the appearance features and spatial distribution patterns of traffic lights in different scenarios under the constraint of a large number of labeled samples. This significantly improves the model's detection accuracy and robustness for traffic light targets. Precise location labeling allows the loss function to measure the deviation between the predicted and ground truth boxes in a fine-grained manner, promoting the model's convergence in bounding box regression and resulting in a traffic light location detection model with more accurate localization and lower false positive and false negative rates.
[0038] In some embodiments, the traffic light recognition method further includes: generating new training data based on the matching results and the corresponding traffic light position prediction results of the current frame image; and using the new training data to train or fine-tune the traffic light position detection model.
[0039] In this embodiment, the matching results can first be used to filter and evaluate the quality of each traffic light position prediction result in the current frame, filtering out obviously unstable or suspicious detection boxes and retaining only those prediction results that are continuous in time series, have smooth spatial location changes, and have high confidence. Subsequently, the current frame image and the retained traffic light position prediction results can be organized into a standardized training sample format, thereby forming one or more new training sample records. Specifically, these newly generated samples can be appended to the original training dataset, or a separate incremental dataset can be created as an input data source for subsequent model training or fine-tuning.
[0040] In this embodiment of the disclosure, during the training or fine-tuning phase, new training data can be input into the model in batches with a small learning rate and an appropriate number of training rounds while maintaining the original parameters of the model. Through forward inference and loss calculation, the model parameters can be fine-tuned based on the latest environmental characteristics reflected by the new data, while maintaining the original performance.
[0041] Thus, the newly generated training data based on the matching results has high reliability and can be used as high-quality pseudo-labels to expand the training samples, enabling the model to continuously adapt to more diverse traffic light samples in real road scenarios and reduce performance degradation caused by scene migration. Through continuous incremental training or fine-tuning, the model can gradually adapt to the perspective, resolution, lighting conditions, and local traffic light shape differences of specific roadside cameras, continuously optimizing its detection accuracy and robustness in long-term operation, thereby significantly improving the stability and reliability of traffic light position detection in complex and variable road environments.
[0042] In some embodiments, matching the initial position prediction result with the historical position prediction result of the corresponding traffic light in at least one historical frame includes: calculating the overlap between the initial position prediction result and each historical position prediction result; and determining the historical position prediction result with an overlap greater than or equal to a preset overlap threshold as the target historical position prediction result that matches the initial position prediction result.
[0043] Here, overlap refers to the degree to which the initial position prediction result of a traffic light in the current frame overlaps with the historical position prediction result of the corresponding traffic light in at least one historical frame on the image plane. It is used to measure the similarity or consistency of two candidate location boxes in spatial location.
[0044] In this embodiment, the initial position prediction result of a traffic light to be matched in the current frame image can be obtained first. Simultaneously, the historical position prediction results of the corresponding traffic light in one or more historical frames temporally associated with the initial position prediction result are retrieved from the historical frame buffer, and these historical prediction boxes are uniformly mapped or represented in the coordinate system of their respective images. Next, for each historical position prediction box, one by one, each historical position prediction box can be selected, and the intersection region of the two rectangles on the image plane can be determined through geometric calculations. If the initial position prediction box intersects with any historical position prediction box, the area of the intersection region is calculated, and the areas of the two rectangles and their union area are also calculated. Finally, the overlap value can be recorded as the ratio of the intersection area to the union area or other equivalent formulas. Specifically, if the two rectangles do not intersect, the overlap is directly recorded as zero.
[0045] In this embodiment of the disclosure, after calculating the overlap between the prediction result of a certain initial position in the current frame and all historical position prediction results, the overlap list corresponding to the initial prediction box can be traversed. All historical position prediction results with an overlap value greater than or equal to a preset threshold are filtered out and marked as the target historical position prediction result of the initial position prediction result. In particular, the preset overlap threshold can be used to distinguish between the same target with similar spatial heights and different targets with inconsistent positions. The specific value can be determined based on factors such as the actual scene, camera viewpoint, and the movement range of the traffic lights.
[0046] Thus, explicit spatial geometric constraints are introduced into the association process of multi-frame temporal information. This enables a simple, intuitive, and computationally low-cost way to effectively associate the initial position prediction result of the current frame with candidate positions in historical frames, avoiding mismatches and trajectory drift problems that may occur when matching based solely on distance or confidence. By setting a reasonable threshold for overlap, only historical position prediction results with high spatial overlap are retained as matching targets. This effectively filters out accidental false detection boxes, noisy boxes far from the trajectory, and rapidly changing abnormal detection results, thereby enhancing the tracking continuity and positional stability of the same traffic light target in the temporal dimension.
[0047] In some embodiments, the traffic light recognition method further includes: when the overlap between the initial position prediction result and each historical initial position prediction result is less than a preset overlap threshold, determining that the current frame image and / or traffic light is abnormal, and generating a warning message.
[0048] In this embodiment, the overlap value can be compared with a preset overlap threshold. When it is found that the overlap between the predicted initial position of a certain current frame and all historical initial position predictions is less than the preset overlap threshold, it is determined that the initial detection position is seriously inconsistent with the historical trajectory. This leads to the conclusion that there is an anomaly in the current frame image and / or the corresponding traffic light. For example, it may be due to severe camera shake, obstruction, strong light interference causing imaging abnormalities, or physical anomalies such as sudden changes in traffic light position, damage, or displacement. After the anomaly determination is completed, corresponding warning information can be generated. The warning information may include at least the timestamp of the anomaly occurrence, the camera or roadside equipment identifier, the anomaly type, the coordinates of the relevant initial position prediction box, and historical trajectory features. The warning information is then reported in real time to the management platform, traffic control center, or vehicle-road cooperative system according to a predetermined communication protocol.
[0049] In this way, anomalies can be detected in a timely manner when the model detection results deviate significantly from the historical trajectory. This effectively avoids providing misleading traffic light location information to autonomous vehicles or traffic control systems when the perceived information is obviously unreliable. It can not only quickly prompt maintenance personnel to conduct troubleshooting and calibration under complex operating conditions, but also trigger safety redundancy strategies in the upper-level logic. This improves the overall safety, reliability, and maintainability of the roadside traffic light recognition system, providing a more robust perception foundation for vehicle-road cooperation and intelligent transportation applications.
[0050] In some embodiments, based on the traffic light position prediction result, the traffic lights in the current frame image are identified to obtain a traffic light identification result, including: extracting the corresponding traffic light region image from the current frame image based on the traffic light position prediction result; performing feature extraction on each traffic light region image to obtain a traffic light feature map set; the traffic light feature map set includes a traffic light feature map corresponding to each traffic light in the current frame image; and determining the structured information of each traffic light based on the traffic light feature map set as the traffic light identification result.
[0051] In this embodiment, a cropping operation can be performed within the image based on the top-left corner coordinates, width, and height of each predicted bounding box in the traffic light position prediction result, within the original image coordinate system of the current frame. This extracts the rectangular region corresponding to the predicted bounding box from the entire frame image, resulting in multiple appropriately sized local image blocks, thereby forming a traffic light region image with uniform format and stable quality. Specifically, the predicted bounding box boundaries can be appropriately expanded outwards during cropping to ensure the traffic light is completely contained within the region. Furthermore, the cropped region image is uniformly scaled to a predetermined input size and preprocessed using methods such as normalization and color space conversion.
[0052] In this embodiment, each traffic light region image can be sequentially input into a pre-constructed feature extraction network or module. The feature extraction network then uses operations such as multi-layer convolution, non-linear activation, pooling, or self-attention calculation to gradually extract high-dimensional feature representations from the original pixel space. The output can be a multi-channel feature map that preserves the spatial structure, or a feature vector after global pooling. Finally, the feature outputs obtained from each traffic light region image can be saved one-to-one according to the traffic light instance to form a traffic light feature map set.
[0053] Here, structured information refers to a set of machine-readable and storable field information output for each traffic light after analyzing and decoding its feature map. In this embodiment, the structured information includes, but is not limited to, the current state of the traffic light, traffic light type, location information, reliability score, and optional phase number, countdown, etc.
[0054] In this embodiment, for each feature record in the traffic light feature map set, the features can be further analyzed and classified through a series of fully connected layers, convolutional layers, or attention modules. For example, the current state category of the traffic light can be determined (e.g., red light, green light, yellow light, left turn arrow, straight arrow, flashing state, etc.), and additional attributes (e.g., visibility, occlusion, brightness level, etc.) can be estimated. Furthermore, a fine-grained position correction aligned with the original image coordinate system or a logical position related to the road topology (associated with lane number, direction, etc.) can be output. The analysis and classification results are converted into discrete labels or continuous numerical values and combined according to a predefined data structure to form a structured information record for the traffic light. Specifically, the structured information may include fields such as traffic light identification number or index, position in the current frame, state category, direction attribute, confidence level, and timestamp.
[0055] Thus, performing subsequent high-precision feature analysis only within the predicted traffic light area can significantly reduce background interference and computational burden, improving the accuracy of the recognition model in judging traffic light status in scenarios with small or multiple targets. The traffic light feature atlas constructed through a deep feature extraction network can fully utilize multi-scale, multi-channel representation capabilities, enhancing the robustness of visual features under complex lighting conditions, inclement weather, and aging or contaminated traffic lights. Designing the recognition results as structured information output allows key attributes such as traffic light status, position, and direction to be directly used in a field-based format, facilitating rapid integration with intersection timing schemes, lane topology information, and vehicle driving strategies, thereby improving the overall system's automated processing capabilities and engineering usability.
[0056] In some embodiments, determining the structured information of each traffic light based on a traffic light feature map set includes: obtaining an enhanced feature map with temporal information based on the traffic light feature map and historical traffic light feature maps; and determining the structured information of each traffic light based on the enhanced feature map.
[0057] Here, temporal information refers to the feature sequence and temporal relationship formed by the changes of the same physical traffic light over time in multiple consecutive frames of images. In this embodiment of the disclosure, the temporal information may include information such as the feature similarity and change trend of the traffic light in adjacent frames, the color or lighting state switching process, the continuity of position and appearance over time, and the order, interval and duration of these features in the time dimension.
[0058] In this embodiment, the traffic light in the current frame can first be associated with the historical feature maps of the same physical traffic light in several historical frames to construct a feature sequence arranged in chronological order. Subsequently, a temporal modeling module can be introduced into this sequence, such as using temporal convolution, recurrent neural networks, temporal attention mechanisms, or a Transformer-based temporal encoding mechanism, to fuse and weight the features of each frame in the temporal dimension, thereby encoding the current frame features themselves and the features of historical frames and their changing trends into a new high-dimensional representation. Specifically, the sequential relationship and duration between features can be characterized through explicit temporal position encoding or implicit recursive state propagation. The final output is an enhanced feature map with temporal information. This feature map retains the spatial and semantic information of the current frame while additionally carrying contextual information from multiple consecutive historical frames, thus more robustly reflecting the true state of the traffic light.
[0059] In this embodiment, an enhanced feature map with temporal information can be input into a decoding and classification module, where multiple output branches are constructed for state recognition, type recognition, and additional attribute estimation. For example, in the state recognition branch, the temporal continuity information contained in the enhanced feature map can be used to determine the current state of the traffic light. For instance, if the light spot is weak in a single frame but shows a bright feature in multiple consecutive frames, the traffic light is likely to be identified as a certain bright color state; or if periodic on / off changes are observed in multiple frames, it is identified as a flashing state. Further, after forward reasoning on the enhanced feature map, the output branches provide the state category, arrow direction, validity flag, confidence level, and necessary coordinate correction information. These output results are then organized according to a pre-defined field structure to form a structured information record corresponding to each traffic light, including the traffic light's spatiotemporal location, current state, directional attribute, reliability index, etc., ultimately serving as the traffic light recognition result for the current frame.
[0060] Thus, by fusing information from multiple frames, the enhanced feature map can effectively smooth out occasional false detections or jitter, thereby significantly reducing single-frame misjudgments and state jumps, and improving recognition stability and robustness in complex environments. Utilizing temporal information can more accurately identify signal behaviors that are inherently time-patterned, such as traffic light switching and yellow light flashing, and can still output reasonable and consistent structured results at phase boundaries and during flashing, avoiding instantaneous misjudgments.
[0061] In some embodiments, determining the structured information of each traffic light based on the enhanced feature map includes: detecting the traffic light group based on the enhanced feature map to obtain the group position frame of each traffic light; detecting the lamp head of the traffic light within the group position frame to obtain the lamp head position frame of each lamp head; identifying the lighting state of the lamp head based on the lamp head position frame; and obtaining structured information based on the lamp head position frame and the lighting state of each lamp head in each traffic light group.
[0062] Here, a traffic light group refers to a set of lights belonging to the same signaling device and facing the same lane or the same traffic participant. Examples include a typical red, yellow, and green traffic light arranged vertically, or an arrow traffic light group consisting of straight, left-turn, and right-turn arrow lights. The light group location bounding box refers to the detection result in an image or feature map, using a rectangular bounding box to describe the spatial location and extent of the entire traffic light group. Its coverage includes all lights in the group and their adjacent structures.
[0063] In this embodiment, a detection sub-network can be pre-constructed, taking the enhanced feature map as input and outputting a series of candidate light group boxes and their confidence scores. Subsequently, post-processing methods such as non-maximum suppression can be used to deduplicate and filter the candidate light group boxes, retaining only a group of rectangular boxes with high confidence, high ranking, and low overlap. These boxes are then regarded as the light group position boxes of each traffic light group in the current frame, with each position box corresponding one-to-one with a logical traffic light entity.
[0064] In this embodiment, a feature sub-image for each light group can be obtained by cropping the corresponding local region from the image or feature map. Then, a finer-grained light head detection sub-network is constructed. The feature sub-image for each light group is input into this sub-network to perform target detection and localization on multiple light heads within the light group. The output of this sub-network is a set of smaller rectangular bounding boxes located within the same light group location frame. Each bounding box corresponds to a specific light head, ultimately resulting in a set of several light head location frames within each light group location frame.
[0065] Here, "lit status" refers to the determination of whether a single lamp head is currently emitting light, at least distinguishing between on and off, and may also include the color and brightness information of the lamp head.
[0066] In this embodiment of the present disclosure, for each lamp head location frame, the corresponding local area of the lamp head can be cropped again from the original image or high-resolution feature map, and these local images of the lamp head can be input into a dedicated lamp head state recognition and classification network. The lamp head state recognition and classification network will focus on analyzing the brightness distribution, color composition, edge contour and contrast with the background of the lamp head area, and output the category result of whether the lamp head is currently lit. Then, this result is simplified or normalized into a lighting status mark, which is stored together with the lamp head location frame as the lighting status of the lamp head.
[0067] In this embodiment, the detected position frames of all light heads within each light group are first organized and sorted. Then, based on their spatial relationships within the light group and the light group structure rules learned during the training phase, each light head is assigned a semantic role. Subsequently, these semantic roles are combined with the light head illumination status obtained in the previous step to deduce the overall traffic signal meaning corresponding to the light group. Finally, the light group position, light group type, spatial position and illumination status of each light head, and the inferred traffic meaning of each traffic light group are encapsulated into a structured information record. The structured information set of all light groups is output as the final structured information for traffic light recognition in that frame.
[0068] Thus, by using the light group's location bounding box to perform coarse-grained scene segmentation, background areas unrelated to the current target can be effectively eliminated, reducing interference from non-signal targets. Through precise positioning and illumination status determination of individual light heads within the light group, various complex signal combinations can be identified, and the specific meaning of passage can be deduced in a clear and interpretable manner, enhancing the reliability and traceability of the recognition results. The hierarchical structure facilitates migration and expansion between different intersection structures and different signal light types. Only adjustments to the light group's structural rules or fine-tuning of the light head status classifier are needed to adapt to various engineering scenarios, thereby significantly improving the accuracy, robustness, and engineering application value of the traffic light recognition system in complex and changing road environments while ensuring real-time performance.
[0069] This disclosure provides a traffic light recognition method applied to a vehicle. Figure 2 This is a flowchart illustrating a traffic light recognition method applied to a vehicle according to an embodiment of this disclosure. This traffic light recognition method can be applied to a traffic light recognition device. The traffic light recognition device is located in an electronic device. The electronic device includes an autonomous driving device, which can be an in-vehicle terminal. In some possible implementations, the traffic light recognition method can also be implemented by a processor calling computer-readable instructions stored in memory. For example... Figure 2 As shown, the traffic light recognition method includes: S201, Obtain the first recognition result and the second recognition result for the target traffic light. The first recognition result is... Figure 1 The first is the roadside recognition result obtained by the traffic light recognition method shown; the second recognition result is the vehicle-side recognition result obtained based on the on-board sensor.
[0070] S202. The first identification result and the second identification result are fused to obtain the fused identification result for the target traffic light.
[0071] Here, the first recognition result refers to the recognition result obtained on the roadside based on the aforementioned method of this disclosure for the target traffic light. The second recognition result refers to the recognition result obtained by the vehicle itself based on the image or point cloud data perceived by the vehicle-mounted sensors, after processing by the vehicle-mounted traffic light recognition algorithm for the target traffic light. In the embodiments of this disclosure, both the first and second recognition results are structured information in data form. The first recognition result comes from a fixed viewpoint on the roadside, and the second recognition result comes from the viewpoint of the vehicle moving with the vehicle.
[0072] In this embodiment, the vehicle receives periodically identification messages from the roadside via Vehicle to Everything (V2X) wireless communication technology. These messages contain structured identification results from the roadside for each intersection's traffic lights. After parsing the message, the vehicle identifies the target traffic light most relevant to its vehicle and marks the corresponding roadside identification result as the first identification result. Further, the vehicle utilizes onboard sensors and a vehicle-side traffic light detection and identification algorithm to detect and identify traffic lights related to the target intersection in either the onboard coordinate system or the image coordinate system. It then determines the traffic light instance corresponding to the target traffic light, outputs the structured identification result of that instance, and marks it as the second identification result.
[0073] In this embodiment, a first and second identification result within the same or similar time window can be selected based on timestamps. After pairing, a confidence-weighted fusion strategy can be adopted, setting basic weights for roadside identification and vehicle-side identification respectively, and dynamically adjusting the confidence weights of both based on current environmental conditions. Specifically, when the identification states of the two are consistent, the confidence of the final result can be improved through weighted averaging or confidence enhancement, and a weighted summation or optimal estimation based on an error model can be performed by combining the position estimates of both sides to obtain a more accurate target location. Furthermore, when the identification states of the two are inconsistent, conflict resolution can be performed according to predefined rules. For example, when the vehicle-side field of view is significantly limited, roadside identification can be prioritized; when the vehicle-side observation is clear at close range and the confidence is significantly higher than that of the roadside, vehicle-side identification can be prioritized; or when the conflict cannot be determined, the result can be marked as uncertain and the confidence reduced. Finally, the fused traffic light position, state, direction, etc., can be organized into a unified structured record as the fused identification result.
[0074] The technical solution of this disclosure achieves multi-source perception and intelligent fusion of the same target traffic light on the vehicle side, fully leveraging the advantages of both the roadside and the vehicle side, thereby significantly improving the reliability and safety of traffic light recognition under various complex conditions. The fusion mechanism, through consistency checks and conflict resolution of the two types of results, can trigger more conservative driving strategies or warnings when serious inconsistencies are detected, improving the safety redundancy and overall robustness of autonomous driving or advanced driver assistance systems in intersection scenarios.
[0075] In some embodiments, the traffic light recognition method further includes: acquiring the vehicle's location information and attitude information; determining the current driving lane based on map data and the location information and attitude information; and determining the target traffic light based on the current driving lane.
[0076] Here, location information refers to the description of the vehicle's current position in geographic space, which can be two-dimensional or three-dimensional coordinates in a map coordinate system, and may also include information such as position accuracy and error. Attitude information refers to the vehicle's orientation and attitude state in space, which may include heading angle, pitch angle, roll angle, etc.
[0077] In this embodiment, positioning technology can be used to obtain the vehicle's approximate geographical location, latitude, longitude, and altitude information. Simultaneously, acceleration and angular velocity data from the inertial measurement unit and wheel odometer information are collected. Algorithms are then used to estimate the relative displacement and turning angle over a short period. Specifically, lane lines and curbs identified by cameras, or LiDAR point clouds, can be matched with a high-definition map to achieve high-precision positioning. Finally, the aforementioned multi-source information can be unified into a single map coordinate system to obtain the vehicle's current position coordinates and corresponding attitude parameters.
[0078] Here, map data refers to pre-built road maps or lane-level maps, which may include structured information such as the geometry of roads and lanes, lane centerlines and boundaries, lane numbers, intersection topology, and the installation locations of traffic lights and the relationships between control lanes.
[0079] In this embodiment, road and lane geometry information within a certain range near the vehicle's current location can first be loaded from a map database, including the spatial trajectory of each lane's centerline, lane width, and lane direction. Then, in the map coordinate system, the vehicle's current location can be projected onto the centerline or lane polygon area of the candidate lanes. The lateral distance between the vehicle and each lane's centerline, the longitudinal projection position, and the angle between the vehicle's heading and the lane direction can be calculated. The lane with the smallest lateral distance, the smallest angle between the heading and the lane direction, and the vehicle's location falling within the effective lane range can then be preferentially selected as the current driving lane.
[0080] In this embodiment, a set of all traffic lights that have a control relationship with the current driving lane can be queried on a map. Based on the vehicle's driving direction and a predetermined path, traffic lights unrelated to the current driving direction are removed from the set. Subsequently, according to the vehicle's longitudinal position in the lane, traffic lights located in front of the vehicle and within a preset distance are considered valid candidates. These are then sorted by distance or estimated arrival time, and the nearest traffic light that is at the main control position of the current intersection on the road topology is selected and marked as the target traffic light.
[0081] Thus, by introducing map and lane semantic constraints, the system can automatically filter out one or a few target traffic lights from numerous traffic lights that are truly relevant to the vehicle's current lane and direction of travel. This effectively avoids erroneous decisions caused by vehicles misreading traffic lights in adjacent lanes, oncoming lanes, or dedicated turning lanes. Lane-level localization and target traffic light selection provide a clear and consistent reference for the subsequent matching and fusion of roadside recognition results and vehicle-side recognition results, significantly reducing ambiguity in complex intersection scenarios with multiple traffic lights and improving accuracy, reliability, and interpretability.
[0082] In some embodiments, the first identification result and the second identification result are fused to obtain a fused identification result for the target traffic light, including: reading the second identification result, and when the second identification result is empty, using the first identification result as the fused identification result.
[0083] In this embodiment, a second recognition result for the target traffic light can be read, and the validity of the second recognition result can be checked. If no vehicle-side recognition output is received within the current time window, or the read data is empty, expires due to timeout, or only contains placeholder markers without valid status information, then the second recognition result is determined to be empty. Furthermore, in the case of an empty result, the fusion strategy is not executed; instead, the currently acquired first recognition result is directly called and output as a fused recognition result.
[0084] In this way, it effectively ensures that the signal light perception capability remains continuous and stable even when the vehicle-side recognition is abnormal or the signal is not triggered, thus avoiding interruption of the recognition link due to waiting for the vehicle-side result or forcibly performing invalid fusion.
[0085] In some embodiments, the first identification result and the second identification result are fused to obtain a fused identification result for the target traffic light, including: reading the second identification result; when the second identification result is not empty, calculating the consistency score between the first identification result and the second identification result; when the consistency score is greater than or equal to a preset threshold, taking the common state of the first identification result and the second identification result as the fused identification result.
[0086] In this embodiment, a second identification result for a target traffic light can be read and its validity checked. If the second identification result is not expired in time, has complete fields in structure, and is semantically usable for fusion, several key fields from the first and second identification results are selected for comparison, and the comparison result is converted into a consistency score. For example, the core light status fields of the two can be compared first, such as light color and traffic meaning, to see if they are completely consistent. If they are consistent, a higher base score is assigned to the consistency score. Second, the traffic light position and orientation fields are compared, and additional points are given to the one with smaller differences. Third, the internal confidence levels of both can be considered, and a higher consistency score is given to results with high confidence and mutual consistency, while the score is appropriately reduced for cases where one party has low confidence. Finally, the above comparison results can be combined into a single consistency score using a pre-designed formula.
[0087] In this embodiment of the disclosure, a common state refers to the set of traffic light states that are simultaneously determined to be the same in the first identification result and the second identification result, that is, the part of the state information that is the same in the intersection of the state fields of the two.
[0088] In this embodiment, the consistency score can first be compared with a preset threshold. When the score reaches or exceeds the threshold, it is determined that the roadside recognition result and the vehicle-side recognition result are highly consistent overall, and there is no substantial conflict in their perception of the target traffic light's state. Further, mutually agreed-upon state information can be extracted from the first and second recognition results as a common state, thereby obtaining the fused recognition result. Specifically, for structured results containing multiple fields, only completely consistent fields can be included in the common state, while differing fields can be marked as uncertain or processed separately.
[0089] In this way, the consistency score quantitatively evaluates the recognition results of the roadside and vehicle sides across multiple dimensions, including light status, location, and confidence level. This effectively avoids the amplification of errors caused by blindly merging results when potential conflicts or deviations exist. By extracting common states as the fusion output, it ensures that the final signal light status only contains the information recognized by both independent sensing sources, making the results more robust and interpretable. The preset threshold, as an adjustment parameter, can be flexibly configured according to scenario complexity and safety redundancy requirements, achieving a balance between safety and usability.
[0090] In some embodiments, fusing the first identification result and the second identification result to obtain a fused identification result for the target traffic light further includes: when the consistency score is less than a preset threshold, determining the confidence levels of the first identification result and the second identification result respectively to obtain a first confidence level and a second confidence level; and performing a weighted fusion of the first identification result and the second identification result based on the first confidence level and the second confidence level, and taking the state corresponding to the fused result as the fused identification result.
[0091] In this embodiment, the consistency score can first be compared with a preset threshold. When the score is lower than the threshold, it is determined that there is a significant inconsistency or conflict between the current roadside recognition and vehicle-side recognition regarding the target traffic light status. Further, for the first recognition result, the classification probability and detection confidence score output by the roadside recognition model can be directly read as the basic confidence score, and weighted and corrected in conjunction with the working status of the roadside equipment to obtain the first confidence score. Similarly, for the second recognition result, the basic confidence score can be obtained from the probability score output by the vehicle-side recognition network, and the confidence score can be dynamically adjusted in conjunction with factors such as the current imaging quality of the vehicle-mounted camera, the distance and viewing angle between the vehicle and the traffic light, and whether it is obstructed by the vehicle in front, to obtain the second confidence score.
[0092] In this embodiment of the disclosure, the first identification result and the second identification result can be represented as the probability distribution of each candidate state, and then linearly weighted and summed with confidence as the weight to obtain the fused comprehensive probability distribution. Finally, the state with the highest comprehensive probability is selected as the final fused identification result.
[0093] In this way, confidence assessment can dynamically determine which side of the information is more reliable based on the current perception conditions, device status, and model output quality, thereby reducing the probability of erroneous decisions when there is a conflict in the state. Weighted fusion can smoothly integrate the information from both sides. Even when the confidence of one side has only a slight advantage, the information contribution of the other side is still retained, avoiding large jumps in results due to occasional misjudgments on one side, and improving the stability and continuity of the recognition output.
[0094] Figure 3 A schematic diagram illustrating the application of the traffic light recognition method in an embodiment of this disclosure is shown, such as... Figure 3 As shown, the left side shows the traffic lights installed in front of the intersection, the right side has a pole with a camera facing the intersection, and below is a roadside recognition unit deployed on the side of the road. The intersection camera continuously collects roadside images including traffic lights and oncoming lanes, and the roadside recognition unit runs a traffic light position detection and structured information prediction model on it to obtain structured recognition results such as the light group type, light color, and light head direction of each traffic light at the current intersection. Figure 3 As multiple vehicles approach an intersection, the roadside recognition unit transmits the target traffic light status obtained from the roadside recognition to the following autonomous vehicles in real time through the vehicle-road cooperative communication link. This enables the vehicle to perceive traffic lights beyond line of sight or to perform secondary calibration with the recognition results from its own camera, thereby improving the accuracy and robustness of the vehicle's recognition of the traffic lights ahead.
[0095] In some implementations, traffic light detection and recognition can be performed on roadside images captured by existing surveillance cameras at intersections. The recognition results can then be sent to the vehicle via vehicle-to-infrastructure communication for beyond-line-of-sight perception or secondary calibration of the vehicle-side recognition results. This improves the accuracy, robustness, and beyond-line-of-sight perception capabilities of autonomous vehicles in intersection scenarios. Figure 4 This illustration shows another schematic flowchart of the traffic light recognition method according to an embodiment of the present disclosure, such as... Figure 4 As shown, it can include two stages: the first stage is traffic light position detection, which is used to determine the approximate position of each traffic light in the roadside image; the second stage is traffic light information prediction and vehicle-side fusion, which is used to identify the structured information of the traffic lights based on the traffic light area image obtained in the first stage, and fuse it with the traffic light recognition results from the vehicle side.
[0096] In the first stage, the position information of each traffic light in the current image, in the image coordinate system, can be obtained based on the intersection images captured by roadside cameras. For example, the intersection image can first be input into a traffic light position detection model, which can be a real-time object detection network used to quickly and accurately locate all candidate traffic light areas in the entire image. Considering that the physical installation position of traffic lights at the intersection does not change significantly in a short period, the invocation of the position detection model does not need to be real-time at the frame level; it only needs to be triggered periodically at preset time intervals, effectively reducing computational and inference costs while ensuring position stability. Subsequently, the initial position prediction results detected at the current moment can be combined with the cached historical traffic light position result set to enhance the results. For example, for each current predicted bounding box, the Intersection over Union (IoU) can be calculated sequentially with each candidate bounding box in the historical result set. When the IoU of a historical candidate bounding box with the initial position prediction result is higher than a preset threshold, the two are considered as different time observations of the same physical traffic light, and a corresponding similarity weight is assigned to the historical candidate bounding box. If the IoU of the initial position prediction result with all historical candidate bounding boxes is lower than the threshold, it is considered that the camera may have shifted significantly or there is an anomaly in the current detection. In this case, the result can be marked as unreliable and an alert can be triggered. For cases where several IoUs exceed the threshold, these similar historical candidate bounding boxes and the current predicted bounding box can be used as samples, and a more smooth and robust traffic light position prediction result can be obtained by weighting the position coordinates according to similarity. Figure 4 The area indicated by the white box in the middle.
[0097] For example, the process of obtaining the traffic light position prediction result can be represented by the following formula: in, This indicates the predicted position of the traffic lights; This indicates the initial position prediction result; The result set representing the location of historical traffic lights One prediction result; This represents the total number of historical traffic light locations. This represents the similarity weight.
[0098] The similarity weight can be expressed by the following formula: in, This represents the IoU threshold.
[0099] Furthermore, in the second stage, the intersection image can be cropped based on the traffic light location prediction results to obtain the traffic light area image on the roadside, i.e. Figure 4 The image shown only contains traffic lights. Subsequently, the current moment's roadside traffic light area image can be input into the feature extraction module to obtain the feature map of the traffic lights in the current frame. Then, the feature map of the current frame's traffic lights can be fused with the traffic light feature maps cached from historical moments to incorporate temporal information. Next, the fused temporal feature map can be used to obtain the structured information of the traffic lights. For example, the structured information includes at least: light group frame detection results (e.g., horizontal three- or four-unit vehicle lights, vertical three- or four-unit vehicle lights, pedestrian lights, bicycle lights, and independent countdown lights, etc.), and light head detection results (including light head colors such as red, yellow, green, or off status, and light head functional attributes such as left turn, straight ahead, right turn, etc.). After obtaining the complete structured recognition results of the roadside, it can also be fused with the vehicle-side traffic light prediction results to fully utilize the advantages of the roadside fixed viewing angle, long-distance visibility, and temporal enhancement, while combining the vehicle-side's close proximity to the driver's viewpoint, high resolution, and flexible viewing angle to obtain the final traffic light prediction result.
[0100] Figure 5 A schematic diagram of the feature fusion process in an embodiment of this disclosure is shown, such as... Figure 5 As shown, the feature map of the traffic light in the current frame and the traffic light feature map cached in the historical time can be input into the cross-attention layer and the feedforward network layer to fuse them and introduce temporal information.
[0101] It should be understood that Figures 3 to 5 The schematic diagrams shown are merely illustrative and not limiting, and are scalable; those skilled in the art can use them as a basis. Figures 3 to 5 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.
[0102] This disclosure provides a traffic light recognition device, such as... Figure 6 As shown, the device may include: an image acquisition module 601, used to acquire a current frame image containing traffic lights; a location retrieval module 602, used to input the current frame image into a pre-trained traffic light location detection model to obtain an initial location prediction result for at least one traffic light in the current frame image; a history correction module 603, used to match the initial location prediction result with the historical location prediction results of corresponding traffic lights in at least one historical frame, and correct the initial location prediction result according to the matching result to obtain a traffic light location prediction result; and a result generation module 604, used to identify the traffic lights in the current frame image based on the traffic light location prediction result to obtain a traffic light identification result.
[0103] In some embodiments, the image acquisition module 601 includes: a video acquisition submodule for acquiring a continuous video stream or image sequence; and an image selection submodule for selecting a target frame from the video stream or image sequence as the current frame image according to a preset sampling interval.
[0104] In some embodiments, the traffic light location detection model is obtained by: acquiring an image training set containing traffic lights; the image training set containing traffic light location annotation information; and using the image training set to train the model to be trained to obtain the traffic light location detection model.
[0105] In some embodiments, the traffic light recognition device further includes: a new data module 605 ( Figure 6 (Not shown in the image), used to generate new training data based on the matching results and the prediction results of the current frame image and the corresponding traffic light position; incremental training module 606 ( Figure 6 (Not shown in the image) is used to train or fine-tune the traffic light position detection model using newly added training data.
[0106] In some embodiments, the history correction module 603 includes: an overlap calculation submodule, used to calculate the overlap between the initial position prediction result and each historical position prediction result; and a history matching submodule, used to determine historical position prediction results with an overlap greater than or equal to a preset overlap threshold as target historical position prediction results that match the initial position prediction result.
[0107] In some embodiments, the traffic light recognition device further includes: an anomaly warning module 607 ( Figure 6 (Not shown in the image) is used to determine that there is an anomaly in the current frame image and / or traffic lights when the overlap between the initial position prediction result and each historical initial position prediction result is less than a preset overlap threshold, and to generate a warning message.
[0108] In some embodiments, the result generation module 604 includes: a region extraction submodule, used to extract the corresponding traffic light region image from the current frame image based on the traffic light position prediction result; a feature extraction submodule, used to extract features from each traffic light region image to obtain a traffic light feature map set; the traffic light feature map set includes a traffic light feature map corresponding to each traffic light in the current frame image; and an information extraction submodule, used to determine the structured information of each traffic light based on the traffic light feature map set, as the traffic light recognition result.
[0109] In some embodiments, the information extraction submodule is configured to: obtain an enhanced feature map with temporal information based on the traffic light feature map and the historical traffic light feature map; and determine the structured information of each traffic light based on the enhanced feature map.
[0110] In some embodiments, the information extraction submodule is configured to: detect the traffic light group based on the enhanced feature map to obtain the position box of each traffic light group; detect the lamp head of the traffic light within the position box of each traffic light group to obtain the position box of each lamp head; identify the lighting state of the lamp head based on the lamp head position box; and obtain structured information based on the lamp head position box and the lighting state of each lamp head in each traffic light group.
[0111] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0112] The traffic light recognition device in this embodiment realizes a complete traffic light recognition link on the roadside, from video frame acquisition, target position detection, temporal position correction to state recognition. It uses historical position prediction results to match and correct the initial detection results of the current frame, effectively introducing a continuity constraint in the time dimension. This not only alleviates the detection jitter and false detection problems in complex scenarios such as changes in lighting, occlusion, and long-distance imaging, but also significantly improves the stability and accuracy of traffic light position estimation. Furthermore, it can provide a more accurate target area for state recognition, thereby improving the reliability and anti-interference capability of traffic light recognition. This results in an overall improvement in the accuracy and robustness of traffic light recognition in complex traffic environments, providing safer and more reliable traffic light perception results for autonomous vehicles or vehicle-road cooperative systems.
[0113] This disclosure provides a traffic light recognition device, such as... Figure 7 As shown, the device may include: a result acquisition module 701, used to acquire a first recognition result and a second recognition result for the target traffic light; the first recognition result is a roadside recognition result obtained by any of the methods in the embodiments of this disclosure; the second recognition result is a vehicle-side recognition result obtained based on an on-board sensor; and a result fusion module 702, used to fuse the first recognition result and the second recognition result to obtain a fused recognition result for the target traffic light.
[0114] In some embodiments, the traffic light recognition device further includes: a vehicle information submodule for acquiring vehicle location information and attitude information; a lane determination submodule for determining the current driving lane based on map data and according to the location information and attitude information; and a target determination submodule for determining the target traffic light according to the current driving lane.
[0115] In some embodiments, the result fusion module 702 includes: a first fusion submodule, used to read a second recognition result, and when the second recognition result is empty, to use the first recognition result as the fused recognition result.
[0116] In some embodiments, the result fusion module 702 includes: a consistency calculation submodule, used to read the second identification result and calculate the consistency score between the first identification result and the second identification result when the second identification result is not empty; and a second fusion submodule, used to take the common state of the first identification result and the second identification result as the fused identification result when the consistency score is greater than or equal to a preset threshold.
[0117] In some embodiments, the result fusion module 702 includes: a confidence calculation submodule, used to determine the confidence levels of the first identification result and the second identification result respectively when the consistency score is less than a preset threshold, and obtain the first confidence level and the second confidence level; and a third fusion submodule, used to perform weighted fusion of the first identification result and the second identification result according to the first confidence level and the second confidence level, and take the state corresponding to the fusion result as the fusion identification result.
[0118] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0119] The traffic light recognition device in this embodiment enables multi-source perception and intelligent fusion of the same target traffic light on the vehicle side, fully leveraging the advantages of both the roadside and vehicle end, thereby significantly improving the reliability and safety of traffic light recognition under various complex conditions. The fusion mechanism, through consistency checks and conflict resolution of the two types of results, can trigger a more conservative driving strategy or warning when serious inconsistencies are detected, enhancing the safety redundancy and overall robustness of autonomous driving or advanced driver assistance systems in intersection scenarios.
[0120] This disclosure provides a scenario illustration of a traffic light recognition method, such as... Figure 8 As shown.
[0121] As previously described, the traffic light recognition method provided in this disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. This electronic device can be installed in or connected to an autonomous vehicle.
[0122] Specifically, the electronic device may perform the following operations: Acquire the current frame image containing traffic lights; input the current frame image into a pre-trained traffic light position detection model to obtain the initial position prediction result of at least one traffic light in the current frame image; match the initial position prediction result with the historical position prediction result of the corresponding traffic light in at least one historical frame, and correct the initial position prediction result according to the matching result to obtain the traffic light position prediction result; based on the traffic light position prediction result, identify the traffic lights in the current frame image to obtain the traffic light identification result.
[0123] This disclosure provides a scenario illustration of a traffic light recognition method, such as... Figure 9 As shown.
[0124] As previously described, the traffic light recognition method provided in this disclosure is applied to an electronic device. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. This electronic device can be installed in or connected to an autonomous vehicle.
[0125] Specifically, the electronic device may perform the following operations: A first recognition result and a second recognition result for the target traffic light are obtained; the first recognition result is the roadside recognition result obtained by any of the methods in the embodiments of this disclosure; the second recognition result is the vehicle-side recognition result obtained based on the vehicle-mounted sensor; the first recognition result and the second recognition result are fused to obtain a fused recognition result for the target traffic light.
[0126] It should be understood that Figure 8 and Figure 9 The scene diagrams shown are merely illustrative and not restrictive; those skilled in the art can interpret them based on... Figure 8 and Figure 9 Even with various obvious changes and / or substitutions to the examples, the resulting technical solutions still fall within the scope of this disclosure.
[0127] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0128] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, a computer program product, and an autonomous vehicle.
[0129] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0130] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 1002 or a computer program loaded from storage unit 1008 into random access memory (RAM) 1003. The RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.
[0131] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0132] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the traffic light recognition method. For example, in some embodiments, the traffic light recognition method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the traffic light recognition method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform a traffic light identification method by any other suitable means (e.g., by means of firmware).
[0133] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0134] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0135] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0136] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0137] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0138] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0139] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0140] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A traffic light recognition method, comprising: Get the current frame image containing the traffic lights; The current frame image is input into a pre-trained traffic light position detection model to obtain the initial position prediction result of at least one traffic light in the current frame image; The initial position prediction result is matched with the historical position prediction result of the corresponding traffic light in at least one historical frame. The initial position prediction result is then corrected based on the matching result to obtain the traffic light position prediction result. Based on the traffic light position prediction result, the traffic lights in the current frame image are identified to obtain the traffic light identification result.
2. The method according to claim 1, wherein, The step of obtaining the current frame image containing the traffic lights includes: To acquire a continuous video stream or image sequence; The target frame is selected from the video stream or image sequence as the current frame image according to a preset sampling interval.
3. The method according to claim 1, wherein, The traffic light position detection model is obtained through the following method: Obtain an image training set containing traffic lights; the image training set includes traffic light location annotation information. The traffic light position detection model is obtained by training the model to be trained using the image training set.
4. The method according to claim 3, wherein, The method further includes: Based on the matching results, new training data is generated according to the current frame image and the corresponding traffic light position prediction results. The newly added training data is used to train or fine-tune the traffic light position detection model.
5. The method according to claim 1, wherein, The matching of the initial position prediction result with the historical position prediction results of the corresponding traffic lights in at least one historical frame includes: Calculate the overlap between the initial location prediction result and each of the historical location prediction results; The historical location prediction results with an overlap greater than or equal to a preset overlap threshold are determined as the target historical location prediction results that match the initial location prediction results.
6. The method according to claim 5, wherein, The method further includes: When the overlap between the initial position prediction result and each of the historical position prediction results is less than the preset overlap threshold, it is determined that the current frame image and / or traffic light is abnormal, and a warning message is generated.
7. The method according to claim 1, wherein, The step of identifying traffic lights in the current frame image based on the traffic light position prediction result to obtain traffic light identification results includes: Based on the traffic light location prediction result, the corresponding traffic light area image is extracted from the current frame image; Feature extraction is performed on each of the traffic light region images to obtain a traffic light feature map set; the traffic light feature map set includes the traffic light feature map corresponding to each traffic light in the current frame image; The structured information of each traffic light is determined based on the traffic light feature map set, and is used as the traffic light recognition result.
8. The method according to claim 7, wherein, The step of determining the structured information of each traffic light based on the traffic light feature map includes: Based on the traffic light feature map and the historical traffic light feature map, an enhanced feature map with time sequence information is obtained; The structured information of each traffic light is determined based on the enhanced feature map.
9. The method according to claim 8, wherein, The determination of the structured information of each traffic light based on the enhanced feature map includes: Based on the enhanced feature map, the traffic light group is detected to obtain the position box of each traffic light group; Within each of the aforementioned lamp group position frames, the lamp head of the signal light is detected to obtain the lamp head position frame for each lamp head. Based on the lamp head position frame, the lighting status of the lamp head is identified; The structured information is obtained based on the position frame of each lamp head in each signal light group and its lighting status.
10. A traffic light recognition method, comprising: Obtain a first identification result and a second identification result for the target traffic light; the first identification result is a roadside identification result obtained by the method described in any one of claims 1-9; the second identification result is a vehicle-side identification result obtained based on an onboard sensor; The first recognition result and the second recognition result are fused to obtain a fused recognition result for the target traffic light.
11. The method according to claim 10, wherein, The method further includes: Obtain the vehicle's location and attitude information; Based on map data, the current driving lane is determined according to the location information and the attitude information; The target traffic light is determined based on the current driving lane.
12. The method according to claim 10, wherein, The step of fusing the first recognition result and the second recognition result to obtain a fused recognition result for the target traffic light includes: Read the second recognition result. When the second recognition result is empty, use the first recognition result as the fusion recognition result.
13. The method according to claim 10, wherein, The step of fusing the first recognition result and the second recognition result to obtain a fused recognition result for the target traffic light includes: Read the second recognition result; if the second recognition result is not empty, calculate the consistency score between the first recognition result and the second recognition result. When the consistency score is greater than or equal to a preset threshold, the common state of the first identification result and the second identification result is taken as the fusion identification result.
14. The method according to claim 13, wherein, The step of fusing the first recognition result and the second recognition result to obtain a fused recognition result for the target traffic light further includes: When the consistency score is less than a preset threshold, the confidence levels of the first identification result and the second identification result are determined respectively to obtain the first confidence level and the second confidence level; Based on the first confidence level and the second confidence level, the first identification result and the second identification result are weighted and fused, and the state corresponding to the fusion result is taken as the fused identification result.
15. A traffic light recognition device, comprising: The image acquisition module is used to acquire the current frame image containing the traffic lights; The location retrieval module is used to input the current frame image into a pre-trained traffic light location detection model to obtain the initial location prediction result of at least one traffic light in the current frame image; The historical correction module is used to match the initial position prediction result with the historical position prediction result of the corresponding traffic light in at least one historical frame, and correct the initial position prediction result according to the matching result to obtain the traffic light position prediction result. The result generation module is used to identify the traffic lights in the current frame image based on the traffic light position prediction result, and obtain the traffic light identification result.
16. A traffic light recognition device, comprising: The result acquisition module is used to acquire a first recognition result and a second recognition result for the target traffic light; the first recognition result is a roadside recognition result obtained by the method described in any one of claims 1-9; the second recognition result is a vehicle-side recognition result obtained based on an onboard sensor; The result fusion module is used to fuse the first recognition result and the second recognition result to obtain a fused recognition result for the target traffic light.
17. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform the method of any one of claims 1-14.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-14.
19. A computer program product comprising a computer program stored on a storage medium, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-14.
20. An autonomous vehicle, including the electronic equipment as claimed in claim 17.