Traffic light identification method, system and device based on artificial intelligence technology
By constructing a dynamic perception data set and reinforcement learning framework, combined with multi-agent collaborative training strategies, the problem of low recognition accuracy of existing traffic light recognition methods in complex scenarios is solved, and more efficient and reliable traffic light recognition is achieved.
Patent Information
- Application Number
- CN202510207649.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
The existing traffic light recognition method based on artificial intelligence has low recognition accuracy in complex traffic scenarios, making it difficult to adapt to different weather and lighting conditions, and fails to make full use of multi-source data and dynamic scene changes.
Dynamic perception data sets are constructed by collecting RGB image sequences and light sensing data, and state representations are generated using the reinforcement learning framework, traffic light state classification actions and confidence are defined, and multi-agent collaborative training strategy is used to optimize the identification model.
The accuracy and efficiency of traffic light recognition are improved, so that the model can adapt to complex traffic scenarios and variable environmental conditions, and the reliability and interpretability of the identification results are enhanced.
Smart Images

Figure CN120148003A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a traffic light recognition method, system and device based on artificial intelligence technology. Background Art
[0002] With the rapid development of artificial intelligence technology, remarkable progress has been made in the field of intelligent transportation. In scenarios such as autonomous driving and intelligent traffic monitoring, accurately recognizing the traffic light status is crucial. Early traffic light recognition mainly relied on traditional image processing methods, such as color threshold segmentation and shape matching. However, these methods performed poorly in complex environments. In recent years, artificial intelligence-based image recognition technology has been widely applied to traffic light recognition. By using deep learning models such as convolutional neural networks (CNNs), it is able to automatically learn image features and improve the recognition accuracy to a certain extent.
[0003] However, the actual traffic scenario is extremely complex. In different weather conditions (such as heavy rain, thick fog), lighting conditions (direct strong light, backlight, weak light at night) and traffic scenarios (road congestion at intersections, complex multi-lane road conditions), the existing artificial intelligence-based recognition methods still face many challenges and are difficult to adapt to different weather and lighting conditions. Moreover, past research has mostly focused on single image feature extraction and lacks in-depth exploration of multi-source data fusion and dynamic scene modeling. Most methods only rely on image data and ignore other information such as light sensing data. However, the existing technology fails to effectively utilize these multi-source data, restricting the improvement of recognition performance. In addition, the actual traffic scenario is dynamically changing, and the vehicle movement trajectory and environmental lighting conditions change at any time. When constructing the model, the existing methods do not adequately consider the dynamic changes of the traffic scenario and do not fully explore the potential relationship between the vehicle movement trajectory and the traffic light status, resulting in a low recognition accuracy in complex traffic scenarios.
[0004] Therefore, the present invention proposes a traffic light recognition method, system and device based on artificial intelligence technology. Summary of the Invention
[0005] The present invention provides a traffic light recognition method, system and device based on artificial intelligence technology, including: collecting RGB image sequences and light sensing data under different conditions to construct a dynamic perception data set, providing a rich and comprehensive data basis for traffic light recognition. The light sensing data can provide key information such as ambient light intensity, which helps to more accurately recognize traffic lights under different light conditions, enabling the model to adapt to different weather and light conditions and enhancing the generalization ability of the model. Constructing a reinforcement learning framework to generate state representations, taking into account the dynamic changes of the traffic scene and fully exploring the potential relationship between vehicle movement trajectories and traffic light states, which helps to deeply mine the potential features and rules in the data and improve the accuracy of recognition. Defining traffic light state classification actions and confidence levels to make the recognition results more reliable and interpretable. Adopting a multi-agent collaborative training strategy to synchronously optimize the local recognition model and the global policy network, improving the overall performance and adaptability of the model. Obtaining the recognition result based on the scene to be recognized and the model, enabling fast and accurate traffic light recognition, which helps to improve traffic efficiency and safety. This method improves the accuracy and efficiency of traffic light recognition based on artificial intelligence technology and adapts to various complex traffic scenes and environmental conditions.
[0006] The present invention provides a traffic light recognition method based on artificial intelligence technology, including:
[0007] S1: Collect RGB image sequences and light sensing data under different weather, light conditions and different traffic scenes, and construct a corresponding dynamic perception data set including traffic light states, vehicle movement trajectories and ambient light conditions based on each RGB image sequence and light sensing data;
[0008] S2: Construct a reinforcement learning framework, and generate state representations of the Markov decision process for each RGB image sequence based on the dynamic perception data set corresponding to each RGB image sequence in the reinforcement learning framework;
[0009] S3: Define traffic light state classification actions and corresponding confidence levels in each RGB image sequence;
[0010] S4: Adopt a multi-agent collaborative training strategy, and synchronously optimize the local recognition model and the global policy network through a distributed reinforcement learning framework and the traffic light state classification actions and corresponding confidence levels and state representations of the Markov decision process of all RGB image sequences to obtain a traffic light recognition model;
[0011] S5: Obtain the traffic light recognition result based on the RGB image sequence of the traffic scene to be recognized, the light sensing data within the corresponding time period and the traffic light recognition model.
[0012] Preferably, for the traffic light recognition method based on artificial intelligence technology, a corresponding dynamic perception dataset including traffic light states, vehicle movement trajectories, and environmental light conditions is constructed based on each RGB image sequence and light sensing data, including:
[0013] Extract the depth information of each RGB image sequence;
[0014] Based on each RGB image sequence, depth information, and light sensing data, construct a corresponding dynamic perception dataset including traffic light states, vehicle movement trajectories, and environmental light conditions.
[0015] Preferably, for the traffic light recognition method based on artificial intelligence technology, a corresponding dynamic perception dataset including traffic light states, vehicle movement trajectories, and environmental light conditions is constructed based on each RGB image sequence, depth information, and light sensing data, including:
[0016] Analyze the traffic light state dynamic perception dataset based on each RGB image sequence;
[0017] Analyze the vehicle movement trajectory dynamic perception dataset based on each RGB image sequence, depth information, and trajectory tracking algorithm;
[0018] Analyze the environmental light condition dynamic perception dataset based on each RGB image sequence and light sensing data;
[0019] Among them, the dynamic perception dataset includes a traffic light state dynamic perception dataset, a vehicle movement trajectory dynamic perception dataset, and an environmental light condition dynamic perception dataset.
[0020] Preferably, for the traffic light recognition method based on artificial intelligence technology, S2: Construct a reinforcement learning framework, and generate the state representation of the Markov decision process for each RGB image sequence based on the corresponding dynamic perception dataset in the reinforcement learning framework, including:
[0021] Construct a reinforcement learning framework based on the deep Q-network and A3C algorithm;
[0022] In the reinforcement learning framework, use an improved convolutional neural network to perform dynamic optimization on the corresponding dynamic perception dataset of each RGB image sequence to obtain the state representation of the Markov decision process for each RGB image sequence.
[0023] Preferably, for the traffic light recognition method based on artificial intelligence technology, use an improved convolutional neural network to perform dynamic optimization on the corresponding dynamic perception dataset of each RGB image sequence in the reinforcement learning framework to obtain the state representation of the Markov decision process for each RGB image sequence, including:
[0024] In the reinforcement learning framework, an improved convolutional neural network is used to extract multi-scale features of the traffic light area from the dynamic perception dataset corresponding to each RGB image sequence, obtaining a sequence of feature maps; and a spatio-temporal attention mechanism is adopted to dynamically model the illumination changes and occlusions in the corresponding traffic light area, and then combined with a long short-term memory network for time series modeling to obtain a complete dynamic perception dataset of the traffic light states for each RGB image sequence;
[0025] Based on the vehicle motion trajectory dynamic perception dataset, analyze and mark the traffic scene state transition probability within the corresponding prediction period to obtain a complete vehicle motion trajectory dynamic perception dataset;
[0026] Among them, the state representation of the Markov decision process for each RGB image sequence includes the complete dynamic perception dataset of the traffic light states for each RGB image sequence and the complete vehicle motion trajectory dynamic perception dataset.
[0027] Preferably, for the traffic light recognition method based on artificial intelligence technology, a spatio-temporal attention mechanism is used to dynamically model the illumination changes and occlusions in the corresponding traffic light area, and then combined with a long short-term memory network for time series modeling to obtain a complete dynamic perception dataset of the traffic light states for each RGB image sequence, including:
[0028] Calculate the spatial attention weight at each position of each feature map in the feature map sequence of each RGB image sequence:
[0029]
[0030] In the formula, S t,i,j is the spatial attention weight of the pixel at the i-th row and j-th column of the t-th frame feature map in the currently calculated feature map sequence, σ represents the attention weight, ω k is the weight parameter of the k-th channel, F t,i,j,k is the feature value of the k-th channel at the position (i, j) of the t-th frame feature map in the currently calculated feature map sequence, F t,k is the sum of the feature values of the k-th channel at all positions of the t-th frame feature map in the currently calculated feature map sequence;
[0031] Perform spatial attention weighting on each feature map in each feature map sequence based on the spatial attention weight to obtain a spatially attention-enhanced feature map for each feature map in each feature map sequence:
[0032] F s,t,i,j,k =S t,i,j ·F t,i,j,k
[0033] In the formula, F s,t,i,j,kis the eigenvalue of the k-th channel at the position (i, j) in the spatial attention enhanced feature map of the t-th frame in the currently calculated feature map sequence;
[0034] Based on the temporal attention mechanism, dynamically model the illumination changes and occlusions in the corresponding traffic light regions in each feature map sequence to obtain all spatio-temporal attention enhanced feature maps in each feature map sequence.
[0035] Preferably, a traffic light recognition method based on artificial intelligence technology, which dynamically models the illumination changes and occlusions in the corresponding traffic light regions in each feature map sequence based on the temporal attention mechanism to obtain all spatio-temporal attention enhanced feature maps in each feature map sequence, includes:
[0036] Unfold the spatial attention enhanced feature maps of each feature map in each feature map sequence along the time dimension and calculate the temporal attention weights of each feature map:
[0037]
[0038] In the formula, T t is the temporal attention weight of the spatial attention enhanced feature map of the t-th frame in the currently calculated feature map sequence, e is the natural constant and its value is 2.71828, T is the total number of frames in the currently calculated feature map sequence, α m,t is the temporal correlation strength between the spatial attention enhanced feature map of the m-th frame and the spatial attention enhanced feature map of the t-th frame in the currently calculated feature map sequence, F s-flat,m is the spatial attention enhanced feature map of the m-th frame unfolded along the time dimension, α m,n is the temporal correlation strength between the spatial attention enhanced feature map of the m-th frame and the spatial attention enhanced feature map of the n-th frame in the previously calculated feature map sequence;
[0039] Perform temporal attention weighting on each spatial attention enhanced feature map unfolded along the time dimension in each feature map sequence to obtain the corresponding spatio-temporal attention enhanced feature map:
[0040] F st,t = T t ·F s-flat,t
[0041] In the formula, F st,t is the spatio-temporal attention enhanced feature map of the t-th frame in the currently calculated feature map sequence, F s-flat,t is the spatial attention enhanced feature map of the t-th frame unfolded along the time dimension;
[0042] Combine all spatio-temporal attention enhanced feature maps in each feature map sequence with a long short-term memory network for time series modeling to obtain a complete dataset for dynamic perception of the traffic light states of each RGB image sequence.
[0043] Preferably, for the traffic light recognition method based on artificial intelligence technology, the state transition probability of the traffic scene within the corresponding prediction period is analyzed and marked based on the dynamic perception dataset of vehicle movement trajectories, and a complete dynamic perception dataset of vehicle movement trajectories is obtained, including:
[0044] Analyze the velocity vector of each vehicle between adjacent time points based on the dynamic perception dataset of vehicle movement trajectories;
[0045] Judge the state of each vehicle at each time point based on the preset state judgment function and the velocity vector of each vehicle between adjacent time points;
[0046] Based on the states of all vehicles in the dynamic perception dataset of vehicle movement trajectories at each time point, count the state transition probability from any state to another state;
[0047] Take the Nth power of the state transition probability from any state to another state in the dynamic perception dataset of vehicle movement trajectories as the state transition probability of the traffic scene within the corresponding prediction period, where N is the total number of time intervals included in the corresponding prediction period;
[0048] Mark the state transition probability of the traffic scene within the corresponding prediction period in the dynamic perception dataset of vehicle movement trajectories to obtain a complete dynamic perception dataset of vehicle movement trajectories.
[0049] The present invention provides a traffic light recognition system based on artificial intelligence technology for executing any of the above traffic light recognition methods based on artificial intelligence technology, including:
[0050] A dynamic modeling module for collecting RGB image sequences and light sensing data under different weather, light conditions and different traffic scenes, and constructing a corresponding dynamic perception dataset including traffic light states, vehicle movement trajectories and environmental light conditions based on each RGB image sequence and light sensing data;
[0051] A state representation generation module for constructing a reinforcement learning framework and generating the state representation of the Markov decision process of each RGB image sequence based on the dynamic perception dataset corresponding to each RGB image sequence in the reinforcement learning framework;
[0052] A classification action definition module for defining the traffic light state classification actions and corresponding confidence levels in each RGB image sequence;
[0053] An identification model optimization module for adopting a multi-agent collaborative training strategy to synchronously optimize the local identification model and the global policy network through the distributed reinforcement learning framework and the traffic light state classification actions and corresponding confidence levels of all RGB image sequences and the state representation of the Markov decision process to obtain a traffic light recognition model;
[0054] An identification execution module, configured to obtain a traffic light identification result based on a sequence of RGB images of a traffic scene to be identified, illumination sensing data within a corresponding time period, and a traffic light identification model.
[0055] The present invention provides a traffic light identification device based on artificial intelligence technology.
[0056] It includes a processor and a storage device.
[0057] The storage device is used to store instructions.
[0058] When the processor executes the instructions, the above-mentioned traffic light identification method based on artificial intelligence technology is implemented.
[0059] The beneficial effects of the present invention compared with the prior art are as follows: collecting RGB image sequences and illumination sensing data under different conditions to construct a dynamic perception data set, providing a rich and comprehensive data basis for traffic light identification. The illumination sensing data can provide key information such as environmental illumination intensity, which helps to more accurately identify traffic lights under different illumination conditions, enabling the model to adapt to different weather and illumination conditions and enhancing the generalization ability of the model. Constructing a reinforcement learning framework to generate state representations, taking into account the dynamic changes of the traffic scene and fully exploring the potential relationship between vehicle movement trajectories and traffic light states, which helps to deeply mine the potential features and laws in the data and improve the accuracy of identification. Defining traffic light state classification actions and confidence levels makes the identification results more reliable and interpretable. Adopting a multi-agent collaborative training strategy to synchronously optimize the local identification model and the global policy network improves the overall performance and adaptability of the model. Obtaining the identification result based on the scene to be identified and the model can achieve fast and accurate traffic light identification, which helps to improve traffic efficiency and safety. This method improves the accuracy and efficiency of traffic light identification based on artificial intelligence technology and adapts to various complex traffic scenes and environmental conditions.
[0060] Other features and advantages of the present invention will be described in the subsequent specification, and some of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.
[0061] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings
[0062] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the accompanying drawings:
[0063] Figure 1Flowchart of the traffic light recognition method based on artificial intelligence technology in the embodiments of the present invention;
[0064] Figure 2 Schematic diagram of the traffic light recognition system based on artificial intelligence technology in the embodiments of the present invention. Detailed implementation manners
[0065] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0066] Embodiment 1:
[0067] The present invention provides a traffic light recognition method based on artificial intelligence technology. Refer to Figure 1 , including:
[0068] S1: First, use a multi-modal sensor to collect RGB image sequences and light sensing data in different weather conditions (such as sunny, rainy, and foggy days), lighting conditions (strong light, weak light, backlight), and different traffic scenarios in real time, and construct a corresponding dynamic perception dataset containing traffic light states, vehicle movement trajectories, and environmental lighting conditions based on each RGB image sequence and light sensing data. Among them, the vehicle movement trajectory is accurately tracked and obtained by using the Kalman filter algorithm, and the environmental lighting condition is determined by joint analysis of the light sensor and image features;
[0069] S2: Construct a reinforcement learning framework based on the Deep Q-Network (DQN) and the Advantage Actor-Critic (A3C) algorithm, and generate state representations of the Markov decision process for each RGB image sequence based on the dynamic perception dataset corresponding to each RGB image sequence in the reinforcement learning framework;
[0070] S3: Define traffic light state classification actions (including red light, green light, yellow light, unknown) and corresponding confidence levels (dynamically adjust the confidence level threshold according to different scenarios through an adaptive algorithm) in each RGB image sequence; generate a dynamic reward signal by using a real-time traffic flow simulator, and the reward function includes recognition accuracy, response delay, and false detection rate penalty terms, and the penalty is graded according to the false detection type and the severity of the consequences.
[0071] S4: Adopt a multi-agent collaborative training strategy. Through a distributed reinforcement learning framework, the traffic light state classification actions and corresponding confidence levels of all RGB image sequences, as well as the state representation of the Markov decision process, synchronously optimize the local recognition model and the global policy network to obtain a traffic light recognition model. Among them, the multi-agent collaborative training strategy adopts a heterogeneous network architecture. The vision-dominated agent focuses on extracting the morphological features of traffic lights under complex lighting conditions by using methods such as histogram equalization and adaptive threshold segmentation. The time-series-dominated agent captures the long-term relationships of time series through a gated recurrent unit (GRU) and focuses on modeling the state persistence during traffic flow mutations. Each agent explores based on an independent copy of the environment, realizes knowledge transfer by sharing policy gradient parameters, introduces an imitation learning model, pre-trains the policy network with a large amount of expert standard data, and uses an adversarial training mechanism to minimize the difference between the results generated by the policy network and the expert annotation results, thereby accelerating convergence;
[0072] S5: Obtain the traffic light recognition result based on the RGB image sequence of the traffic scene to be recognized, the light sensing data within the corresponding time period, and the traffic light recognition model;
[0073] When deploying the lightweight model, use the hierarchical knowledge distillation technology to compress the global policy network into a real-time inference model that can be allowed by edge computing devices. This model can be cascaded and verified with traditional recognition methods (generating candidate regions based on RGB or HSV threshold segmentation). The specific method is as follows:
[0074] When the confidence level of the reinforcement learning model is lower than the preset threshold, trigger an auxiliary judgment process based on morphological filtering using OpenCV and threshold segmentation in the HSV color space;
[0075] If the two-level judgment results conflict, start a historical state backtracking mechanism based on a long short-term memory network (LSTM), and make a final determination by combining historical recognition results and state transition probabilities.
[0076] In this embodiment, the RGB image sequence in the traffic scene refers to a series of continuous RGB-format images collected in the traffic environment, which contains information such as different weather conditions, lighting, and traffic conditions.
[0077] In this embodiment, the light sensing data is data related to the environmental light intensity, direction, etc. obtained through a specific light sensor.
[0078] In this embodiment, the dynamic perception dataset containing traffic light states, vehicle movement trajectories, and environmental light conditions is a dataset constructed by integrating information such as the display of traffic lights, the paths and speeds of vehicle driving, and the lighting conditions of the environment at that time, and can reflect the dynamic changes of these factors over time.
[0079] In this embodiment, the reinforcement learning framework is a machine learning framework that learns the optimal policy through the interaction between the agent and the environment and based on the reward feedback.
[0080] In this embodiment, the state representation of the Markov decision process of the RGB image sequence is a representation form that transforms the RGB image sequence in the reinforcement learning framework into a form that can describe its state. This representation takes into account the association and probability distribution of the front and back states.
[0081] In this embodiment, defining the traffic light state classification actions and corresponding confidence levels in each RGB image sequence includes: clearly indicating the category of the traffic light state (such as red light, green light, yellow light) in each RGB image sequence, and giving an estimate of the credibility of this judgment.
[0082] In this embodiment, obtaining the traffic light recognition result based on the RGB image sequence of the traffic scene to be recognized, the light sensing data within the corresponding time period, and the traffic light recognition model means: using the RGB image sequence of the traffic scene that has not been judged and the corresponding light sensing data, inputting them into the established traffic light recognition model, so as to obtain the judgment result on the traffic light state.
[0083] In this embodiment, the traffic light recognition result refers to the judgment conclusion on the traffic light state finally obtained through specific methods and models after processing and analyzing the given RGB image sequence of the traffic scene and the light sensing data within the corresponding time period. For example, clearly indicating whether the traffic light at this time is a red light, a green light, or a yellow light.
[0084] In this embodiment, when the vision-dominated agent extracts the morphological features of traffic lights under complex lighting conditions, it uses the adaptive histogram equalization algorithm to enhance the contrast of the image, and combines the adaptive threshold segmentation algorithm to accurately extract the contour features of the traffic light area.
[0085] In this embodiment, the time-series-dominated agent uses the gated recurrent unit to model the state persistence during traffic flow mutations. By adjusting the number of hidden layer nodes and the time step of the GRU, it optimizes the processing ability of time series data and accurately captures the changing trend of traffic states.
[0086] In this embodiment, the hierarchical knowledge distillation technology transfers the knowledge of the teacher model layer by layer to the student model. While retaining the model performance, it compresses the model parameter quantity to 15% of the original version, enabling the model to achieve a real-time processing performance of 83fps on edge computing devices (such as the Jetson Nano platform).
[0087] In this embodiment, the traffic flow simulator adopts a microscopic traffic flow model to simulate traffic scenarios under different traffic flows, vehicle speeds, and vehicle densities, generating more realistic and dynamic reward signals to improve the training results of the reinforcement learning model.
[0088] The beneficial effects of the above technologies are as follows: Collect RGB image sequences and light sensing data under different conditions to construct a dynamic perception dataset, providing a rich and comprehensive data basis for traffic light recognition. The light sensing data can provide key information such as ambient light intensity, which helps to more accurately recognize traffic lights under different lighting conditions, enabling the model to adapt to different weather and lighting conditions and enhancing the generalization ability of the model. Construct a reinforcement learning framework to generate state representations, taking into account the dynamic changes in traffic scenarios and fully exploring the potential relationship between vehicle movement trajectories and traffic light states, which helps to deeply mine the potential features and laws in the data and improve the accuracy of recognition. Define traffic light state classification actions and confidence levels to make the recognition results more reliable and interpretable. Adopt a multi-agent collaborative training strategy to synchronously optimize the local recognition model and the global policy network, improving the overall performance and adaptability of the model. Obtain recognition results based on the scene to be recognized and the model, enabling fast and accurate traffic light recognition, which helps to improve traffic efficiency and safety. This method improves the accuracy and efficiency of traffic light recognition based on artificial intelligence technology and adapts to various complex traffic scenarios and environmental conditions.
[0089] Embodiment 2:
[0090] Based on the traffic light recognition method based on artificial intelligence technology in Embodiment 1, construct a corresponding dynamic perception dataset including traffic light states, vehicle movement trajectories, and environmental lighting conditions based on each RGB image sequence and light sensing data, including:
[0091] Extract the depth information of each RGB image sequence;
[0092] Construct a corresponding dynamic perception dataset including traffic light states, vehicle movement trajectories, and environmental lighting conditions based on each RGB image sequence, depth information, and light sensing data.
[0093] In this embodiment, extracting the depth information of each RGB image sequence means obtaining relevant information that can reflect the distance between the objects in the image and the camera from each collected RGB image sequence, which helps to more comprehensively understand the scene and object relationships in the image.
[0094] The beneficial effects of the above technical solutions are as follows: Extracting the depth information of the RGB image sequence adds an important dimension to the construction of the dynamic perception dataset and enriches the data features. Combining the RGB image sequence, depth information, and light sensing data to construct the dataset can more comprehensively reflect the real situation of the traffic scene. Making the dynamic perception dataset contain more details and key information helps to improve the accuracy and reliability of traffic light recognition. It provides a richer and more valuable data basis for subsequent recognition and analysis. Enhancing the integrity and effectiveness of the data improves the performance and adaptability of the traffic light recognition method.
[0095] Embodiment 3:
[0096] Based on the traffic light recognition method using artificial intelligence technology on the basis of Embodiment 2, a corresponding dynamic perception dataset including traffic light states, vehicle movement trajectories, and environmental light conditions is constructed based on each RGB image sequence, depth information, and light sensing data, including:
[0097] Analyzing the dynamic perception dataset of traffic light states based on each RGB image sequence;
[0098] Analyzing the dynamic perception dataset of vehicle movement trajectories based on each RGB image sequence, depth information, and a trajectory tracking algorithm;
[0099] Analyzing the dynamic perception dataset of environmental light conditions based on each RGB image sequence and light sensing data;
[0100] Among them, the dynamic perception dataset includes the dynamic perception dataset of traffic light states, the dynamic perception dataset of vehicle movement trajectories, and the dynamic perception dataset of environmental light conditions.
[0101] In this embodiment, analyzing the dynamic perception dataset of traffic light states based on each RGB image sequence means determining the display states of traffic lights only based on each RGB image sequence and organizing this state information into a dataset that can reflect its dynamic changes.
[0102] In this embodiment, analyzing the dynamic perception dataset of vehicle movement trajectories based on each RGB image sequence, depth information, and a trajectory tracking algorithm means combining each RGB image sequence, the depth information therein, and then using the trajectory tracking algorithm to infer the movement paths and trajectories of vehicles and summarizing this information into a dataset that can reflect its dynamic changes.
[0103] In this embodiment, analyzing the dynamic perception dataset of environmental light conditions based on each RGB image sequence and light sensing data means determining the environmental light conditions at that time according to each RGB image sequence and the corresponding light sensing data and then organizing them into a dataset that can reflect the dynamic changes of light conditions.
[0104] In this embodiment, the dynamic perception data set of traffic light states refers to a set that contains relevant information about the changes in traffic light states over time, and this information can reflect the dynamic change characteristics of traffic light states. For example, the traffic light state data collected at different time periods at an intersection, including the duration of red lights, the duration of green lights, the moments when yellow lights appear, etc., form a data set that can show the change rules of traffic light states at different times of each day.
[0105] In this embodiment, the dynamic perception data set of vehicle movement trajectories refers to a set of data that integrates relevant information about vehicle movement trajectories at different times and in different scenarios, and it can reflect the dynamic changes in vehicle movement trajectories. For example, the driving trajectories of different vehicles within a month are recorded, including information such as the starting points of the vehicles, driving directions, turning points, speed changes, etc., and are summarized into a data set that can reflect the dynamic changes in vehicle movement trajectories.
[0106] In this embodiment, the dynamic perception data set of environmental light conditions is a set of data that contains specific information and change information about environmental light at different times and in different scenarios, and it can reflect the dynamic change characteristics of environmental light conditions. For example, data such as light intensity, light angle, and shadow distribution are used to construct a data set that can reflect the dynamic changes in environmental light conditions in this scenario over a period of time.
[0107] The beneficial effects of the above technical solutions are as follows: By separately analyzing and constructing the dynamic perception data sets of traffic light states, vehicle movement trajectories, and environmental light conditions, the data classification becomes clearer and more targeted. Analyzing the dynamic perception data set of traffic light states for RGB image sequences can accurately capture the changes in traffic lights. Combining RGB image sequences, depth information, and trajectory tracking algorithms to obtain the dynamic perception data set of vehicle movement trajectories can comprehensively reflect the movement of vehicles. Obtaining the dynamic perception data set of environmental light conditions through RGB image sequences and light sensing data provides a basis for subsequent consideration of the influence of light. The constructed multiple dynamic perception data sets provide comprehensive and detailed data support for the accurate recognition of traffic lights, improving the accuracy and reliability of recognition.
[0108] Embodiment 4:
[0109] Based on the traffic light recognition method using artificial intelligence technology in Embodiment 1, S2: Construct a reinforcement learning framework, and based on the dynamic perception data set corresponding to each RGB image sequence in the reinforcement learning framework, generate the state representation of the Markov decision process for each RGB image sequence, including:
[0110] Construct a reinforcement learning framework based on the deep Q-network and the A3C algorithm;
[0111] In the reinforcement learning framework, an improved convolutional neural network (CNN) is used to perform dynamic optimization on the dynamic perception dataset corresponding to each RGB image sequence, and the state representation of the Markov decision process of each RGB image sequence is obtained. Among them, the CNN includes a depthwise separable convolutional layer to reduce the computational amount, and a spatio-temporal attention mechanism is used to dynamically model the illumination changes and occlusions in the corresponding traffic light area. Then, combined with a long short-term memory network (LSTM) for time series modeling, a complete dataset for dynamic perception of the traffic light state of each RGB image sequence is obtained.
[0112] In this embodiment, constructing a reinforcement learning framework based on the deep Q-network and the A3C algorithm means combining the deep Q-network and the A3C algorithm to build a structure and system for reinforcement learning to achieve effective interaction and learning between the agent and the environment.
[0113] In this embodiment, the improved convolutional neural network uses a depthwise separable convolutional layer to replace the traditional convolutional layer to reduce the number of model parameters and the computational amount. At the same time, a batch normalization layer is introduced into the network to accelerate the network convergence and improve the stability and generalization ability of the model.
[0114] The beneficial effects of the above technical solutions are as follows: Constructing a reinforcement learning framework based on the deep Q-network and the A3C algorithm integrates the advantages of the two algorithms and improves the learning efficiency and performance. Using the improved convolutional neural network for dynamic optimization can better extract image features and enhance the accuracy of state representation. Processing the dynamic perception dataset in the reinforcement learning framework makes full use of the spatio-temporal information in the data and improves the understanding of traffic scenarios. The generated state representation of the Markov decision process provides an effective feature representation for subsequent decision-making and recognition. It improves the intelligence and adaptability of the traffic light recognition method and can cope with complex and changeable traffic conditions.
[0115] Embodiment 5:
[0116] Based on the traffic light recognition method of artificial intelligence technology in Embodiment 4, in the reinforcement learning framework, an improved convolutional neural network is used to perform dynamic optimization on the dynamic perception dataset corresponding to each RGB image sequence, and the state representation of the Markov decision process of each RGB image sequence is obtained, including:
[0117] In the reinforcement learning framework, an improved convolutional neural network is used to extract multi-scale features of the traffic light area from the dynamic perception dataset corresponding to each RGB image sequence, and a sequence of feature maps is obtained;
[0118] Based on the vehicle motion trajectory dynamic perception dataset, the traffic scene state transition probability in the corresponding prediction period is analyzed and marked to obtain a complete vehicle motion trajectory dynamic perception dataset;
[0119] Among them, the state representation of the Markov decision process of each RGB image sequence includes the complete dynamic perception dataset of traffic light states and the complete dynamic perception dataset of vehicle motion trajectories for each RGB image sequence.
[0120] In this embodiment, an improved convolutional neural network is used in the reinforcement learning framework to extract multi-scale features of the traffic light area from the dynamic perception dataset corresponding to each RGB image sequence, obtaining a sequence of feature maps. In the set reinforcement learning architecture, an optimized convolutional neural network is used to obtain the characteristics of the traffic light area at different levels and scales from the dynamic perception dataset related to each RGB image sequence, thereby obtaining a series of feature maps.
[0121] In this embodiment, the prediction period refers to a preset time interval for prediction or estimation, such as the next 5 minutes, 10 minutes, etc.
[0122] In this embodiment, the traffic scene state transition probability within the prediction period is the likelihood of the traffic scene changing from one state to another within the set prediction time period. For example, the probability of the traffic changing from a smooth state to a congested state within the next 3 minutes.
[0123] The beneficial effects of the above technical solutions are as follows: By using an improved convolutional neural network to extract multi-scale features of the traffic light area and through spatio-temporal attention mechanism and long short-term memory network modeling, it can more accurately capture the dynamic changes of traffic light states and improve the recognition accuracy. Obtaining the complete dynamic perception dataset of traffic light states fully considers complex situations such as light changes and occlusions, enhancing the robustness of the model. Analyzing the dynamic perception dataset of vehicle motion trajectories to obtain and mark the state transition probability provides an important basis for understanding the traffic scene. The complete dynamic perception dataset of vehicle motion trajectories enriches the state representation of the Markov decision process, making the decision-making more comprehensive and scientific. It improves the construction of the state representation and enhances the adaptability and reliability of the traffic light recognition method in complex traffic scenarios.
[0124] Embodiment 6:
[0125] Based on the traffic light recognition method of artificial intelligence technology on the basis of Embodiment 5, a spatio-temporal attention mechanism is used to dynamically model the light changes and occlusions of the corresponding traffic light area, and then combined with a long short-term memory network for time series modeling to obtain the complete dynamic perception dataset of traffic light states for each RGB image sequence, including:
[0126] Calculate the spatial attention weights at each position of each feature map in the feature map sequence of each RGB image sequence:
[0127]
[0128] Where S t,i,j is the spatial attention weight of the pixel at the i-th row and j-th column in the t-th frame feature map in the currently calculated feature map sequence, σ represents the attention weight (σ is the Sigmoid function, which can map the output value to the range [0,1], and this weight represents the importance degree of this position), ω k is the weight parameter of the k-th channel, which makes the output image highlight the features that are more sensitive to illumination changes and occlusions, F t,i,j,k is the feature value of the k-th channel at the position (i,j) in the t-th frame feature map in the currently calculated feature map sequence, F t,k is the sum of the feature values of the k-th channel at all positions in the t-th frame feature map in the currently calculated feature map sequence;
[0129] Assume that in a simplified scenario, the feature map has 3 channels (k = 1, 2, 3), corresponding to color, brightness, and texture information respectively. If the illumination change has a greater impact on the color information, and the occlusion is more sensitive to the texture information, we may set ω 1 = 0.6, ω 2 = 0.2, ω 3 = 0.2. In this way, when calculating the spatial attention weight, the feature values of the color channel will be given a greater weight, so as to highlight the features related to illumination changes. For example, under direct strong light, the color channel may contain more information about the color distortion of traffic lights. Through the larger ω 1 weight, the model can pay more attention to these key information;
[0130] Perform spatial attention weighting on each feature map in each feature map sequence based on the spatial attention weight to obtain the spatially attention-enhanced feature map of each feature map in each feature map sequence:
[0131] F s,t,i,j,k = S t,i,j ·F t,i,j,k
[0132] Where F s,t,i,j,k is the feature value of the k-th channel at the position (i,j) in the t-th frame spatially attention-enhanced feature map in the currently calculated feature map sequence;
[0133] Dynamically model the illumination changes and occlusions corresponding to the traffic light area in each feature map sequence based on the temporal attention mechanism to obtain all spatio-temporal attention-enhanced feature maps in each feature map sequence.
[0134] The beneficial effects of the above technical solutions are as follows: By calculating the spatial attention weights, the features at important positions in the feature map can be highlighted, and the attention to key regions can be improved. Weighting based on the spatial attention weights to obtain a spatially attention-enhanced feature map enhances the feature expression ability. Using the temporal attention mechanism to dynamically model the illumination changes and occlusions in the traffic light area can better handle complex environmental changes. Obtaining a spatio-temporal attention-enhanced feature map provides a more effective feature representation for accurately identifying the traffic light state. It improves the model's ability to capture and process features in the traffic light area, enhancing the accuracy and robustness of traffic light recognition.
[0135] Example 7:
[0136] Based on the traffic light recognition method based on artificial intelligence technology in Example 6, the illumination changes and occlusions in the traffic light area corresponding to each feature map sequence are dynamically modeled based on the temporal attention mechanism to obtain all spatio-temporal attention-enhanced feature maps in each feature map sequence, including:
[0137] Unfolding the spatially attention-enhanced feature map of each feature map in each feature map sequence along the time dimension and then calculating the temporal attention weight of each feature map:
[0138]
[0139] where T t is the temporal attention weight of the spatially attention-enhanced feature map of the t-th frame in the currently calculated feature map sequence, e is the natural constant and its value is 2.71828, T is the total number of frames in the currently calculated feature map sequence, α m,t is the temporal correlation strength between the spatially attention-enhanced feature map of the m-th frame and the spatially attention-enhanced feature map of the t-th frame in the currently calculated feature map sequence, F s-flat,m is the spatially attention-enhanced feature map of the m-th frame unfolded along the time dimension, α m,n is the temporal correlation strength between the spatially attention-enhanced feature map of the m-th frame and the spatially attention-enhanced feature map of the n-th frame in the previously calculated feature map sequence;
[0140] Consider an image sequence with 5 time frames (T = 5). Suppose the current time frame t = 3, and we want the model to pay more attention to the correlation between the previous frame (m = 3) and the current frame (m = 2) because illumination changes and occlusions usually have strong continuity between adjacent frames. Then we can set α 2,3 = 0.5, α 3,3 = 0.3, and for other frames, such as α 1,3 = 0.1, α 4,3 = 0.05, α 5,3 = 0.05. In this way, when calculating the temporal attention weight T 3When the second and third frames contribute more to the feature pair T 3 it helps to capture the dynamic changes in time;
[0141] Perform temporal attention weighting on each spatially attention-enhanced feature map expanded along the temporal dimension in each feature map sequence to obtain the corresponding spatio-temporal attention-enhanced feature map:
[0142] F st,t = T t ·F s-flat,t
[0143] where F st,t is the spatio-temporal attention-enhanced feature map of the t-th frame in the currently calculated feature map sequence, and F s-flat,t is the spatially attention-enhanced feature map of the t-th frame expanded along the temporal dimension;
[0144] Combine all the spatio-temporal attention-enhanced feature maps in each feature map sequence with a long short-term memory network for time series modeling to obtain a complete dataset for dynamic perception of the traffic light states of each RGB image sequence.
[0145] In this embodiment, combining all the spatio-temporal attention-enhanced feature maps in each feature map sequence with a long short-term memory network for time series modeling to obtain a complete dataset for dynamic perception of the traffic light states of each RGB image sequence means: combining all the feature maps that have undergone spatio-temporal attention enhancement in each feature map sequence with a long short-term memory network to construct a time series model, thereby obtaining a comprehensive and complete dynamic perception dataset of the traffic light states in each RGB image sequence. For example, after extracting and enhancing features from a series of consecutive RGB images, use a long short-term memory network to capture the changing patterns and correlations of these features over time, and finally form a dataset that can fully reflect the dynamic changes of the traffic light states over time.
[0146] The beneficial effects of the above technical solutions are as follows: By calculating the temporal attention weights, the importance and correlation strength of the feature maps in the temporal dimension can be captured. Performing temporal attention weighting on the spatially attention-enhanced feature maps to obtain spatio-temporal attention-enhanced feature maps takes into account the temporal factors more comprehensively. Combining with a long short-term memory network for time series modeling makes full use of the time series features of the data, improving the ability to understand and predict the dynamic changes of the traffic light states. Obtaining a complete dataset for dynamic perception of the traffic light states provides richer and more accurate information for accurately identifying the traffic light states. It enhances the model's ability to handle dynamic situations such as illumination changes and occlusions in the traffic light area, improving the accuracy and reliability of traffic light recognition.
[0147] Embodiment 8:
[0148] Based on Example 5, for the traffic light recognition method based on artificial intelligence technology, the state transition probability of the traffic scene within the corresponding prediction period is analyzed and marked based on the dynamic perception dataset of vehicle movement trajectories, and a complete dynamic perception dataset of vehicle movement trajectories is obtained, including:
[0149] Analyze the velocity vector of each vehicle between adjacent time points based on the dynamic perception dataset of vehicle movement trajectories;
[0150] Judge the state of each vehicle at each time point based on the preset state judgment function and the velocity vector of each vehicle between adjacent time points;
[0151] Based on the states of all vehicles in the dynamic perception dataset of vehicle movement trajectories at each time point, count the state transition probability from any state to another state;
[0152] Take the Nth power of the state transition probability from any state to another state in the dynamic perception dataset of vehicle movement trajectories as the state transition probability of the traffic scene within the corresponding prediction period, where N is the total number of time intervals included in the corresponding prediction period;
[0153] Mark the state transition probability of the traffic scene within the corresponding prediction period in the dynamic perception dataset of vehicle movement trajectories to obtain a complete dynamic perception dataset of vehicle movement trajectories.
[0154] In this embodiment, the velocity vector of each vehicle between adjacent time points is analyzed based on the dynamic perception dataset of vehicle movement trajectories. From the dynamic perception dataset of vehicle movement trajectories, the magnitude and direction of the velocity corresponding to each vehicle at adjacent front and rear time points are calculated to form a velocity vector, that is, the vector obtained by dividing the displacement vector of the vehicle between adjacent time points by the time interval between the adjacent time points.
[0155] In this embodiment, the preset state judgment function is a mathematical function set in advance for judging the vehicle state according to certain conditions or rules. For example: when the modulus of the velocity vector is less than the first velocity threshold, it is determined that the state of the vehicle is S1; when the modulus of the velocity vector is greater than the first velocity threshold and less than the second velocity threshold, it is determined that the state of the vehicle is S2; when the modulus of the velocity vector is greater than or equal to the second velocity threshold and less than the third velocity threshold, it is determined that the state of the vehicle is S3, and so on.
[0156] In this embodiment, judging the state of each vehicle at each time point based on the preset state judgment function and the velocity vector of each vehicle between adjacent time points uses the preset state judgment function and combines the velocity vector of each vehicle at adjacent time points to determine the state of each vehicle at each specific time point, such as driving, stopping, accelerating, decelerating, etc.
[0157] In this embodiment, based on the states of all vehicles in the dynamic perception dataset of vehicle movement trajectories at each time point, the state transition probability from any state S j to another state S k is statistically calculated as follows: According to the state conditions of all vehicles in the dynamic perception dataset of vehicle movement trajectories at each time point, the probability of a vehicle transitioning from one state to another is calculated. The statistical process is as follows: The number of times n j of transitioning from state S k to state S jk is divided by the total number of times n j of transitioning out of state S j , and the quotient is regarded as the state transition probability from state S j to state S k .
[0158] The beneficial effects of the above technical solutions are as follows: Analyzing the dynamic perception dataset of vehicle movement trajectories to obtain the velocity vector provides basic data for judging the vehicle state. By presetting the state judgment function to determine the vehicle state at each time point, accurate classification of the vehicle state is achieved. Statistical state transition probabilities can reflect the change rules of vehicle states in traffic scenarios. Taking the Nth power of the state transition probability as the probability within the prediction period takes into account the influence of the time interval and makes the prediction more accurate. Marking the state transition probabilities to obtain a complete dataset provides more comprehensive information for subsequent analysis and decision-making. The dynamic perception dataset of vehicle movement trajectories is improved, enhancing the understanding and prediction ability of traffic scenarios and contributing to more accurate traffic light recognition.
[0159] Embodiment 9:
[0160] The present invention provides a traffic light recognition system based on artificial intelligence technology for implementing any one of the traffic light recognition methods based on artificial intelligence technology in Embodiments 1 to 8, with reference to Figure 2 , and includes:
[0161] A dynamic modeling module for collecting RGB image sequences and light sensing data under different weather, lighting conditions, and different traffic scenarios, and constructing a corresponding dynamic perception dataset including traffic light states, vehicle movement trajectories, and environmental lighting conditions based on each RGB image sequence and light sensing data;
[0162] A state representation generation module for constructing a reinforcement learning framework and generating state representations of the Markov decision process for each RGB image sequence based on the dynamic perception dataset corresponding to each RGB image sequence in the reinforcement learning framework;
[0163] A classification action definition module for defining traffic light state classification actions and corresponding confidence levels in each RGB image sequence;
[0164] An identification model optimization module, which is used to adopt a multi-agent collaborative training strategy, and synchronously optimize the local identification model and the global policy network through a distributed reinforcement learning framework, the traffic light state classification actions and corresponding confidence levels of all RGB image sequences, and the state representation of the Markov decision process, so as to obtain a traffic light identification model;
[0165] An identification execution module, which is used to obtain a traffic light identification result based on the RGB image sequence of the traffic scene to be identified, the light sensing data within the corresponding time period, and the traffic light identification model.
[0166] The beneficial effects of the above technologies are as follows: RGB image sequences and light sensing data under different conditions are collected to construct a dynamic perception data set, which provides a rich and comprehensive data basis for traffic light identification. The light sensing data can provide key information such as environmental light intensity, which helps to more accurately identify traffic lights under different light conditions, enabling the model to adapt to different weather and light conditions and enhancing the generalization ability of the model. A reinforcement learning framework is constructed to generate state representations, taking into account the dynamic changes of the traffic scene and fully exploring the potential relationship between vehicle movement trajectories and traffic light states, which helps to deeply mine the potential features and laws in the data and improve the accuracy of identification. Defining traffic light state classification actions and confidence levels makes the identification results more reliable and interpretable. Adopting a multi-agent collaborative training strategy to synchronously optimize the local identification model and the global policy network improves the overall performance and adaptability of the model. Obtaining the identification result based on the scene to be identified and the model can achieve fast and accurate traffic light identification, which helps to improve traffic efficiency and safety. This method improves the accuracy and efficiency of traffic light identification and adapts to various complex traffic scenes and environmental conditions.
[0167] Embodiment 10:
[0168] The present invention provides a traffic light identification device based on artificial intelligence technology,
[0169] including a processor and a storage device;
[0170] The storage device is used to store instructions;
[0171] When the processor executes the instructions, it implements the traffic light identification method based on artificial intelligence technology as described in any one of claims 1 to 8.
[0172] The beneficial effects of the above technical solutions are as follows: By storing instructions in a storage device, the stability and repeatability of the method are ensured. The processor executes the instructions to implement the traffic light recognition method, achieving automated and intelligent processing. The proposed traffic light recognition method based on artificial intelligence technology can be accurately applied to ensure the accuracy and reliability of recognition. This device has high efficiency and convenience and can quickly process a large amount of data. It provides an integrated solution for traffic light recognition, helping to improve the efficiency and safety of traffic management.
[0173] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.
Claims
1. A traffic light recognition method based on artificial intelligence technology, characterized in that: include: S1: Collect RGB image sequences and light sensing data under different weather conditions, lighting conditions and different traffic scenes, and build a corresponding dynamic perception dataset containing traffic light status, vehicle motion trajectory and ambient lighting conditions based on each RGB image sequence and light sensing data; S2: construct a reinforcement learning framework, and generate a state representation of the Markov decision process of each RGB image sequence based on the dynamic perception dataset corresponding to each RGB image sequence in the reinforcement learning framework; S3: define the traffic light state classification action and corresponding confidence in each RGB image sequence; S4: Adopting a multi-agent collaborative training strategy, the distributed reinforcement learning framework and the traffic light state classification actions and corresponding confidences of all RGB image sequences and the state representation of the Markov decision process are used to simultaneously optimize the local recognition model and the global policy network to obtain the traffic light recognition model; S5: Obtain a traffic light recognition result based on the RGB image sequence of the traffic scene to be recognized, the light sensor data in the corresponding time period, and the traffic light recognition model.
2. The traffic light recognition method based on artificial intelligence technology according to claim 1 is characterized in that: Based on each RGB image sequence and light sensing data, a corresponding dynamic perception dataset containing traffic light status, vehicle motion trajectory and ambient light conditions is constructed, including: Extract the depth information of each RGB image sequence; Based on each RGB image sequence and depth information as well as light sensing data, a corresponding dynamic perception dataset including traffic light status, vehicle motion trajectory and ambient lighting conditions is constructed.
3. The traffic light recognition method based on artificial intelligence technology according to claim 2 is characterized in that: Based on each RGB image sequence and depth information as well as light sensing data, a corresponding dynamic perception dataset containing traffic light status, vehicle motion trajectory and ambient lighting conditions is constructed, including: A traffic light status dynamic perception dataset is generated based on each RGB image sequence; The vehicle motion trajectory dynamic perception dataset is analyzed based on each RGB image sequence and depth information as well as the trajectory tracking algorithm; A dynamic perception dataset of ambient lighting conditions is generated based on each RGB image sequence and lighting sensor data; Among them, the dynamic perception dataset includes the traffic light status dynamic perception dataset, the vehicle motion trajectory dynamic perception dataset, and the ambient lighting condition dynamic perception dataset.
4. The traffic light recognition method based on artificial intelligence technology according to claim 1 is characterized in that: S2: Construct a reinforcement learning framework, and generate a state representation of the Markov decision process of each RGB image sequence based on the dynamic perception dataset corresponding to each RGB image sequence in the reinforcement learning framework, including: Build a reinforcement learning framework based on deep Q network and A3C algorithm; In the reinforcement learning framework, an improved convolutional neural network is used to dynamically optimize the dynamic perception dataset corresponding to each RGB image sequence to obtain the state representation of the Markov decision process of each RGB image sequence.
5. The traffic light recognition method based on artificial intelligence technology according to claim 4 is characterized in that: In the reinforcement learning framework, an improved convolutional neural network is used to dynamically optimize the dynamic perception dataset corresponding to each RGB image sequence to obtain the state representation of the Markov decision process of each RGB image sequence, including: In the reinforcement learning framework, an improved convolutional neural network is used to extract multi-scale features of the traffic light area in the dynamic perception dataset corresponding to each RGB image sequence to obtain a feature map sequence. The spatiotemporal attention mechanism is used to dynamically model the illumination changes and occlusions in the corresponding traffic light area, and then the long short-term memory network is combined to perform time series modeling to obtain a complete dataset of dynamic perception of the traffic light status of each RGB image sequence. Based on the vehicle motion trajectory dynamic perception data set, the traffic scene state transition probability within the corresponding prediction period is analyzed and marked to obtain a complete vehicle motion trajectory dynamic perception data set; Among them, the state representation of the Markov decision process of each RGB image sequence includes a complete data set of dynamic perception of traffic light states and a complete data set of dynamic perception of vehicle motion trajectories for each RGB image sequence.
6. The traffic light recognition method based on artificial intelligence technology according to claim 5 is characterized in that: The spatiotemporal attention mechanism is used to dynamically model the illumination changes and occlusions in the corresponding traffic light area, and then combined with the long short-term memory network for time series modeling, to obtain a complete data set of dynamic perception of the traffic light state for each RGB image sequence, including: Calculate the spatial attention weight at each position of each feature map in the feature map sequence of each RGB image sequence: In the formula, S t,i,j is the spatial attention weight of the pixel in the i-th row and j-th column of the feature map of the t-th frame in the currently calculated feature map sequence, σ represents the attention weight, ω k is the weight parameter of the kth channel, F t,i,j,k is the feature value of the kth channel at position (i, j) in the feature map of the tth frame in the currently calculated feature map sequence, F t,k is the sum of the feature values of the kth channel at all positions in the feature map of the tth frame in the currently calculated feature map sequence; Based on the spatial attention weight, each feature map in each feature map sequence is spatially weighted to obtain the spatial attention enhanced feature map of each feature map in each feature map sequence: F s,t,i,j,k =S t,i,j ·F t,i,j,k In the formula, F s,t,i,j,k is the feature value of the kth channel at position (i, j) in the spatial attention enhanced feature map of the tth frame in the currently calculated feature map sequence; Based on the temporal attention mechanism, the illumination changes and occlusions of the corresponding traffic light area in each feature map sequence are dynamically modeled to obtain all spatiotemporal attention enhanced feature maps in each feature map sequence.
7. The traffic light recognition method based on artificial intelligence technology according to claim 6 is characterized in that: Based on the temporal attention mechanism, the illumination changes and occlusions of the corresponding traffic light area in each feature map sequence are dynamically modeled to obtain all spatiotemporal attention enhanced feature maps in each feature map sequence, including: Expand the spatial attention enhancement feature map of each feature map in each feature map sequence according to the time dimension and calculate the time attention weight of each feature map: Where, T t is the temporal attention weight of the spatial attention enhanced feature map of the tth frame in the currently calculated feature map sequence, e is a natural constant and its value is 2.71828, T is the total number of frames in the currently calculated feature map sequence, and α m,t is the temporal correlation strength between the spatial attention enhanced feature map of the mth frame and the spatial attention enhanced feature map of the tth frame in the currently calculated feature map sequence, F s-flat,m is the spatial attention enhancement feature map of the mth frame expanded in the time dimension, α m,n is the temporal correlation strength between the spatial attention enhanced feature map of the mth frame and the spatial attention enhanced feature map of the nth frame in the feature map sequence calculated previously; Perform temporal attention weighting on each spatial attention enhancement feature map expanded along the time dimension in each feature map sequence to obtain the corresponding spatiotemporal attention enhancement feature map: F st,t =T t ·F s-flat,t In the formula, F st,t is the spatiotemporal attention enhanced feature map of the tth frame in the currently calculated feature map sequence, F s-flat,t The spatial attention enhanced feature map of the t-th frame expanded in the time dimension; All spatiotemporal attention-enhanced feature maps in each feature map sequence are combined with long short-term memory networks for time series modeling to obtain a complete dataset of dynamic perception of traffic light status for each RGB image sequence.
8. The traffic light recognition method based on artificial intelligence technology according to claim 5 is characterized in that: Based on the vehicle motion trajectory dynamic perception data set, the traffic scene state transition probability within the corresponding prediction period is analyzed and marked to obtain the complete vehicle motion trajectory dynamic perception data set, including: Analyze the velocity vector of each vehicle between adjacent time points based on the vehicle motion trajectory dynamic perception dataset; Determine the state of each vehicle at each time point based on a preset state determination function and the speed vector of each vehicle between adjacent time points; Based on the state of all vehicles in the vehicle motion trajectory dynamic perception data set at each time point, the state transition probability from any state to another state is calculated; The Nth power of the state transition probability of any state in the vehicle motion trajectory dynamic perception data set transferring to another state is regarded as the state transition probability of the traffic scene in the corresponding prediction period, where N is the total number of time intervals included in the corresponding prediction period; The traffic scene state transition probability within the corresponding prediction period is marked in the vehicle motion trajectory dynamic perception dataset to obtain a complete vehicle motion trajectory dynamic perception dataset.
9. A traffic light recognition system based on artificial intelligence technology, characterized in that: The method for executing the traffic light recognition method based on artificial intelligence technology according to any one of claims 1 to 8 comprises: The dynamic modeling module is used to collect RGB image sequences and light sensor data under different weather, lighting conditions and different traffic scenes, and build a corresponding dynamic perception dataset containing traffic light status, vehicle movement trajectory and ambient lighting conditions based on each RGB image sequence and light sensor data; A state representation generation module is used to build a reinforcement learning framework, and generate a state representation of the Markov decision process of each RGB image sequence based on the dynamic perception data set corresponding to each RGB image sequence in the reinforcement learning framework; A classification action definition module is used to define the traffic light state classification action and the corresponding confidence in each RGB image sequence; The recognition model optimization module is used to adopt a multi-agent collaborative training strategy, through a distributed reinforcement learning framework and the traffic light state classification actions and corresponding confidences of all RGB image sequences and the state representation of the Markov decision process, to simultaneously optimize the local recognition model and the global strategy network to obtain the traffic light recognition model; The recognition execution module is used to obtain the traffic light recognition result based on the RGB image sequence of the traffic scene to be recognized, the light sensor data in the corresponding time period, and the traffic light recognition model.
10. A traffic light recognition device based on artificial intelligence technology, characterized in that: Includes processors and storage devices; The storage device is used to store instructions; When the processor executes the instruction, the traffic light recognition method based on artificial intelligence technology as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Traffic light control method based on multi-channel vehicle detection and three-dimensional feature labeling
CN115565388A
Traffic signal lamp strategy evaluation evolution method and system
CN118015860A
Power distribution network overhead line radar external damage prevention device based on AI and vehicle intelligent identification
CN118262286A
Sensing and calculation control fused intelligent traffic light adaptive control method
CN118711367A
Apparatus, method, and computer program for identifying state of signal light, and controller
US20210312198A1