Edge computing traffic light intelligent decision-making method and system integrated with AI video analysis
By integrating AI video analysis on edge computing devices, collecting and processing traffic video streams in real time, generating traffic state vectors and optimizing signal light control strategies, the existing intelligent traffic light control system's response lag and insufficient flexibility in control strategies are solved, and efficient and safe traffic management is achieved.
Patent Information
- Application Number
- CN202510496389.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing intelligent traffic light control system has problems such as poor real-time response, insufficient information utilization, and insufficient flexibility in control strategies, and it is difficult to adapt to real-time fluctuations in traffic flow and dynamic changes in pedestrian behavior.
The intelligent decision-making method of edge computing traffic lights with integrated AI video analysis is adopted. The edge computing device collects traffic video streams in real time, extracts dynamic traffic features, calls the pre-trained spatio-temporal analysis model for multimodal fusion processing, generates traffic state vectors, matches candidate strategies in the decision rule base based on this vector, and adjusts parameters through the strategy optimization model to generate target signal light control parameters.
It realizes the rapid generation of control strategies that adapt to the current road conditions during peak periods or sudden congestion scenarios, improves the real-time and safety of traffic signal control response, and enhances the intelligence of the traffic management system.
Smart Images

Figure CN120089002A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and more particularly, to an edge computing traffic light intelligent decision-making method and system integrating AI video analysis. Background Art
[0002] Intelligent traffic light control aims to optimize the traffic efficiency and safety at intersections by dynamically adjusting signal light parameters. Existing technologies usually adopt a control method with a fixed cycle control strategy. For example, it relies on geomagnetic coils to detect traffic flow and set a fixed green light duration. Such methods are difficult to adapt to the real-time fluctuations of traffic flow and the dynamic changes of pedestrian behavior. Some improved solutions generate control instructions by analyzing intersection surveillance video data through a cloud server, but the network latency caused by transmitting video data to the cloud results in a lag in signal light response, especially in scenarios such as rainy, snowy weather or sudden changes in traffic flow during morning and evening rush hours, where congestion cannot be relieved in a timely manner. In addition, although existing intelligent decision-making systems based on machine learning can process video data, the analysis dimension is single, the generated signal light strategies have safety blind spots, and most systems cannot adapt to the long-term evolution law of traffic flow. In summary, existing technologies have defects such as poor real-time response, insufficient information utilization, and lack of flexibility in control strategies, resulting in low traffic efficiency at intersections, increased risks for pedestrians to cross the street, and difficulty in meeting the precise management and control requirements of complex urban traffic scenarios. Summary of the Invention
[0003] In view of this, the present application provides an edge computing traffic light intelligent decision-making method and system integrating AI video analysis.
[0004] According to one aspect of the embodiments of the present disclosure, there is provided an edge computing traffic light intelligent decision-making method integrating AI video analysis. The method includes: collecting a real-time traffic video stream of a target intersection through an edge computing device, and extracting a set of dynamic traffic features from the real-time traffic video stream, the set of dynamic traffic features including traffic flow distribution features, vehicle behavior trajectory features, and pedestrian movement trend features; invoking a pre-trained spatio-temporal analysis model to perform multi-modal fusion processing on the set of dynamic traffic features to generate a traffic state vector of the target intersection, the traffic state vector being used to characterize the congestion level, vehicle passing priority, and pedestrian safety risk index of the target intersection; matching a candidate control strategy in a preset decision rule library based on the traffic state vector, and adjusting the parameters of the candidate control strategy through a strategy optimization model to generate target signal light control parameters; sending the target signal light control parameters to a traffic signal control terminal of the target intersection, and real-time monitoring the traffic state change data of the target intersection to update the weight parameters of the spatio-temporal analysis model.
[0005] According to another aspect of the embodiments of the present disclosure, an edge computing traffic light intelligent decision-making system is provided, including: one or more processors; and one or more memories, wherein computer-readable code is stored in the memories, and when the computer-readable code is run by the one or more processors, the one or more processors are caused to execute the method as described above.
[0006] The edge computing traffic light intelligent decision-making method integrating AI video analysis provided by the present invention collects the traffic video stream of the target intersection in real time through an edge computing device and extracts a set of dynamic traffic features, calls a pre-trained spatio-temporal analysis model to perform multi-modal fusion processing on traffic flow distribution features, vehicle behavior trajectory features, and pedestrian movement trend features to generate a traffic state vector, based on the vector matching decision rule library to select candidate strategies and dynamically adjust through a strategy optimization model to generate target signal control parameters, and finally updates the model weights by combining real-time feedback data, so that traffic signal control can make decisions by using multi-dimensional information such as the real-time vehicle passing state, pedestrian behavior patterns, and historical traffic laws in the video stream at the same time, effectively solving the problems of response lag and insufficient safety risk estimation caused by traditional timing control; through real-time edge processing and spatio-temporal feature fusion, the latency of video data transmission to the cloud for processing is significantly reduced, ensuring that control strategies adapted to the current road conditions can still be quickly generated during peak hours or in case of sudden congestion; using a strategy optimization model to dynamically weight and simulate verify candidate control strategies, achieving a balanced optimization of traffic efficiency and safety under the premise of complying with traffic rules constraints, avoiding the subjective deviation of manual experience tuning; by continuously monitoring feedback data such as vehicle passing rate and pedestrian waiting queue after the signal light is executed to construct an error index and trigger a model update mechanism, the system can adapt to seasonal changes and long-term evolution trends of traffic flow, maintaining the accuracy and robustness of the decision-making model, thereby comprehensively improving the traffic efficiency of urban intersections, the level of pedestrian safety protection, and the intelligence of the traffic management system. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 is a schematic structural diagram of a traffic application scenario provided by the present application;
[0008] Figure 2 is a schematic flowchart of an edge computing traffic light intelligent decision-making method integrating AI video analysis provided by the present application;
[0009] Figure 3 is a schematic structural diagram of an edge computing traffic light intelligent decision-making system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0010] To facilitate a clearer understanding of the present application, first, a media data processing system for implementing the media data processing method of the present application is introduced, as Figure 1As shown in the figure, the traffic application scenario includes an edge computing device 10 and a terminal cluster. The terminal cluster may include one or more terminals, and the number of terminals will not be limited here. As Figure 1 shown, the terminal cluster may specifically include terminals 1, 2, …, n. It can be understood that terminals 1, 2, 3, …, n can all be network-connected to the edge computing device 10, so that each terminal can perform data interaction with the edge computing device 10 through the network connection. The edge computing device 10 is the edge computing traffic light intelligent decision-making system provided in the embodiments of the present application.
[0011] It can be understood that the edge computing device 10 may refer to a device that performs edge computing for traffic light intelligent decision-making. The terminal may specifically refer to a sensor that collects traffic environment information, such as a geomagnetic coil sensor, an image sensor, etc., but is not limited thereto. Each terminal and the edge computing device 10 can be directly or indirectly connected through wired or wireless communication methods. At the same time, the number of terminals and the edge computing device 10 can be one or at least two, and the present application does not limit this here. The information collected by the terminal is sent to the edge computing device 10 for edge computing and decision-making.
[0012] Further, please refer to Figure 2 , which is a schematic flowchart of an edge computing traffic light intelligent decision-making method integrating AI video analysis provided in the embodiments of the present application. As Figure 2 shown, this method can be executed by the Figure 1 edge computing device 10 therein. Among them, the edge computing traffic light intelligent decision-making method integrating AI video analysis may include the following steps:
[0013] Step S100: Collect the real-time traffic video stream of the target intersection through the edge computing device, and extract the dynamic traffic feature set from the real-time traffic video stream. The dynamic traffic feature set includes traffic flow distribution features, vehicle behavior trajectory features, and pedestrian movement trend features.
[0014] The edge computing device is a computing device that performs data processing and analysis close to the data source. In the embodiments of the present invention, it is used to collect the real-time traffic video stream of the target intersection. The real-time traffic video stream refers to a data stream that records the traffic conditions of the target intersection in the form of continuous video frames in real time. The dynamic traffic feature set is a set of a series of features extracted from the real-time traffic video stream that can reflect the dynamic changes of traffic. Among them, the traffic flow distribution feature describes the traffic flow size and distribution in different regions and different time periods of the target intersection; the vehicle behavior trajectory feature records the driving trajectory, speed change and other behavior information of the vehicle at the intersection; the pedestrian movement trend feature reflects the movement direction, density change and other trends of pedestrians at the intersection.
[0015] In the embodiments of the present invention, the edge computing device can be installed near the target intersection, such as on a traffic signal pole, to ensure the collection of traffic videos at the intersection. The collected real-time traffic video stream is transmitted to the processing module of the edge computing device through a high-speed data interface for subsequent processing. For example, at an intersection, the edge computing device can collect the driving conditions of vehicles on each lane in all directions and the movement of pedestrians on the crosswalk. By analyzing these video streams, traffic flow distribution characteristics can be extracted, such as a large traffic volume in a certain direction lane during the morning rush hour; vehicle behavior trajectory characteristics, such as a vehicle turning and accelerating at the intersection; and pedestrian movement trend characteristics, such as pedestrians quickly moving from one side of the crosswalk to the other when the green light is on.
[0016] As an implementation manner, in step S100, extracting the set of dynamic traffic characteristics from the real-time traffic video stream may specifically include:
[0017] Step S110: Perform frame splitting on the real-time traffic video stream to obtain a continuous sequence of video frames, and perform multi-object detection on each video frame in the video frame sequence to obtain a set of vehicle position coordinates, a set of pedestrian position coordinates, and traffic marker bounding box data.
[0018] Frame splitting is the process of splitting a continuous real-time traffic video stream into a series of independent video frames at a preset time interval. The obtained continuous sequence of video frames is composed of these split video frames arranged in chronological order. Multi-object detection refers to the process of simultaneously detecting multiple different objects in each video frame. In the embodiments of the present invention, the mainly detected objects include, for example, vehicles, pedestrians, and traffic markers. The set of vehicle position coordinates refers to the set composed of the position coordinates of all detected vehicles in the video frame, the set of pedestrian position coordinates is the set of the position coordinates of all detected pedestrians, and the traffic marker bounding box data refers to the bounding box information used to identify the position and range of traffic markers (such as traffic lights, traffic signs, etc.) in the video frame.
[0019] In an embodiment of the present invention, the video processing module of the edge computing device performs frame division processing on the collected real-time traffic video stream. For example, the video stream is divided into frames at a frame rate of 25 frames per second, obtaining a series of consecutive video frames. For each video frame, a multi-object detection algorithm based on deep learning, such as the YOLO (You Only Look Once) algorithm, is used to detect vehicles, pedestrians, and traffic markers. Through this algorithm, the position coordinates of the vehicle can be obtained. For example, the coordinates of a vehicle in the video frame are (x1, y1); the position coordinates of a pedestrian, such as the coordinates of a certain pedestrian are (x2, y2); and the bounding box data of the traffic marker, such as the bounding box coordinates of a traffic signal are (x3, y3, x4, y4), where (x3, y3) is the upper left corner coordinate of the bounding box, and (x4, y4) is the lower right corner coordinate.
[0020] Step S120: Perform spatio-temporal association on the set of vehicle position coordinates through a trajectory tracking algorithm, generate the driving trajectory sequence of each vehicle, and calculate the vehicle speed change curve, acceleration distribution characteristics, and lane occupancy frequency according to the driving trajectory sequence.
[0021] The trajectory tracking algorithm is an algorithm used to track the motion trajectory of a target object in consecutive video frames. In an embodiment of the present invention, it is used to perform spatio-temporal association on the set of vehicle position coordinates. Spatio-temporal association means associating the position coordinates of the same vehicle in different video frames in chronological order to generate the driving trajectory sequence of the vehicle. The driving trajectory sequence is a sequence composed of the position coordinates of the vehicle at different times, which records the driving path of the vehicle at the target intersection. The vehicle speed change curve is a curve describing the change of vehicle speed over time, obtained by calculating the change of the vehicle position in the driving trajectory sequence. The acceleration distribution characteristics refer to the distribution of vehicle acceleration in different time periods, which can be obtained by taking the derivative of the speed change curve. The lane occupancy frequency refers to the frequency of vehicle occupancy in each lane, obtained by counting the number of times the vehicle appears in different lanes.
[0022] In the embodiments of the present invention, trajectory tracking algorithms such as Kalman filtering can be used to perform spatio-temporal association on the vehicle position coordinate set. For example, in a series of consecutive video frames, based on information such as the position and speed of the vehicle, it is determined whether a certain vehicle in different frames is the same vehicle, and its position coordinates are associated to generate the driving trajectory sequence of the vehicle. When calculating the vehicle speed change curve according to the driving trajectory sequence, first calculate the distance difference between the vehicle positions at two adjacent time points, and then divide it by the time interval to obtain the average speed during this time period. Calculate the speeds in multiple time periods in turn, so as to draw the speed change curve. For the acceleration distribution characteristics, numerically differentiate the speed change curve to obtain the acceleration values in different time periods, and then analyze the distribution of the acceleration. When calculating the lane occupancy frequency, count the number of frames in which the vehicle appears in each lane, and then divide it by the total number of frames to obtain the occupancy frequency of each lane.
[0023] Step S130: Perform group behavior analysis on the pedestrian position coordinate set, and extract the pedestrian moving direction consistency feature, the pedestrian density fluctuation feature, and the crosswalk stay duration.
[0024] Group behavior analysis refers to the process of comprehensively analyzing the behaviors of a group of pedestrians. In the embodiments of the present invention, it is used to analyze the pedestrian position coordinate set. The pedestrian moving direction consistency feature refers to the degree of consistency of the directions of a group of pedestrians during the moving process, which can be measured by existing methods such as calculating the included angle of the pedestrian moving directions. The pedestrian density fluctuation feature refers to the change of the pedestrian density in different time periods, which can be obtained by calculating the change of the number of pedestrians in different regions. The crosswalk stay duration refers to the length of time that pedestrians stay on the crosswalk, which can be calculated by recording the time when pedestrians enter and leave the crosswalk.
[0025] Exemplarily, when performing group behavior analysis on the set of pedestrian position coordinates, pedestrians are first divided into different groups. For the feature of pedestrian movement direction consistency, calculate the included angle of the movement directions of pedestrians in each group. If the included angle is small, it indicates a high consistency in the movement directions of pedestrians in that group. For the feature of pedestrian density fluctuation, divide the target intersection into multiple small areas, count the number of pedestrians in each area at different time periods, calculate the change rate of the number of pedestrians, and thus obtain the fluctuation of pedestrian density. When calculating the residence time on the crosswalk, determine whether a pedestrian enters the crosswalk based on the pedestrian position coordinates, and record the entry and exit times. The difference between the two is the residence time. For example, on a crosswalk, a group of pedestrians enters during a certain period. By analyzing the change in their position coordinates, it is found that the movement directions of most pedestrians are basically the same, indicating a high consistency in the movement directions of pedestrians. At the same time, by counting the number of pedestrians in the crosswalk area, it is found that the number of pedestrians first increases and then decreases over time, reflecting the fluctuation feature of pedestrian density. For a certain pedestrian, record the time when they enter and leave the crosswalk, and the obtained residence time is 10 seconds.
[0026] Step S140: Normalize the vehicle speed change curve, acceleration distribution feature, lane occupancy frequency, pedestrian movement direction consistency feature, pedestrian density fluctuation feature, and residence time on the crosswalk to generate a set of dynamic traffic features.
[0027] Normalization is the process of converting data in different ranges to the same range. In the embodiments of the present invention, it is used to uniformly process data of different features such as the vehicle speed change curve, acceleration distribution feature, lane occupancy frequency, pedestrian movement direction consistency feature, pedestrian density fluctuation feature, and residence time on the crosswalk, so that they are comparable. The set of dynamic traffic features is a set composed of these features after normalization processing, and it can more accurately reflect the traffic dynamic situation of the target intersection.
[0028] In the embodiments of the present invention, min-max normalization can be used to process each feature. For the vehicle speed change curve, first find the minimum and maximum values in the curve, then subtract the minimum value from each speed value, and divide by the difference between the maximum value and the minimum value to obtain the normalized speed value. For the acceleration distribution feature, lane occupancy frequency, pedestrian movement direction consistency feature, pedestrian density fluctuation feature, and residence time on the crosswalk, similar methods are also used for normalization processing. For example, the minimum value in the vehicle speed change curve is 10 km / h, the maximum value is 60 km / h, and the speed value at a certain moment is 30 km / h. Then the normalized speed value is (30 - 10) / (60 - 10) = 0.4. Combine all the features after normalization processing to obtain the set of dynamic traffic features.
[0029] Step S200: Invoke a pre-trained spatio-temporal analysis model to perform multi-modal fusion processing on the dynamic traffic feature set, generating a traffic state vector for the target intersection. The traffic state vector is used to characterize the congestion level, vehicle passing priority, and pedestrian safety risk index of the target intersection.
[0030] The pre-trained spatio-temporal analysis model is a model trained on a large amount of historical traffic data, which can analyze and process the input dynamic traffic feature set. Multi-modal fusion processing refers to fusing different types of dynamic traffic features (such as vehicle-related features and pedestrian-related features) to obtain more comprehensive and accurate information. The traffic state vector is a vector containing multiple elements, where each element corresponds to a traffic state index of the target intersection. In the embodiments of the present invention, it mainly includes the congestion level, vehicle passing priority, and pedestrian safety risk index. The congestion level is used to describe the congestion degree of the target intersection, which can be divided into different levels such as mild congestion, moderate congestion, and severe congestion; the vehicle passing priority is used to determine the passing order of different vehicles at the intersection; the pedestrian safety risk index is used to evaluate the safety risk degree of pedestrians at the intersection.
[0031] In the embodiments of the present invention, the dynamic traffic feature set generated in step S100 is input into the pre-trained spatio-temporal analysis model. The spatio-temporal analysis model can be a combined model based on a convolutional neural network (CNN) and a recurrent neural network (RNN), a spatio-temporal graph convolutional network (ST-GCN), a Transformer architecture model, etc. The spatio-temporal analysis model performs multi-modal fusion processing on the dynamic traffic feature set, comprehensively considering various feature information of vehicles and pedestrians. For example, the model combines features such as the speed change of vehicles, lane occupancy, and the moving direction and density of pedestrians to determine the congestion level of the target intersection. If the vehicle speed is generally low, the lane occupancy rate is high, pedestrians move slowly and are dense, then the model may determine that the intersection is in a severe congestion level. For the vehicle passing priority, the model will determine according to factors such as the driving direction of the vehicle and whether it is an emergency vehicle. For the pedestrian safety risk index, the model will consider factors such as the moving trend of pedestrians and the distance from vehicles for evaluation. Finally, the model outputs a traffic state vector containing the congestion level, vehicle passing priority, and pedestrian safety risk index.
[0032] As an implementation, the pre-training process of the above spatio-temporal analysis model includes:
[0033] Step S201: Collect a historical traffic video dataset and the corresponding labeled traffic state labels. The labeled traffic state labels include manually labeled congestion levels, vehicle passing priority coefficients, and pedestrian accident records.
[0034] A historical traffic video dataset refers to a collection of traffic video data for a target intersection or other similar intersections collected over a period of time in the past. Labeling traffic status tags are tags used to label the traffic status corresponding to each video segment in the historical traffic video dataset. Among them, the congestion level manually labeled is the degree of congestion judged and labeled manually according to the traffic conditions in the video; the vehicle passing priority coefficient is a value used to measure the passing priority of vehicles at the intersection; the pedestrian accident record refers to whether a pedestrian accident occurs in the video and the relevant information about the accident.
[0035] Step S202: Perform spatio-temporal slicing on the historical traffic video dataset to generate a training sample sequence, and perform data augmentation processing on each training sample in the training sample sequence. The data augmentation processing includes simulating light changes, adding occluders, and perturbing the video frame sampling rate.
[0036] Spatio-temporal slicing is a process of dividing the historical traffic video dataset according to the time and space dimensions. The generated training sample sequence is a sequence composed of these segmented video segments. Data augmentation processing refers to performing a series of transformations on the training samples to increase the diversity and quantity of the training data and improve the generalization ability of the model. Simulating light changes is to simulate traffic scenarios under different lighting conditions by adjusting parameters such as the brightness and contrast of the video; adding occluders is to add some simulated occluders in the video, such as billboards, trees, etc., to simulate the occlusion situations that may occur in actual traffic; perturbing the video frame sampling rate is to randomly adjust the sampling rate of the video to simulate video acquisition under different frame rates.
[0037] In the embodiment of the present invention, when performing spatio-temporal slicing on the historical traffic video dataset, the video is segmented according to a preset time interval (such as every 5 minutes) and spatial region (such as dividing the intersection into multiple sub-regions) to obtain a series of training samples. For each training sample, data augmentation processing is performed. For example, use image processing algorithms to simulate light changes in the video, increase or decrease the brightness of the video by a preset ratio; simulate occlusion situations by adding virtual occluder images in the video; randomly adjust the sampling rate of the video, such as adjusting the original video sampling rate of 25 frames per second to 20 frames per second or 30 frames per second. In this way, more training samples in different scenarios can be obtained, improving the adaptability of the model to various actual traffic situations.
[0038] Step S203: Construct an initial spatio-temporal analysis model. The initial spatio-temporal analysis model includes a time feature extraction network and a spatial feature extraction network connected in parallel, as well as a cross-modal fusion layer.
[0039] The initial spatio-temporal analysis model is a basic model for analyzing and processing traffic characteristics. It consists of a time feature extraction network, a space feature extraction network, and a cross-modal fusion layer. The time feature extraction network is used to extract information on traffic characteristics in the time dimension, such as the periodic fluctuation characteristics of traffic flow and the characteristics of sudden abnormal events. The space feature extraction network is used to extract information on traffic characteristics in the space dimension, such as the vehicle density gradient characteristics and the pedestrian aggregation hotspot characteristics. The cross-modal fusion layer is used to fuse the features extracted by the time feature extraction network and the space feature extraction network to obtain more comprehensive spatio-temporal feature information.
[0040] In an embodiment of the present invention, a deep learning architecture can be used to construct the initial spatio-temporal analysis model. The time feature extraction network can use a temporal convolutional network (TCN), which can effectively capture the change information of traffic characteristics in the time dimension. The space feature extraction network can use a region segmentation network, such as a U-Net network, which can divide the target intersection into multiple sub-regions and extract the spatial features of each sub-region. The cross-modal fusion layer can adopt a cross-modal attention mechanism to achieve feature fusion by assigning weights to the time and space features. For example, the dynamic traffic feature set is input into the time feature extraction network and the space feature extraction network. The time feature extraction network extracts the periodic fluctuation characteristics of traffic flow in the time dimension, and the space feature extraction network extracts the density gradient characteristics of vehicles in different sub-regions. Then the cross-modal fusion layer fuses these two features to obtain more comprehensive spatio-temporal feature information.
[0041] Step S204: Iteratively train the initial spatio-temporal analysis model through a multi-task loss function, which includes the weighted sum of the congestion level prediction error, the vehicle passing priority classification error, and the pedestrian safety risk regression error.
[0042] The multi-task loss function is a function used to measure the error of the model on multiple tasks. In an embodiment of the present invention, it consists of the weighted sum of the congestion level prediction error, the vehicle passing priority classification error, and the pedestrian safety risk regression error. The congestion level prediction error refers to the difference between the congestion level predicted by the model and the actually labeled congestion level. The vehicle passing priority classification error refers to the difference between the classification result of the vehicle passing priority by the model and the actually labeled passing priority. The pedestrian safety risk regression error refers to the difference between the pedestrian safety risk index predicted by the model and the actually labeled pedestrian safety risk index. The weighted sum means assigning different weights to these three errors and then adding them to obtain the total loss value. Iterative training means continuously adjusting the parameters of the model to gradually reduce the value of the multi-task loss function, thereby improving the performance of the model.
[0043] In the embodiments of the present invention, optimization algorithms such as stochastic gradient descent are used to iteratively train the initial spatio-temporal analysis model. In each training iteration, the training samples are input into the model, and the model outputs the predicted congestion level, vehicle passing priority, and pedestrian safety risk indicators. Then, the errors of these three indicators are calculated, namely, the congestion level prediction error, the vehicle passing priority classification error, and the pedestrian safety risk regression error. The value of the multi-task loss function is calculated according to the preset weights. The parameters of the model are adjusted through the optimization algorithm to gradually reduce the value of the multi-task loss function.
[0044] Step S205: When the convergence speed of the multi-task loss function on the validation set is lower than the preset threshold, freeze the parameters of the time feature extraction network and only update the weights of the space feature extraction network.
[0045] The validation set is a data set used to evaluate the performance of the model during the training process, and it is independent of the training set. The convergence speed refers to the speed at which the multi-task loss function decreases with the increase of the training iteration times on the validation set. The preset threshold is a preset value used to judge whether the convergence speed is sufficient. Freezing the parameters of the time feature extraction network means not updating the parameters of the time feature extraction network during the training process and only updating the weights of the space feature extraction network. This can reduce the training complexity of the model and at the same time utilize the time feature information already learned by the time feature extraction network, focusing on optimizing the performance of the space feature extraction network.
[0046] In the embodiments of the present invention, during the training process, the model is regularly evaluated using the validation set, and the value of the multi-task loss function on the validation set is calculated. If it is found that the convergence speed of the multi-task loss function on the validation set is lower than the preset threshold, it means that the time feature extraction network has learned relatively stable time feature information. At this time, the parameters of the time feature extraction network can be frozen. For example, assume that the preset threshold is 0.01. In a certain training stage, the convergence speed of the multi-task loss function on the validation set is 0.005, which is lower than the preset threshold. Then, the parameters of the time feature extraction network can be frozen, and only the weights of the space feature extraction network are updated. This can avoid excessive adjustment of the parameters of the time feature extraction network in subsequent training and at the same time make the model pay more attention to the learning and optimization of space features.
[0047] As an implementation manner, step S200, calling the pre-trained spatio-temporal analysis model to perform multi-modal fusion processing on the dynamic traffic feature set to generate the traffic state vector of the target intersection, may specifically include:
[0048] Step S210: Input the dynamic traffic feature set into the time feature extraction branch of the spatio-temporal analysis model, and extract the periodic fluctuation features and sudden abnormal event features of the traffic flow in the time dimension through the time convolutional network.
[0049] The time feature extraction branch is the part of the spatio-temporal analysis model used to extract the information of traffic features in the time dimension, and it can be composed of a temporal convolutional network (TCN). The temporal convolutional network is a convolutional neural network that can capture the temporal dependencies in sequential data. In the embodiments of the present invention, it is used to extract the periodic fluctuation features and sudden abnormal event features of traffic flow in the time dimension. The periodic fluctuation features refer to the regular changes of traffic flow within a preset time period, such as the increase in traffic volume during the morning and evening rush hours every day; the sudden abnormal event features refer to the sudden abnormal changes in traffic flow, such as the sharp decrease in traffic volume caused by a traffic accident.
[0050] In the embodiments of the present invention, the dynamic traffic feature set generated in step S100 is input into the time feature extraction branch of the spatio-temporal analysis model. The temporal convolutional network will process the dynamic traffic feature set and extract the features of traffic flow in the time dimension through convolutional operations. For example, for traffic flow data, the temporal convolutional network can identify the periodic increase in traffic volume during the morning and evening rush hours every day, and the sudden decrease in traffic volume caused by a traffic accident at a certain moment and other features. By extracting these features, the model can better understand the change law of traffic flow in the time dimension and provide more accurate information for subsequent traffic state analysis.
[0051] Step S220: Input the dynamic traffic feature set into the spatial feature extraction branch of the spatio-temporal analysis model, divide the target intersection into multiple sub-regions through a region segmentation network, and respectively extract the vehicle density gradient features, pedestrian aggregation hot spot features, and signal light visibility distribution of each sub-region.
[0052] The spatial feature extraction branch is the part of the spatio-temporal analysis model used to extract the information of traffic features in the spatial dimension, and it is composed of a region segmentation network. The region segmentation network is a network that can segment and identify different regions in an image or video. In the embodiments of the present invention, it is used to divide the target intersection into multiple sub-regions. The vehicle density gradient feature refers to the density change of vehicles in different sub-regions, which is obtained by calculating the difference in vehicle density between adjacent sub-regions; the pedestrian aggregation hot spot feature refers to the area and degree of pedestrian aggregation at the intersection, which is obtained by counting the number of pedestrians in the sub-region; the signal light visibility distribution refers to the visibility of the signal light in different sub-regions, which is obtained by analyzing factors such as the brightness and occlusion of the signal light in the video.
[0053] In the embodiment of the present invention, the dynamic traffic feature set is input into the spatial feature extraction branch of the spatio-temporal analysis model. The regional segmentation network divides the target intersection into multiple sub-regions. For example, a crossroads is divided into four quadrants and each lane and other sub-regions. For each sub-region, vehicle density gradient features, pedestrian aggregation hot spot features, and signal visibility distributions are extracted. For example, by counting the number of vehicles in each sub-region and calculating the difference in vehicle density between adjacent sub-regions, vehicle density gradient features are obtained; the number of pedestrians in the sub-region is counted, the regions with a large number of pedestrians are found, and pedestrian aggregation hot spot features are determined; the brightness and occlusion conditions of the traffic lights in each sub-region in the video are analyzed to evaluate the signal visibility distribution. These spatial feature information can help the model better understand the spatial traffic conditions of the target intersection.
[0054] Step S230: Through the cross-modal attention mechanism, weight distribution is performed on the periodic fluctuation features, sudden abnormal event features, vehicle density gradient features, pedestrian aggregation hot spot features, and signal visibility distributions in the time dimension to generate a fused global spatio-temporal feature map.
[0055] The cross-modal attention mechanism is a mechanism that can perform weight distribution on features of different modalities (such as time and space modalities), and weights them according to the importance of the features, so as to highlight important feature information. In the embodiment of the present invention, it is used to perform weight distribution on different types of features such as periodic fluctuation features, sudden abnormal event features, vehicle density gradient features, pedestrian aggregation hot spot features, and signal visibility distributions in the time dimension. The global spatio-temporal feature map is a feature map fused by these features after weight distribution, and it contains comprehensive information about the target intersection in the time and space dimensions.
[0056] In the embodiment of the present invention, the cross-modal attention mechanism performs weight distribution on these features according to their importance. For example, when the traffic flow is in normal periodic fluctuations, the periodic fluctuation features in the time dimension may have a higher weight; while when a sudden abnormal event occurs, the weight of the sudden abnormal event features will increase. For spatial features, such as in the pedestrian aggregation hot spot area, the weight of the pedestrian aggregation hot spot features will be relatively high. Through this weight distribution, different types of features are fused to generate a fused global spatio-temporal feature map. This global spatio-temporal feature map can more comprehensively and accurately reflect the traffic state of the target intersection and provide richer information for the subsequent generation of the traffic state vector.
[0057] Step S240: Map the global spatio-temporal feature map to a preset traffic state space, and output a traffic state vector including congestion level score, vehicle passing priority weight, and pedestrian safety risk index through a fully connected layer.
[0058] The preset traffic state space is a predefined space for representing traffic states, which contains various possible traffic state indicators. The fully connected layer is a neural network layer that connects each element in the input feature map to each neuron in the output layer, and maps the input features to the output space through linear transformation and activation functions. The congestion level score is a quantitative score for the congestion degree of the target intersection. The vehicle passing priority weight is a weight value used to determine the vehicle passing order. The pedestrian safety risk index is a quantitative indicator for the safety risk degree of pedestrians at the intersection. The traffic state vector is a vector composed of the congestion level score, the vehicle passing priority weight, and the pedestrian safety risk index, which can comprehensively reflect the traffic state of the target intersection.
[0059] In the embodiment of the present invention, the global spatio-temporal feature map generated in step S230 is input into the fully connected layer. The fully connected layer performs linear transformation and activation function processing on the global spatio-temporal feature map, and maps it to the preset traffic state space. Through this mapping, a traffic state vector containing the congestion level score, the vehicle passing priority weight, and the pedestrian safety risk index is output. For example, according to the information in the global spatio-temporal feature map, the fully connected layer calculates that the congestion level score of the target intersection is 7 points (out of 10), the passing priority weight of a certain vehicle is 0.8, and the pedestrian safety risk index is 0.3, and combines these values into the traffic state vector [7, 0.8, 0.3].
[0060] Step S300: Based on the traffic state vector, match the candidate control strategies in the preset decision rule library, and adjust the parameters of the candidate control strategies through the policy optimization model to generate the target signal control parameters.
[0061] The preset decision rule library is a pre-established database containing various traffic states and corresponding control strategies, which formulates corresponding signal control strategies according to different traffic states. The candidate control strategies refer to the control strategies matched from the decision rule library corresponding to the current traffic state vector, and these strategies contain some basic control parameters, such as the green light duration, the yellow light transition interval, etc. The policy optimization model is a trained model used to adjust and optimize the parameters of the candidate control strategies to improve traffic passing efficiency and safety. The target signal control parameters are the final parameters obtained after being adjusted by the policy optimization model and are used to control the signals at the target intersection.
[0062] In an embodiment of the present invention, the traffic state vector generated in step S200 is matched with a preset decision rule library. For example, if the traffic state vector shows that the target intersection is moderately congested, and the vehicle passing priority is high while the pedestrian safety risk is low, then a corresponding candidate control strategy is matched from the decision rule library, such as increasing the green light duration in the vehicle passing direction. These candidate control strategies are input into the policy optimization model, and the policy optimization model adjusts the parameters of the candidate control strategies according to the current traffic state vector. For example, by analyzing historical traffic data and real-time traffic conditions, the policy optimization model adjusts the green light duration in the vehicle passing direction from the original 30 seconds to 40 seconds, and finally generates the target signal control parameters.
[0063] As an implementation manner, the training process of the above policy optimization model includes the following steps:
[0064] Step S301: Extract multiple historical traffic state vectors and corresponding actually executed control parameters from the historical traffic video dataset, and calculate the true passing efficiency and safety indicators of each actually executed control parameter within a subsequent time window.
[0065] The historical traffic state vector is a traffic state vector corresponding to different time periods extracted from the historical traffic video dataset, which contains information such as the congestion level, vehicle passing priority, and pedestrian safety risk indicators at that time. The actually executed control parameter refers to the signal control parameter actually adopted in the historical traffic scenario, such as the green light duration, yellow light transition interval, etc. The subsequent time window refers to a period of time after adopting a certain actually executed control parameter, which is used to evaluate the effect of this control parameter. The true passing efficiency refers to the actual passing efficiency of the traffic system after adopting a certain actually executed control parameter, such as the reduction rate of vehicle average delay, the improvement amplitude of green light utilization rate, etc.; the safety indicator refers to the safety condition of the traffic system after adopting this control parameter, such as the occurrence rate of pedestrian crossing conflict events, the emergency braking frequency, etc.
[0066] In an embodiment of the present invention, historical traffic state vectors and corresponding actual execution control parameters are extracted from a historical traffic video dataset. For example, for the morning rush hour of a certain day, the traffic state vector of this period is extracted, such as the congestion level is severe congestion, the vehicle passing priority is high, the pedestrian safety risk is low, and the actual signal light control parameters at that time, such as the green light duration is 25 seconds and the yellow light transition interval is 3 seconds. Then, the true passing efficiency and safety indicators of the actual execution control parameters within a subsequent time window (such as the next 10 minutes) are calculated. By statistically analyzing the driving data of vehicles and the movement data of pedestrians, indicators such as the average vehicle delay reduction rate, the improvement amplitude of green light utilization rate, the incidence rate of pedestrian crossing conflict events, and the emergency braking frequency are calculated. For example, the calculated average vehicle delay reduction rate is 15%, the improvement amplitude of green light utilization rate is 10%, the incidence rate of pedestrian crossing conflict events is 0.5%, and the emergency braking frequency is 2 times per minute.
[0067] Step S302: Construct the neural network structure of the policy optimization model. The neural network structure includes a policy generator and a discriminator. The policy generator is used to generate control parameters according to the input traffic state vector, and the discriminator is used to evaluate the distribution consistency between the generated control parameters and the true control parameters.
[0068] The neural network structure of the policy optimization model consists of a policy generator and a discriminator. The policy generator is a neural network model that generates corresponding signal light control parameters according to the input traffic state vector. The discriminator is also a neural network model that is used to evaluate the distribution consistency between the control parameters generated by the policy generator and the true control parameters, that is, to determine whether the generated control parameters have similar distribution characteristics to the actually adopted control parameters.
[0069] In an embodiment of the present invention, the neural network structure of the policy optimization model is constructed using the architecture of a generative adversarial network (GAN). The policy generator can adopt the structure of a multi-layer perceptron (MLP). It receives the traffic state vector as input and outputs the generated signal light control parameters through a series of linear transformations and activation functions. The discriminator can also adopt the structure of an MLP. It receives the generated control parameters and the true control parameters as input and judges their distribution consistency through feature analysis of the input parameters. For example, the policy generator generates a set of signal light control parameters, such as the green light duration is 35 seconds and the yellow light transition interval is 4 seconds, according to the input traffic state vector [7, 0.8, 0.3]; the discriminator will compare this set of generated control parameters with the true control parameters in the historical data and evaluate their distribution consistency.
[0070] Step S303: Alternately optimize the policy generator and the discriminator through an adversarial training algorithm until the discriminator can no longer distinguish the difference between the generated control parameters and the true control parameters.
[0071] In an embodiment of the present invention, the goal of the policy generator is to generate control parameters that are as close as possible to the true control parameters, while the goal of the discriminator is to accurately distinguish between the generated control parameters and the true control parameters. When the discriminator cannot distinguish the difference between the generated control parameters and the true control parameters, it indicates that the policy generator has learned the distribution characteristics of the true control parameters, and the model has achieved better performance.
[0072] In an embodiment of the present invention, the adversarial training algorithm is used to alternately optimize the policy generator and the discriminator. In each training iteration, first, the parameters of the discriminator are fixed, and the parameters of the policy generator are updated to make the control parameters generated by the policy generator closer to the true control parameters to deceive the discriminator. Then, the parameters of the policy generator are fixed, and the parameters of the discriminator are updated to enable the discriminator to better identify the difference between the generated control parameters and the true control parameters.
[0073] Step S304: Integrate the trained policy generator with the preset traffic rule constraints to generate a policy optimization model that can simultaneously meet efficiency optimization and rule compliance.
[0074] The preset traffic rule constraints refer to a series of pre-set traffic rules and restrictions, such as the minimum green light duration, the maximum red light waiting time, and the pedestrian priority principle. Integrating the trained policy generator with the preset traffic rule constraints means adding these traffic rule constraints during the process of the policy generator generating control parameters to ensure that the generated control parameters can both improve traffic passing efficiency and comply with traffic rules.
[0075] As an implementation, in step S303, alternately optimizing the policy generator and the discriminator through the adversarial training algorithm may specifically include:
[0076] Step S3031: Introduce a traffic rule penalty term into the loss function of the policy generator. The traffic rule penalty term is used to perform gradient penalty on the control parameters that violate the minimum green light duration, the maximum red light waiting time, and the pedestrian priority principle.
[0077] The loss function is a function used to measure the difference between the model output and the target value. In the policy generator, it is used to measure the difference between the generated control parameters and the true control parameters. The traffic rule penalty term is an additional term introduced into the loss function to penalize the control parameters that violate traffic rules. Gradient penalty means adjusting the gradient of the loss function to enable the model to avoid generating control parameters that violate traffic rules during the training process. The minimum green light duration refers to the shortest time for the signal light to turn green, the maximum red light waiting time refers to the longest time for vehicles or pedestrians to wait for the red light, and the pedestrian priority principle refers to giving priority to the right of way of pedestrians in traffic control.
[0078] In an embodiment of the present invention, a traffic rule penalty term is introduced into the loss function of the policy generator. For example, when the green light duration in the generated control parameters is less than the minimum green light duration, the traffic rule penalty term will increase the value of the loss function, so that the policy generator can avoid generating such control parameters in subsequent training. Specifically, the traffic rule penalty term can be set as a function related to the degree of violation of traffic rules. For example, for the case of violating the minimum green light duration, the value of the penalty term can be calculated according to the difference between the green light duration and the minimum green light duration. Through this gradient penalty mechanism, the policy generator can learn to generate control parameters that conform to traffic rules.
[0079] Step S3032: Add a feature matching constraint to the loss function of the discriminator, so that the generated control parameters maintain statistical distribution consistency with the real control parameters in the hidden layer feature space.
[0080] The feature matching constraint is a constraint condition added to the loss function of the discriminator, which requires that the generated control parameters have similar statistical distribution characteristics with the real control parameters in the hidden layer feature space. The hidden layer feature space refers to the feature representation space of the hidden layer in the neural network. By analyzing the distribution of the generated control parameters and the real control parameters in this space, their similarity can be evaluated. Statistical distribution consistency means that the distributions of the generated control parameters and the real control parameters in the hidden layer feature space have similar statistical characteristics, such as mean, variance, etc.
[0081] In an embodiment of the present invention, a feature matching constraint is added to the loss function of the discriminator. When the discriminator judges the generated control parameters and the real control parameters, it not only considers their surface features, but also considers their distribution in the hidden layer feature space. For example, by calculating the difference between the mean and variance of the generated control parameters and the real control parameters in the hidden layer feature space, this difference is added as a feature matching constraint term to the loss function of the discriminator. In this way, during the training process, the discriminator prompts the control parameters generated by the policy generator to maintain statistical distribution consistency with the real control parameters in the hidden layer feature space, thereby improving the quality of the generated control parameters.
[0082] Step S3033: Update the weight parameters of the policy generator and the discriminator through a moving average mechanism, and perform a simulation environment test on the generated control parameters after each iteration to screen out valid samples.
[0083] The sliding average mechanism is a method for updating the weight parameters of a model. By performing a weighted average on historical weight parameters, the parameter updates of the model become smoother and more stable. In the embodiments of the present invention, it is used to update the weight parameters of the policy generator and the discriminator. The simulation environment test refers to testing the generated control parameters in a virtual traffic environment to evaluate their effects in actual traffic scenarios. An effective sample refers to the generated control parameters that, after being tested in the simulation environment, are considered to be able to improve traffic efficiency and safety.
[0084] In the embodiments of the present invention, the sliding average mechanism is used to update the weight parameters of the policy generator and the discriminator. For example, in each training iteration, the weight of the model is updated based on the weighted average of the current weight parameters and historical weight parameters, where the weight of the historical weight parameters gradually decreases over time. After each iteration, the generated control parameters are input into the simulation environment test. The simulation environment can be a virtual traffic scenario constructed based on real traffic data. By simulating the control process of traffic lights in this scenario, the effects of the generated control parameters are evaluated. For example, metrics such as the average delay time of vehicles and the crossing time of pedestrians are calculated, and the generated control parameters that can improve these metrics are selected as effective samples.
[0085] Step S3034: When the policy generator reaches the preset traffic efficiency improvement rate in the simulation environment and the number of violations is lower than the tolerance threshold, terminate the adversarial training process.
[0086] The preset traffic efficiency improvement rate is a pre-set indicator for measuring the degree of improvement in traffic efficiency by the control parameters generated by the policy generator. It represents the improvement ratio of traffic efficiency relative to the original after adopting the generated control parameters. The number of violations refers to the number of times the generated control parameters violate traffic rules, and the tolerance threshold is a pre-set maximum number of allowed violations. When the policy generator reaches the preset traffic efficiency improvement rate in the simulation environment and the number of violations is lower than the tolerance threshold, it indicates that the policy generator has learned the ability to generate better control parameters. At this time, the adversarial training process can be terminated. In the embodiments of the present invention, during the adversarial training process, the control parameters generated by the policy generator are continuously tested in the simulation environment.
[0087] As an implementation, in step S300, the candidate control strategy is parameter-adjusted by the policy optimization model to generate the target signal light control parameters, which may specifically include:
[0088] Step S310: Select at least two candidate control strategies with the highest matching degree to the traffic state vector from the decision rule library. The candidate control strategies include the green light duration baseline parameter, the yellow light transition interval threshold, and the pedestrian dedicated phase trigger condition.
[0089] The decision rule library is a pre-established database containing various traffic states and corresponding control strategies. In this step, according to the traffic state vector generated in step S200, at least two candidate control strategies with the highest matching degree are selected from the decision rule library. Candidate control strategies refer to these selected control strategies, which contain some basic control parameters. For example, the green light duration baseline parameter refers to the basic duration when the signal light is green; the yellow light transition interval threshold refers to the time interval when the yellow light is on when the signal light switches from green to red; the pedestrian dedicated phase trigger condition refers to the situation under which the pedestrian dedicated signal light phase is triggered.
[0090] In the embodiment of the present invention, the traffic state vector is input into the decision rule library for matching. For example, the traffic state vector shows that the congestion level of the target intersection is moderately congested, the vehicle passing priority is relatively high, and the pedestrian safety risk is relatively low. According to this information, two candidate control strategies with the highest matching degree are selected from the decision rule library. For one strategy, the green light duration baseline parameter is 30 seconds, the yellow light transition interval threshold is 3 seconds, and the pedestrian dedicated phase trigger condition is that the number of pedestrians exceeds 10; for the other strategy, the green light duration baseline parameter is 35 seconds, the yellow light transition interval threshold is 4 seconds, and the pedestrian dedicated phase trigger condition is that the number of pedestrians exceeds 15.
[0091] Step S320: Input each candidate control strategy and the traffic state vector into the strategy evaluation module in the strategy optimization model, and calculate the expected traffic efficiency gain and safety risk attenuation coefficient of each candidate control strategy.
[0092] The strategy evaluation module in the strategy optimization model can be a multi-layer perceptron structure, which is used to evaluate the effect of candidate control strategies. It receives the candidate control strategies and the traffic state vector as inputs, and calculates the expected traffic efficiency gain and safety risk attenuation coefficient of each candidate control strategy through the analysis of historical traffic data and real-time traffic conditions. The expected traffic efficiency gain refers to the improvement amplitude of traffic efficiency after adopting a certain candidate control strategy compared with the original; the safety risk attenuation coefficient refers to the reduction degree of the safety risk of the traffic system after adopting this strategy compared with the original.
[0093] In the embodiment of the present invention, the selected candidate control strategy and traffic state vector in step S310 are input into the policy evaluation module. The policy evaluation module analyzes the execution effects of each candidate control strategy in historical traffic scenarios, and combines the current traffic state vector to calculate the expected traffic efficiency gain and safety risk attenuation coefficient. For example, for the first candidate control strategy, by analyzing historical data, it is found that in a similar traffic state, after adopting this strategy, the average vehicle delay reduction rate is 12%, and the occurrence rate of pedestrian crossing conflict events is reduced by 0.3%. Then the calculated expected traffic efficiency gain is 12%, and the safety risk attenuation coefficient is 0.3%; for the second candidate control strategy, the calculated expected traffic efficiency gain is 15%, and the safety risk attenuation coefficient is 0.4%.
[0094] Step S330: Dynamically weight each candidate control strategy according to the expected traffic efficiency gain and safety risk attenuation coefficient to generate an initial set of control parameters.
[0095] Dynamically weighting means assigning different weights to each candidate control strategy according to the magnitudes of the expected traffic efficiency gain and safety risk attenuation coefficient, and then combining the parameters of each candidate control strategy according to the weights to generate an initial set of control parameters. The initial set of control parameters is a set composed of the parameters of the candidate control strategies after dynamic weighting, which comprehensively considers the two factors of traffic efficiency and safety risk. In the embodiment of the present invention, different weights are set for the expected traffic efficiency gain and safety risk attenuation coefficient respectively. For example, the weight of the expected traffic efficiency gain is 0.6, and the weight of the safety risk attenuation coefficient is 0.4. For the first candidate control strategy, its expected traffic efficiency gain is 12%, and the safety risk attenuation coefficient is 0.3%. Then the comprehensive weight of this strategy is 0.6×12% + 0.4×0.3% = 7.32%; for the second candidate control strategy, its expected traffic efficiency gain is 15%, and the safety risk attenuation coefficient is 0.4%. Then the comprehensive weight of this strategy is 0.6×15% + 0.4×0.4% = 9.16%. According to these comprehensive weights, the parameters of the two candidate control strategies are combined to generate an initial set of control parameters. For example, the green light duration may be the result of weighted averaging of the green light duration baseline parameters of the two strategies according to the comprehensive weights.
[0096] Step S340: Simulate the execution effects of the initial set of control parameters in historical traffic scenarios through a reinforcement learning algorithm, and select the control parameters that meet the preset efficiency threshold and have a safety risk lower than the critical value as the target signal light control parameters.
[0097] The reinforcement learning algorithm is an algorithm that continuously learns the optimal strategy through the interaction between the intelligent agent and the environment. In the embodiment of the present invention, it is used to simulate the execution effect of the initial control parameter set in the historical traffic scene. The historical traffic scene is the actual traffic situation in the past period of time. By simulating the execution of the initial control parameters in these scenes, its effect can be evaluated. The preset efficiency threshold is a preset minimum standard for measuring traffic efficiency, and the critical value is a preset maximum standard for measuring traffic safety risks. The target signal light control parameter is a control parameter selected from the initial control parameter set that meets the preset efficiency threshold and the safety risk is lower than the critical value. In the embodiment of the present invention, the reinforcement learning algorithm is used to simulate the execution effect of the initial control parameter set in the historical traffic scene. Load the pre-stored historical traffic scene data set, which contains the traffic flow state snapshots and the corresponding signal light control parameter execution records for multiple time periods. Each control parameter in the initial control parameter set is aligned and matched with the traffic flow state snapshot in time and space to generate a simulation task queue containing the signal light control parameter adjustment instruction and the corresponding scene identifier. In the virtual simulation environment of the reinforcement learning algorithm, the simulation task queue is executed one by one, the phase switching logic of the virtual signal light is dynamically modified according to the signal light control parameter adjustment instruction, and the modified vehicle traffic trajectory change data and pedestrian movement path offset are recorded. Based on the vehicle traffic trajectory change data, the traffic efficiency index corresponding to each control parameter is calculated, such as the average vehicle delay reduction rate, the green light utilization rate improvement and the queue length attenuation coefficient; the safety risk index is extracted according to the pedestrian movement path offset, such as the pedestrian crossing conflict incident rate, the emergency braking frequency and the pedestrian waiting time exceeding the limit ratio. The traffic efficiency index is compared with the preset efficiency threshold, and the first candidate parameter set with the traffic efficiency index exceeding the threshold is screened out, and the safety risk index is compared with the critical value, and the parameters with the safety risk index higher than the critical value are eliminated from the first candidate parameter set to generate the second candidate parameter set. If the second candidate parameter set is empty, the parameters in the initial control parameter set are optimized by gradient perturbation, a new control parameter set is generated, and the spatiotemporal alignment matching and virtual simulation process are re-executed until a non-empty second candidate parameter set is obtained. Finally, according to the weighted score ranking of the traffic efficiency index and the safety risk index of the parameters in the second candidate parameter set, the parameter with the highest score is selected as the target traffic light control parameter.
[0098] As an implementation manner, in step S320, the expected traffic efficiency gain and safety risk attenuation coefficient of each candidate control strategy are calculated, which may specifically include:
[0099] Step S321: Obtain the historical execution record corresponding to the candidate control strategy, and extract from the historical execution record the traffic flow state change sequence triggered by the candidate control strategy within the historical time window. The traffic flow state change sequence includes the vehicle average delay change rate, the cumulative value of the number of stops, and the pedestrian conflict event frequency.
[0100] The historical execution record refers to the actual execution record after adopting a certain candidate control strategy in past traffic scenarios, which contains various change information of the traffic flow state. The historical time window refers to a set time period used to analyze the execution effect of the candidate control strategy. The traffic flow state change sequence is a series of index sequences reflecting the traffic flow state changes extracted from the historical execution record. Among them, the vehicle average delay change rate refers to the change ratio of the vehicle average delay time after adopting the candidate control strategy; the cumulative value of the number of stops refers to the total number of vehicle stops during this time period; the pedestrian conflict event frequency refers to the frequency of pedestrian-vehicle conflict events during this time period.
[0101] In the embodiment of the present invention, the historical execution record corresponding to each candidate control strategy is obtained from the historical traffic dataset. For example, for a certain candidate control strategy, find the execution record of adopting this strategy under similar past traffic states. Extract the traffic flow state change sequence within the historical time window (such as 10 minutes) from this record. For example, the vehicle average delay change rate decreases from the original 20% to 15%, the cumulative value of the number of stops decreases from the original 50 times to 40 times, and the pedestrian conflict event frequency decreases from the original 1 time / minute to 0.5 times / minute.
[0102] Step S322: Perform time window segmentation processing on the traffic flow state change sequence to generate multiple consecutive sub-time period feature sets. Each sub-time period feature set includes the fluctuation amplitude of the vehicle average delay change rate, the growth slope of the cumulative value of the number of stops, and the distribution density of the pedestrian conflict event frequency.
[0103] The time window segmentation processing is a process of segmenting the traffic flow state change sequence according to the time interval. The multiple consecutive sub-time period feature sets generated are sets composed of the traffic flow state features within these segmented sub-time periods. The fluctuation amplitude of the vehicle average delay change rate refers to the difference between the maximum value and the minimum value of the vehicle average delay change rate within each sub-time period; the growth slope of the cumulative value of the number of stops refers to the change rate of the cumulative value of the number of stops with time within each sub-time period; the distribution density of the pedestrian conflict event frequency refers to the distribution of the pedestrian conflict event frequency within each sub-time period.
[0104] In an embodiment of the present invention, time window segmentation processing is performed on the traffic flow state change sequence. For example, a 10-minute historical time window is segmented by minute, resulting in 10 sub-time periods. For each sub-time period, the fluctuation amplitude of the average vehicle delay change rate, the growth slope of the cumulative parking count value, and the distribution density of the pedestrian conflict event frequency are calculated. For example, within the first sub-time period, the fluctuation amplitude of the average vehicle delay change rate is 2%, the growth slope of the cumulative parking count value is -1 time / minute, and the distribution density of the pedestrian conflict event frequency is 0.2 times / minute.
[0105] Step S323: Input the sub-time period feature set into the pre-trained traffic efficiency prediction model. Through the multi-scale temporal convolutional layer in the traffic efficiency prediction model, traffic efficiency correlation features with different time granularities are extracted, and traffic efficiency metrics after the simulated execution of the candidate control strategy are generated based on the traffic efficiency correlation features.
[0106] The pre-trained traffic efficiency prediction model is a model trained on a large amount of historical traffic data and is used to predict the traffic efficiency metrics after the simulated execution of the candidate control strategy. The multi-scale temporal convolutional layer is a convolutional layer in the traffic efficiency prediction model that can extract traffic efficiency correlation features with different time granularities, that is, analyze the relationship between the traffic flow state and traffic efficiency from different time scales. The traffic efficiency metric refers to a quantitative metric reflecting traffic efficiency, such as the reduction rate of average vehicle delay, the improvement amplitude of green light utilization rate, etc.
[0107] In an embodiment of the present invention, the sub-time period feature set generated in step S322 is input into the pre-trained traffic efficiency prediction model. The multi-scale temporal convolutional layer processes these feature sets and extracts traffic efficiency correlation features with different time granularities. For example, on a shorter time scale, it may be found that the sudden change in the average vehicle delay change rate is related to traffic efficiency; on a longer time scale, it may be found that the long-term trend of the cumulative parking count value affects traffic efficiency. Based on these traffic efficiency correlation features, the model generates traffic efficiency metrics after the simulated execution of the candidate control strategy, such as predicting that the reduction rate of average vehicle delay is 13% and the improvement amplitude of green light utilization rate is 8%.
[0108] Step S324: Synchronously input the sub-time period feature set into the safety risk assessment model. Through the event prediction network in the safety risk assessment model, the spatio-temporal correlation pattern between the pedestrian conflict event frequency and the sudden change in vehicle acceleration is identified, and the safety risk metrics corresponding to the candidate control strategy are output.
[0109] The safety risk assessment model is a model used to evaluate the safety risks of a traffic system. It predicts the safety risk indicators corresponding to candidate control strategies by analyzing various characteristics of the traffic flow state. The event prediction network is a network in the safety risk assessment model, which is used to identify the spatio-temporal correlation patterns between the frequency of pedestrian conflict events and the sudden change in vehicle acceleration, that is, to analyze at what time and spatial location there is a correlation between the frequency of pedestrian conflict events and the sudden change in vehicle acceleration. The safety risk indicator refers to a quantitative indicator reflecting the degree of safety risk of the traffic system, such as the incidence rate of pedestrian crossing conflict events, the emergency braking frequency, etc.
[0110] In the embodiment of the present invention, the sub-time period feature sets are synchronously input into the safety risk assessment model. The event prediction network will analyze these feature sets to identify the spatio-temporal correlation patterns between the frequency of pedestrian conflict events and the sudden change in vehicle acceleration. For example, it is found that when the vehicle acceleration suddenly increases, at certain set intersection locations, the frequency of pedestrian conflict events will increase. Based on this correlation pattern, the model outputs the safety risk indicators corresponding to the candidate control strategies, such as predicting that the incidence rate of pedestrian crossing conflict events is 0.4% and the emergency braking frequency is 1.5 times per minute.
[0111] Step S325: Calculate the expected traffic efficiency gain based on the difference between the traffic efficiency indicator and the historical benchmark traffic efficiency, and calculate the safety risk attenuation coefficient based on the deviation degree of the safety risk indicator relative to the preset safety threshold, where the expected traffic efficiency gain and the safety risk attenuation coefficient constitute the evaluation weight parameters of the candidate control strategy.
[0112] The historical benchmark traffic efficiency refers to the traffic efficiency indicator of the traffic system before the candidate control strategy is adopted. The expected traffic efficiency gain refers to the improvement amplitude of the traffic efficiency indicator relative to the historical benchmark traffic efficiency after the candidate control strategy is adopted, which is calculated by the difference between the two. The preset safety threshold is a preset standard value used to measure the safety risk of the traffic system. The safety risk attenuation coefficient refers to the reduction degree of the safety risk indicator relative to the preset safety threshold after the candidate control strategy is adopted, which is obtained by calculating the deviation degree of the safety risk indicator and the preset safety threshold. The evaluation weight parameter is a parameter used to evaluate the pros and cons of the candidate control strategy, and is composed of the expected traffic efficiency gain and the safety risk attenuation coefficient.
[0113] In an embodiment of the present invention, assume that the historical benchmark traffic efficiency is that the average vehicle delay reduction rate is 8%, and the preset safety threshold is that the occurrence rate of pedestrian crossing conflict events is 0.6%. For a certain candidate control strategy, its traffic efficiency index is that the average vehicle delay reduction rate is 13%, and the safety risk index is that the occurrence rate of pedestrian crossing conflict events is 0.4%. Then the expected traffic efficiency gain is 13% - 8% = 5%, and the safety risk attenuation coefficient is (0.6% - 0.4%) / 0.6% ≈ 33.3%. The expected traffic efficiency gain and the safety risk attenuation coefficient constitute the evaluation weight parameters of the candidate control strategy, which are used for subsequent evaluation and selection of the candidate control strategy.
[0114] As an implementation manner, in step S340, the execution effect of the initial control parameter set in the historical traffic scenario is simulated through a reinforcement learning algorithm, and the control parameter that meets the preset efficiency threshold and has a safety risk lower than the critical value is selected as the target signal light control parameter, which may specifically include:
[0115] Step S341: Load the pre-stored historical traffic scenario data set, which contains traffic flow state snapshots of multiple time periods and corresponding execution records of signal light control parameters.
[0116] The pre-stored historical traffic scenario data set is a data set that is pre-stored in the database and contains traffic flow state snapshots of multiple time periods and corresponding execution records of signal light control parameters. The traffic flow state snapshot refers to the state information of the traffic flow at a certain time point, such as the positions and speeds of vehicles, the number and positions of pedestrians, etc.; the execution record of the signal light control parameter refers to the actually adopted signal light control parameter during this time period, such as the green light duration, the yellow light transition interval, etc.
[0117] In an embodiment of the present invention, the pre-stored historical traffic scenario data set is loaded from the database. This data set can be obtained by collecting traffic data at the target intersection or other similar intersections for a long time. For example, the data set contains traffic flow state snapshots and corresponding execution records of signal light control parameters at different time periods every day in the past month, and these records can be used to simulate the execution effect of the initial control parameter set in the historical traffic scenario subsequently.
[0118] Step S342: Spatially and temporally align and match each control parameter in the initial control parameter set with the traffic flow state snapshot to generate a simulation task queue containing signal light control parameter adjustment instructions and corresponding scenario identifiers.
[0119] Spatiotemporal alignment matching refers to matching each control parameter in the initial control parameter set with the snapshot of the traffic flow state in time and space to ensure that the control parameters can be accurately applied to the corresponding traffic scenarios. The simulation task queue is a queue consisting of tasks containing signal light control parameter adjustment instructions and corresponding scene identifiers, which are used to execute in the virtual simulation environment of the reinforcement learning algorithm. Signal light control parameter adjustment instructions refer to instructions for adjusting the control parameters of signal lights, such as adjusting the green light duration from 30 seconds to 35 seconds; the corresponding scene identifier is information used to identify the specific traffic scenario where the control parameter is applied, such as time, location, etc.
[0120] In an embodiment of the present invention, each control parameter in the initial control parameter set is temporally and spatially aligned with the traffic flow state snapshot. For example, for a certain control parameter, a corresponding traffic flow state snapshot is found, which records the traffic flow state of the target intersection at a certain point in time. A simulation task queue containing a signal light control parameter adjustment instruction and a corresponding scene identifier is generated, such as a task in the task queue: in the traffic scene of time A and location B, adjust the green light duration of the signal light from 30 seconds to 35 seconds.
[0121] Step S343: Execute the simulation task queue one by one in the virtual simulation environment of the reinforcement learning algorithm, dynamically modify the phase switching logic of the virtual traffic light according to the traffic light control parameter adjustment instruction, and record the modified vehicle traffic trajectory change data and pedestrian movement path offset.
[0122] The virtual simulation environment of the reinforcement learning algorithm is a virtual environment that simulates real traffic scenes, in which different traffic light control parameters can be tested and evaluated. In this step, the tasks in the simulation task queue are executed one by one, and the phase switching logic of the virtual traffic light is dynamically modified according to the traffic light control parameter adjustment instruction. The phase switching logic of the virtual traffic light refers to the rules and time arrangements for the traffic light to switch between different colors, such as the duration of the green light on, the duration of the yellow light transition, etc. The modified vehicle traffic trajectory change data and pedestrian movement path offset are recorded. The vehicle traffic trajectory change data refers to the changes in the vehicle's driving trajectory after the traffic light control parameters are adjusted, such as changes in driving speed, driving direction, etc.; the pedestrian movement path offset refers to the deviation of the pedestrian's moving path relative to the original after the traffic light control parameters are adjusted.
[0123] In an embodiment of the present invention, a simulation task queue is executed in a virtual simulation environment of a reinforcement learning algorithm. For example, for a task, the green light duration of a virtual traffic light is adjusted from 30 seconds to 35 seconds. After the adjustment, the traffic trajectory change data of vehicles in the virtual environment is observed. For example, if a vehicle originally needs to stop and wait, after the green light duration is increased, the vehicle can directly pass through the intersection and the driving speed is also increased; at the same time, the pedestrian's moving path offset is recorded. For example, if a pedestrian originally needs to wait for a long time on the roadside, after the traffic light is adjusted, the pedestrian can pass through the crosswalk faster and the moving path offset is reduced.
[0124] Step S344: Calculate the traffic efficiency index corresponding to each control parameter based on the vehicle traffic trajectory change data. The traffic efficiency index includes the average vehicle delay reduction rate, the green light utilization rate improvement rate and the queue length attenuation coefficient.
[0125] The traffic efficiency index is a quantitative index used to measure the traffic efficiency. In the embodiment of the present invention, it includes the average vehicle delay reduction rate, the green light utilization rate improvement and the queue length attenuation coefficient. The average vehicle delay reduction rate refers to the reduction ratio of the average vehicle delay time relative to the original after adopting a certain control parameter; the green light utilization rate improvement refers to the improvement ratio of the utilization efficiency of the green light time relative to the original after adopting the control parameter; the queue length attenuation coefficient refers to the attenuation ratio of the vehicle queue length relative to the original after adopting the control parameter. When calculating these traffic efficiency indicators based on the vehicle traffic trajectory change data, it is first necessary to clarify that the average vehicle delay time refers to the time that the vehicle waits for passage at the intersection. By analyzing the vehicle's stay time at the intersection in the vehicle traffic trajectory change data, the average vehicle delay time before and after the control parameter is adopted is calculated, and then the average vehicle delay reduction rate is calculated. For example, before adopting a certain control parameter, the average vehicle delay time is 30 seconds, and after the adoption, the average delay time becomes 20 seconds, then the average vehicle delay reduction rate is (30-20) / 30\approx 33.3\%.
[0126] The green light utilization rate refers to the ratio of the number of vehicles that actually pass through the intersection during the green light time to the number of vehicles that can theoretically pass through during the green light time. The number of vehicles that actually pass through during the green light time and the theoretical number of vehicles that can pass through calculated based on traffic flow theory are counted through the vehicle traffic trajectory change data, and the green light utilization rate before and after the control parameters are adopted is compared to obtain the improvement of the green light utilization rate. For example, the green light utilization rate was 60% before the control parameters were adopted, and it became 70% after the adoption, so the improvement of the green light utilization rate is 70%-60%=10%.
[0127] The queue length refers to the length of the vehicle queue waiting to pass at an intersection. The queue lengths before and after adopting the control parameters are determined based on the vehicle passing trajectory change data, and the queue length attenuation coefficient is calculated. For example, if the queue length before adopting the control parameter is 50 meters and it becomes 30 meters after adoption, then the queue length attenuation coefficient is (50 - 30) / 50 = 40%.
[0128] Step S345: Synchronously extract safety risk indicators according to the pedestrian movement path offset. The safety risk indicators include the incidence rate of pedestrian crossing conflict events, the emergency braking frequency, and the proportion of pedestrians with excessive waiting time.
[0129] The pedestrian movement path offset reflects the change in the pedestrian movement path after the signal control parameter adjustment. Based on this, safety risk indicators are extracted. The incidence rate of pedestrian crossing conflict events refers to the ratio of the number of conflict events between pedestrians and vehicles to the total number of pedestrian crossings within a set time period. By analyzing the pedestrian movement path offset and the vehicle passing trajectory change data, it is judged whether there are conflict events between pedestrians and vehicles, the number of conflict events and the total number of pedestrian crossings are counted, and the incidence rate of pedestrian crossing conflict events is calculated. For example, within one hour, the total number of pedestrian crossings is 200 times, and there are 2 conflict events between pedestrians and vehicles, then the incidence rate of pedestrian crossing conflict events is 2 / 200 = 1%.
[0130] The emergency braking frequency refers to the number of times a vehicle makes an emergency brake due to pedestrians or other traffic conditions within a set time period. By analyzing the speed change in the vehicle passing trajectory change data and the pedestrian movement path offset, the emergency braking behavior of the vehicle is identified, the number of emergency brakes is counted, and the emergency braking frequency is obtained. For example, if a vehicle makes 10 emergency brakes within one hour, then the emergency braking frequency is 10 times / hour.
[0131] The proportion of pedestrians with excessive waiting time refers to the ratio of the number of pedestrians whose waiting time exceeds the preset waiting time to the total number of pedestrians crossing the street. The waiting time of pedestrians is determined according to the pedestrian movement path offset, compared with the preset waiting time, the number of pedestrians with excessive waiting time and the total number of pedestrians crossing the street are counted, and the proportion of pedestrians with excessive waiting time is calculated. For example, the preset pedestrian waiting time is 30 seconds. In a statistics, the total number of pedestrians crossing the street is 150, and among them, 30 have a waiting time exceeding 30 seconds. Then the proportion of pedestrians with excessive waiting time is 30 / 150 = 20%.
[0132] Step S346: Compare the traffic efficiency indicators with the preset efficiency threshold, screen out the first candidate parameter set whose traffic efficiency indicators exceed the threshold, and compare the safety risk indicators with the critical value, and eliminate the parameters in the first candidate parameter set whose safety risk indicators are higher than the critical value to generate the second candidate parameter set.
[0133] The preset efficiency threshold is a pre-set minimum standard for measuring traffic efficiency, and the critical value is a pre-set maximum standard for measuring safety risks. Compare the traffic efficiency index calculated in step S344 with the preset efficiency threshold. For example, the preset threshold for the reduction rate of average vehicle delay is 20%, the threshold for the improvement amplitude of green light utilization rate is 5%, and the threshold for the queue length decay coefficient is 30%. For the traffic efficiency index corresponding to each control parameter, if the reduction rate of average vehicle delay, the improvement amplitude of green light utilization rate, and the queue length decay coefficient all exceed the corresponding thresholds, then include this control parameter in the first candidate parameter set.
[0134] Next, compare the safety risk index extracted in step S345 with the critical value. Assume that the critical value for the incidence rate of pedestrian crossing conflict events is 1.5%, the critical value for the emergency braking frequency is 12 times per hour, and the critical value for the proportion of pedestrians waiting time exceeding the limit is 25%. Screen out the parameters in the first candidate parameter set whose safety risk index is lower than the critical value, and eliminate the parameters whose safety risk index is higher than the critical value to generate the second candidate parameter set. For example, there are 5 control parameters in the first candidate parameter set. After comparing the traffic efficiency indexes, 3 parameters meet the requirements and enter the first candidate parameter set; after comparing the safety risk indexes, the incidence rate of pedestrian crossing conflict events of one of the parameters is 2%, which is higher than the critical value of 1.5%, so it is eliminated, and finally a second candidate parameter set containing 2 parameters is obtained.
[0135] Step S347: If the second candidate parameter set is empty, perform gradient perturbation optimization on the parameters in the initial control parameter set, generate a new control parameter set and re-execute the space-time alignment matching and virtual simulation process until a non-empty second candidate parameter set is obtained.
[0136] If the second candidate parameter set is empty, it means that the parameters in the initial control parameter set cannot simultaneously meet the requirements of traffic efficiency and safety risks. At this time, perform gradient perturbation optimization on the parameters in the initial control parameter set. Gradient perturbation optimization is a gradient-based optimization method that makes small adjustments to the control parameters to make the control parameters change in a more optimal direction. For example, for the green light duration parameter, make a small increase or decrease according to the gradient information of the traffic efficiency and safety risk indexes.
[0137] After generating a new set of control parameters, re - execute the spatio - temporal alignment matching in step S342, match each control parameter in the new set of control parameters with the traffic flow state snapshot, and generate a new simulation task queue. Then, execute the new simulation task queue in the virtual simulation environment of the reinforcement learning algorithm, repeat the process of steps S343 - S346, that is, dynamically modify the phase - switching logic of the virtual traffic lights, record the vehicle passing trajectory change data and the pedestrian movement path offset, calculate the passing efficiency index and the safety risk index, and conduct comparison and screening until a non - empty second candidate parameter set is obtained. For example, after multiple gradient perturbation optimizations and re - simulations, a non - empty second candidate parameter set containing 3 parameters is finally obtained.
[0138] Step S348: According to the weighted score ranking of the passing efficiency index and the safety risk index of the parameters in the second candidate parameter set, select the parameter with the highest score as the target traffic light control parameter.
[0139] To comprehensively consider the passing efficiency and safety risk, conduct a weighted score for the passing efficiency index and the safety risk index of the parameters in the second candidate parameter set. First, set different weights for the passing efficiency index and the safety risk index. For example, the weight of the passing efficiency index is 0.6, and the weight of the safety risk index is 0.4. For the vehicle average delay reduction rate, the green - light utilization improvement rate, and the queue - length decay coefficient in the passing efficiency index, weights can also be set respectively. For example, the weight of the vehicle average delay reduction rate is 0.4, the weight of the green - light utilization improvement rate is 0.3, and the weight of the queue - length decay coefficient is 0.3; for the pedestrian - crossing conflict event occurrence rate, the emergency braking frequency, and the pedestrian waiting - time over - limit ratio in the safety risk index, weights are also set. For example, the weight of the pedestrian - crossing conflict event occurrence rate is 0.5, the weight of the emergency braking frequency is 0.3, and the weight of the pedestrian waiting - time over - limit ratio is 0.2.
[0140] Calculate the weighted score of each parameter. For example, for a certain parameter, the vehicle average delay reduction rate is 25%, the green - light utilization improvement rate is 8%, the queue - length decay coefficient is 35%, the pedestrian - crossing conflict event occurrence rate is 1%, the emergency braking frequency is 10 times per hour, and the pedestrian waiting - time over - limit ratio is 15%. Then the passing efficiency index score = 0.4×25%+0.3×8%+0.3×35% = 0.1 + 0.024+0.105 = 0.229; the safety risk index score = 0.5×(1 - 1%)+0.3×(1 - 10 / 12)+0.2×(1 - 15%) = 0.5×0.99+0.3×(1 - 0.833)+0.2×0.85 = 0.495+0.05+0.17 = 0.715; the total weighted score = 0.6×0.229+0.4×0.715 = 0.1374+0.286 = 0.4234.
[0141] Sort the parameters in the second candidate parameter set according to the weighted scores, and select the parameter with the highest score as the target signal light control parameter. For example, if there are 3 parameters in the second candidate parameter set, after weighted scoring, the score of parameter A is the highest, then parameter A is used as the target signal light control parameter.
[0142] Step S400: Send the target signal light control parameter to the traffic signal control terminal at the target intersection, and monitor the traffic state change data at the target intersection in real time to update the weight parameters of the spatio-temporal analysis model.
[0143] The traffic signal control terminal is a device installed at the target intersection for controlling signal lights. Send the target signal light control parameter generated in step S300 to this terminal, so that the signal lights perform phase switching according to the target signal light control parameter. For example, if the target signal light control parameter is that the green light duration is 40 seconds, the yellow light transition interval is 4 seconds, and the trigger condition for the pedestrian dedicated phase is that the number of pedestrians exceeds 12, after receiving these parameters, the traffic signal control terminal will adjust the control logic of the signal lights.
[0144] Monitor the traffic state change data at the target intersection in real time. These data reflect the actual traffic conditions at the target intersection after adopting the target signal light control parameter. By monitoring these data, update the weight parameters of the spatio-temporal analysis model to improve the prediction accuracy of the model for traffic states. For example, devices such as cameras and sensors installed at the intersection collect traffic videos, vehicle speeds, pedestrian numbers, etc. in real time, and use these data for subsequent analysis and model update.
[0145] As an implementation manner, in step S400, monitoring the traffic state change data at the target intersection in real time to update the weight parameters of the spatio-temporal analysis model may specifically include:
[0146] Step S410: After the target signal light control parameter is executed, continuously collect the feedback video stream at the target intersection, and extract the vehicle passing rate change feature, pedestrian waiting queue length feature and signal light switching delay time in the feedback video stream.
[0147] The feedback video stream refers to the traffic video stream continuously collected by the camera installed at the target intersection after the target signal light control parameter is executed. The vehicle passing rate change feature refers to the change of the vehicle passing rate at the intersection over time after adopting the target signal light control parameter. By analyzing the number and time of vehicles passing through the intersection in the feedback video stream, calculate the change of the vehicle passing rate. For example, within a certain period of time, the original vehicle passing rate was 80%, and after adopting the new control parameter, the vehicle passing rate became 85%, then the vehicle passing rate change feature reflects this change trend.
[0148] The pedestrian waiting queue length feature refers to the change in the length of the queue of pedestrians waiting in front of the crosswalk over time after the target signal control parameters are executed. By performing image processing and analysis on the feedback video stream, the boundaries of the pedestrian waiting queue are identified, the length of the queue is measured, and the queue lengths at different time points are statistically analyzed to obtain the pedestrian waiting queue length feature. For example, before the green light comes on, the length of the pedestrian waiting queue gradually increases, and after the green light comes on, the queue length gradually decreases.
[0149] The signal light switching delay time refers to the difference between the actual switching time and the theoretical switching time of the signal light. By analyzing the change time of the signal light color in the feedback video stream and comparing it with the theoretical switching time in the target signal control parameters, the signal light switching delay time is calculated. For example, if the target signal control parameters stipulate that the green light switches to the yellow light after 40 seconds, but it actually switches at 42 seconds, then the signal light switching delay time is 2 seconds.
[0150] Step S420: Perform a difference analysis on the vehicle passing rate change feature, the pedestrian waiting queue length feature, and the signal light switching delay time with the traffic state vector to generate a set of model error indicators.
[0151] The traffic state vector is generated in step S200 and is used to characterize the congestion level, vehicle passing priority, and pedestrian safety risk indicators of the target intersection. The vehicle passing rate change feature, the pedestrian waiting queue length feature, and the signal light switching delay time extracted in step S410 are subjected to a difference analysis with the traffic state vector to evaluate the prediction accuracy of the spatio-temporal analysis model.
[0152] Specifically, first, perform a time window alignment process on the vehicle passing rate change feature, the pedestrian waiting queue length feature, and the signal light switching delay time extracted from the feedback video stream to generate feedback time series data that is consistent with the prediction time range of the traffic state vector. For example, the prediction time range of the traffic state vector is one time point every 5 minutes. The vehicle passing rate change feature, the pedestrian waiting queue length feature, and the signal light switching delay time are also statistically analyzed and sorted every 5 minutes to obtain the feedback time series data.
[0153] Then, compare the vehicle passing rate change feature in the feedback time series data frame by frame with the vehicle passing priority weight in the traffic state vector to generate a vehicle passing difference vector. The vehicle passing difference vector includes the deviation amplitude between the actual passing rate and the predicted priority weight within each time window. For example, within a certain time window, the traffic state vector predicts that the passing rate corresponding to the vehicle passing priority weight is 80%, while the actual vehicle passing rate is 75%. Then the deviation amplitude within this time window is 80% - 75% = 5%.
[0154] Next, associate and map the pedestrian waiting queue length feature in the feedback time series data with the pedestrian safety risk index in the traffic state vector, and extract the abnormal fluctuation frequency and the proportion of the duration outside the corresponding threshold interval of the safety risk index for the pedestrian waiting queue length. For example, when the pedestrian safety risk index is within a certain threshold interval, the pedestrian waiting queue length should fluctuate within a set range. If it exceeds this range, it is considered an abnormal fluctuation, and the frequency and the proportion of the duration of the abnormal fluctuation are statistically analyzed. Then, perform a coupling analysis on the signal light switching delay time in the feedback time series data and the congestion level score in the traffic state vector to identify the cumulative delay deviation amount of the signal light switching delay time in different intervals of the congestion level score. For example, in the interval with a higher congestion level score, the signal light switching delay time may be longer, and the cumulative delay deviation amount in different intervals of the congestion level score is statistically analyzed. Finally, construct a three-dimensional error space based on the deviation amplitude, abnormal fluctuation frequency and the proportion of the duration, and the cumulative delay deviation amount in the vehicle passing difference vector, and calculate the Euclidean distance and covariance relationship of the error indicators in each dimension in the three-dimensional error space. Generate a set of model error indicators including vehicle passing consistency error, pedestrian safety prediction error and signal response delay error according to the Euclidean distance and covariance relationship. This set is used to quantify the local deviation and global offset trend of the spatio-temporal analysis model in real-time decision-making.
[0155] Step S421: Perform time window alignment processing on the vehicle passing rate change feature, pedestrian waiting queue length feature, and signal light switching delay time extracted from the feedback video stream to generate feedback time series data consistent with the prediction time range of the traffic state vector.
[0156] The time window alignment processing is to make the vehicle passing rate change feature, pedestrian waiting queue length feature, and signal light switching delay time extracted from the feedback video stream match the prediction time range of the traffic state vector. The traffic state vector is predicted at time intervals, for example, the traffic state is predicted every 5 minutes. Therefore, the feature data extracted from the feedback video stream also needs to be sorted and statistically analyzed at the same time interval.
[0157] For the vehicle passing rate change feature, count the number of vehicles passing through the intersection in each time window and calculate the vehicle passing rate. For example, within a 5-minute time window, the number of vehicles passing through the intersection is 100, and the theoretically passable number of vehicles in this time window is 120, then the vehicle passing rate is 100 / 120 ≈ 83.3%.
[0158] For the pedestrian waiting queue length feature, measure the length of the pedestrian waiting queue at the end of each time window. For example, at the end of a 5-minute time window, the length of the pedestrian waiting queue is 20 meters.
[0159] For the signal light switching delay time, record the difference between the actual switching time and the theoretical switching time of the signal light within each time window. For example, within a 5-minute time window, the signal light has a theoretical switching time at the 3rd minute, but it actually switches at 3 minutes and 20 seconds. Then the signal light switching delay time within this time window is 20 seconds.
[0160] Combine the data statistically counted according to time windows to generate feedback time series data that is consistent with the prediction time range of the traffic state vector for subsequent difference analysis.
[0161] Step S422: Compare the vehicle passing rate change characteristics in the feedback time series data with the vehicle passing priority weights in the traffic state vector frame by frame to generate a vehicle passing difference vector, where the vehicle passing difference vector includes the deviation amplitude of the actual passing rate from the predicted priority weight within each time window.
[0162] The vehicle passing priority weight is a value in the traffic state vector used to represent the passing priority of vehicles at intersections, and it has a certain relationship with the vehicle passing rate. By comparing the vehicle passing rate change characteristics in the feedback time series data with the vehicle passing priority weights in the traffic state vector frame by frame, the differences between the predicted vehicle passing situation by the model and the actual situation can be found.
[0163] Within each time window, compare the actual vehicle passing rate with the passing rate predicted based on the vehicle passing priority weight. For example, the vehicle passing priority weight in the traffic state vector indicates that the vehicle passing rate should be 85% within a certain time window, while the actual vehicle passing rate is 80%. Then the deviation amplitude within this time window is 85% - 80% = 5%. Record the deviation amplitude of each time window to form a vehicle passing difference vector. This vector reflects the prediction error of the vehicle passing situation in different time windows and helps analyze the accuracy of the model in vehicle passing prediction.
[0164] Step S423: Correlate and map the pedestrian waiting queue length characteristics in the feedback time series data with the pedestrian safety risk index in the traffic state vector, and extract the abnormal fluctuation frequency and the proportion of the duration outside the corresponding threshold interval of the pedestrian safety risk index for the pedestrian waiting queue length.
[0165] The pedestrian safety risk index is an indicator in the traffic state vector used to evaluate the safety risk level of pedestrians at intersections, and it has a relationship with the pedestrian waiting queue length. Generally speaking, when the pedestrian safety risk index is in different threshold intervals, the pedestrian waiting queue length should fluctuate within the corresponding range.
[0166] First, determine different threshold intervals of the pedestrian safety risk index and the reasonable ranges of the pedestrian waiting queue lengths corresponding to each interval. For example, when the pedestrian safety risk index is low, the pedestrian waiting queue length should be short; when the pedestrian safety risk index is high, the pedestrian waiting queue length may be long. Then, map the pedestrian waiting queue length features in the feedback time series data to these threshold intervals. For each time window, determine whether the pedestrian waiting queue length exceeds the reasonable range of the corresponding safety risk index threshold interval. If it exceeds, it is considered an abnormal fluctuation. Count the frequency of abnormal fluctuations, that is, the ratio of the number of time windows with abnormal fluctuations to the total number of time windows. At the same time, calculate the proportion of the duration of abnormal fluctuations, that is, the ratio of the total duration of abnormal fluctuations to the total time. For example, among 10 time windows, there are 2 time windows with abnormal fluctuations, the total duration of abnormal fluctuations is 20 minutes, and the total time is 100 minutes. Then the frequency of abnormal fluctuations is 2 / 10 = 20%, and the proportion of the duration of abnormal fluctuations is 20 / 100 = 20%. These indicators can reflect the accuracy of the model in predicting pedestrian safety risks.
[0167] Step S424: Perform a coupling analysis on the signal light switching delay time in the feedback time series data and the congestion level score in the traffic state vector to identify the cumulative delay deviation amount of the signal light switching delay time in different intervals of the congestion level score.
[0168] The congestion level score is an indicator in the traffic state vector used to describe the congestion degree of the target intersection. There is an association between the signal light switching delay time and the congestion level. Generally speaking, when the congestion level is high, the signal light switching delay time may be longer.
[0169] Divide the signal light switching delay time in the feedback time series data according to different intervals of the congestion level score. For example, divide the congestion level score into three intervals: mild congestion, moderate congestion, and severe congestion. For each interval, count the cumulative value of the signal light switching delay time and compare it with the theoretically signal light switching delay time in this interval to obtain the cumulative delay deviation amount.
[0170] For example, in the mild congestion interval, the cumulative value of the theoretically signal light switching delay time is 10 minutes, while the actually counted cumulative delay time is 12 minutes. Then the cumulative delay deviation amount in this interval is 12 - 10 = 2 minutes. Through this coupling analysis, it is possible to understand the difference between the actual situation of the signal light switching delay time under different congestion levels and the model prediction, which helps to discover the deficiencies of the model in signal light control.
[0171] Step S425: Construct a three-dimensional error space based on the deviation amplitude, abnormal fluctuation frequency and duration ratio, and cumulative delay deviation in the vehicle traffic difference vector, and calculate the Euclidean distance and covariance relationship of the error indicators in each dimension in the three-dimensional error space.
[0172] The three-dimensional error space is a space composed of three dimensions: the deviation amplitude, abnormal fluctuation frequency and duration ratio, and cumulative delay deviation in the vehicle traffic difference vector. In this space, each point represents the error situation within a time window or a statistical period.
[0173] The Euclidean distance is an indicator used to measure the distance between two points in the three-dimensional error space, which reflects the overall difference degree between different error situations. For two different error points (x_1, y_1, z_1) and (x_2, y_2, z_2), where x represents the deviation amplitude in the vehicle traffic difference vector, y represents the abnormal fluctuation frequency and duration ratio, and z represents the cumulative delay deviation, the Euclidean distance d = \sqrt{(x_2 - x_1)^2 + (y_2 - y_1)^2 + (z_2 - z_1)^2}. By calculating the Euclidean distance between different time windows or statistical periods, the change trend and stability of the error can be analyzed.
[0174] The covariance relationship is used to measure the correlation between the dimensions in the three-dimensional error space. For example, calculate the covariance between the deviation amplitude in the vehicle traffic difference vector and the abnormal fluctuation frequency and duration ratio. If the covariance is positive, it indicates a positive correlation between these two dimensions, that is, when the error in one dimension increases, the error in the other dimension may also increase; if the covariance is negative, it indicates a negative correlation; if the covariance is close to 0, it indicates a weak correlation between the two. By analyzing the covariance relationship, the mutual influence between the errors in each dimension can be understood, providing a basis for the update and optimization of the model.
[0175] Step S426: Generate a set of model error indicators including vehicle traffic consistency error, pedestrian safety prediction error, and signal response delay error according to the Euclidean distance and covariance relationship. The set of model error indicators is used to quantify the local deviation and global offset trend of the spatio-temporal analysis model in real-time decision-making.
[0176] The vehicle traffic consistency error reflects the difference degree between the vehicle traffic situation predicted by the model and the actual situation, which can be determined by the deviation amplitude in the vehicle traffic difference vector and the relevant analysis in the three-dimensional error space. For example, the vehicle traffic consistency error is comprehensively calculated according to the average deviation amplitude of the vehicle traffic difference vector and its covariance relationship with other dimensions.
[0177] The pedestrian safety prediction error reflects the accuracy of the model in predicting pedestrian safety risks, which is mainly calculated based on the abnormal fluctuation frequency, the proportion of the duration, and the relationship with other dimensions. For example, the higher the abnormal fluctuation frequency, the larger the proportion of the duration, and the stronger the correlation with other error dimensions, the greater the pedestrian safety prediction error. The signal response delay error reflects the prediction accuracy of the model in signal control, which is determined by the cumulative delay deviation and the analysis results in the three-dimensional error space. For example, the larger the cumulative delay deviation and the more complex the covariance relationship with other error dimensions, the greater the signal response delay error.
[0178] Combine the vehicle passing consistency error, the pedestrian safety prediction error, and the signal response delay error to generate a set of model error metrics. This set can quantify the local deviation of the spatio-temporal analysis model in real-time decision-making, such as the error situation in a certain time window or a certain traffic scenario; at the same time, it can also reflect the global deviation trend, such as the overall change trend of the model error over time. By analyzing the set of model error metrics, the weight parameters of the spatio-temporal analysis model can be updated targeted to improve the performance of the model.
[0179] Step S430: Adjust the attention weight allocation parameters and the bias terms of the fully connected layer in the spatio-temporal analysis model according to the set of model error metrics, and fine-tune the adjusted spatio-temporal analysis model using the incremental learning algorithm.
[0180] The attention weight allocation parameters are used to weight different features in the spatio-temporal analysis model to highlight important feature information. The bias terms of the fully connected layer are learnable parameters in the fully connected layer, which have a certain impact on the output results of the model. According to the set of model error metrics generated in step S420, adjust the attention weight allocation parameters and the bias terms of the fully connected layer in the spatio-temporal analysis model.
[0181] For example, if the set of model error metrics shows that the vehicle passing consistency error is large, it indicates that the model may be insufficient in processing vehicle passing-related features. At this time, the attention weight allocation parameters can be adjusted to increase the weight of vehicle passing-related features, so that the model pays more attention to these features. For the bias terms of the fully connected layer, they can be adjusted according to the direction and magnitude of the error to reduce the error.
[0182] Fine-tune the adjusted spatio-temporal analysis model using the incremental learning algorithm. In the embodiment of the present invention, the newly collected traffic video stream data and the set of model error metrics are used as inputs to train the adjusted spatio-temporal analysis model. Through incremental learning, the model can adapt to new traffic conditions without forgetting the knowledge learned before, and further improve the accuracy of the model.
[0183] Step S440: When any index in the model error index set exceeds the preset update threshold, trigger the full-parameter retraining process of the spatio-temporal analysis model, and generate updated model weight parameters based on the latest collected traffic video stream.
[0184] The preset update threshold is a pre-set standard for determining whether to perform full-parameter retraining on the spatio-temporal analysis model. If any of the vehicle passing consistency error, pedestrian safety prediction error, or signal response delay error in the model error index set exceeds the preset update threshold, it indicates that the performance of the model has deteriorated to the extent that retraining is required.
[0185] Trigger the full-parameter retraining process of the spatio-temporal analysis model, that is, start again from step S201, collect the data of the latest collected traffic video stream and the corresponding labeled traffic state tags, perform spatio-temporal slicing processing and data augmentation processing on the historical traffic video data set, construct an initial spatio-temporal analysis model, and perform iterative training on the model through a multi-task loss function, etc. During the retraining process, use the latest traffic data to update the weight parameters of the model to enable the model to better adapt to the new traffic situation.
[0186] For example, the preset update threshold for vehicle passing consistency error is 15%. When the vehicle passing consistency error in the model error index set reaches 18%, it exceeds the preset update threshold. At this time, trigger the full-parameter retraining process. Through retraining, generate updated model weight parameters to improve the accuracy and reliability of the spatio-temporal analysis model.
[0187] As an implementation manner, the method provided by the embodiment of the present application further includes the step of deploying a model lightweight module in an edge computing device, including:
[0188] Step S500: Perform channel pruning on the trained spatio-temporal analysis model, and remove the output channels in the convolutional layer whose absolute value of the weight is lower than the preset threshold.
[0189] Channel pruning is a model compression technology used to reduce the number of model parameters and computational complexity, and improve the running efficiency of the model. In the trained spatio-temporal analysis model, the convolutional layer contains multiple output channels, and each channel corresponds to a feature map. By analyzing the absolute value of the weight of each output channel in the convolutional layer, remove the output channels whose absolute value of the weight is lower than the preset threshold.
[0190] The preset threshold is a criterion set in advance for determining whether to remove a certain output channel. For example, if the preset threshold is set to 0.01, for a certain output channel in the convolutional layer, if the absolute value of its weight is less than 0.01, it is considered that the contribution of this channel to the model is small and it is removed. Through channel pruning, the number of model parameters and computational complexity can be reduced without significantly degrading the model performance, making the model more suitable for running on edge computing devices.
[0191] Step S600: Perform quantization-aware training on the pruned model, convert the floating-point weight parameters into a preset integer format with a certain number of bits, and optimize the rounding error during quantization.
[0192] Quantization-aware training is a method that considers quantization operations during model training, used to convert floating-point weight parameters into a preset integer format with a certain number of bits, while optimizing the rounding error during quantization. After channel pruning, quantization-aware training is performed on the pruned model.
[0193] The preset integer format with a certain number of bits refers to the number of integer bits into which the floating-point weight parameters are converted, such as an 8-bit integer format. During quantization-aware training, the training objective of the model is not only to minimize the prediction error but also to minimize the rounding error during quantization. By simulating quantization operations during training, the model can learn weight parameters that are more suitable for quantization. For example, using a quantization-aware training algorithm, in each training iteration, the floating-point weight parameters are quantized, the error after quantization is calculated, and it is incorporated into the loss function for optimization. In this way, the floating-point weight parameters are converted into a preset integer format with a certain number of bits, while reducing the quantization error and improving the quantization performance of the model.
[0194] Step S700: Convert the quantized model into an instruction set format supported by the hardware accelerator, and load the model parameters into the cache area of the edge computing device through memory mapping.
[0195] Different hardware accelerators support different instruction set formats. Converting the quantized model into an instruction set format supported by the hardware accelerator enables accelerated inference using the hardware accelerator on the edge computing device. Memory mapping is to map the content of a file or device to the address space of a process. Through memory mapping, the quantized model parameters are loaded into the cache area of the edge computing device, enabling the model to quickly access these parameters during inference and improving the inference efficiency. For example, convert the quantized model into the CUDA instruction set format supported by the GPU, and then load the model parameters into the GPU cache of the edge computing device through memory mapping. When performing model inference, the GPU can directly read the model parameters from the cache, accelerating the inference speed.
[0196] Step S800: Enable dynamic computational graph optimization during the model inference phase, and automatically skip redundant convolution operations according to the resolution of the real-time video stream.
[0197] Dynamic computational graph optimization is an optimization method that dynamically adjusts the computational graph according to the characteristics of the input data during the model inference phase. In edge computing devices, the resolution of the real-time video stream may change. According to the resolution of the real-time video stream, dynamic computational graph optimization can automatically skip redundant convolution operations, reducing unnecessary computational effort. For example, when the resolution of the real-time video stream is low, some convolution operations may have little impact on the output result of the model, and these convolution operations can be skipped. Through dynamic computational graph optimization, while ensuring the accuracy of model inference, the inference efficiency of the model is improved, and the computational burden on edge computing devices is reduced.
[0198] As an implementation, the deployment process of the model lightweight module further includes:
[0199] Step S900: Configure a dual-model operation mode in the edge computing device. The dual-model operation mode includes a high-precision mode and an energy-saving mode. The high-precision mode uses the unpruned spatio-temporal analysis model, and the energy-saving mode uses the pruned and quantized lightweight model.
[0200] To meet the requirements of different application scenarios, a dual-model operation mode is configured in the edge computing device. The high-precision mode uses the unpruned spatio-temporal analysis model, which has high accuracy but large computational effort and high energy consumption. The energy-saving mode uses the pruned and quantized lightweight model, which has fewer parameters and less computational effort, lower energy consumption, but the accuracy may decrease. For example, in the case of large traffic flow and high requirements for the accuracy of traffic state prediction, select the high-precision mode and use the unpruned spatio-temporal analysis model for traffic state analysis; in the case of small traffic flow and high requirements for energy consumption, select the energy-saving mode and use the pruned and quantized lightweight model for analysis.
[0201] Step S1000: Monitor the computational resource occupancy rate and the remaining battery power of the edge computing device in real time. When the resource occupancy rate exceeds the first threshold or the battery power is lower than the second threshold, automatically switch to the energy-saving mode.
[0202] Real-time monitor the computing resource occupancy rate of edge computing devices, such as CPU usage rate, GPU usage rate, etc., as well as the remaining battery power. The first threshold is a pre-set standard for judging whether the computing resource occupancy rate is too high, and the second threshold is a pre-set standard for judging whether the battery power is too low. When the computing resource occupancy rate exceeds the first threshold or the battery power is lower than the second threshold, it indicates that the computing resources of the edge computing device are tense or the battery power is insufficient. At this time, it automatically switches to the energy-saving mode. For example, set the first threshold to 80% and the second threshold to 20%. When the CPU usage rate of the edge computing device reaches 85%, exceeding the first threshold, or the battery power drops to 15%, lower than the second threshold, it automatically switches from the high-precision mode to the energy-saving mode and uses the lightweight model after pruning and quantization for calculation to reduce the occupancy of computing resources and energy consumption.
[0203] Step S1100: Start the model output corrector in the energy-saving mode, and use the statistical distribution of historical traffic state vectors to compensate for the deviation of the output of the lightweight model.
[0204] In the energy-saving mode, since the lightweight model after pruning and quantization is used, there may be a certain deviation in its output results. To improve the output accuracy of the lightweight model, start the model output corrector. The model output corrector uses the statistical distribution of historical traffic state vectors to compensate for the deviation of the output of the lightweight model. First, collect a large number of historical traffic state vectors and analyze their statistical distribution characteristics, such as mean, variance, etc. Then, calculate the deviation compensation value according to the output result of the lightweight model and the statistical distribution of historical traffic state vectors. For example, if the congestion level score output by the lightweight model is lower than the mean of the historical statistics distribution, adjust the congestion level score upward according to the deviation situation to compensate for the deviation. In this way, the output accuracy of the lightweight model in the energy-saving mode is improved.
[0205] Step S1200: When it is detected that the traffic scene complexity exceeds the processing capacity of the lightweight model, trigger the cloud collaborative inference mechanism, upload part of the video stream data to the cloud server, and fuse the output of the cloud model to generate the final decision parameters.
[0206] The traffic scene complexity refers to the number of traffic elements in the traffic scene, the complexity of the motion state, etc. When it is detected that the traffic scene complexity exceeds the processing capacity of the lightweight model, it indicates that the lightweight model may not be able to accurately analyze the current traffic state.
[0207] At this time, trigger the cloud collaborative inference mechanism and upload part of the video stream data to the cloud server. The cloud server has stronger computing power and more complex models and can handle more complex traffic scenes. The cloud server uses the cloud model to analyze the uploaded video stream data and outputs the analysis result of the cloud model.
[0208] Fuse the output results of the lightweight model and the output results of the cloud model to generate the final decision parameters. For example, the weighted average method can be used to assign different weights to the output results of the lightweight model and the cloud model according to their accuracy and reliability, and then perform weighted average to obtain the final decision parameters. Through the cloud collaborative inference mechanism, the processing ability and decision-making accuracy for complex traffic scenarios are improved while ensuring relatively low energy consumption of the edge computing device.
[0209] Please refer to Figure 3 , which is a schematic structural diagram of an edge computing traffic light intelligent decision-making system provided by an embodiment of the present application, including: a processor 101 and a memory 103. Among them, the processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the edge computing traffic light intelligent decision-making system 100 may further include a transceiver 104. It should be noted that in practical applications, the transceiver 104 is not limited to one, and the structure of the edge computing traffic light intelligent decision-making system 100 does not constitute a limitation to the embodiments of the present application. The memory 103 stores computer-readable code, and when the computer-readable code is run by the one or more processors 101, the one or more processors 101 execute the method provided by the embodiments of the present application.
Claims
1. An edge computing traffic light intelligent decision-making method integrating AI video analysis, characterized in that: The method comprises: Collecting a real-time traffic video stream of a target intersection through an edge computing device, and extracting a dynamic traffic feature set from the real-time traffic video stream, wherein the dynamic traffic feature set includes traffic flow distribution features, vehicle behavior trajectory features, and pedestrian movement trend features; Calling a pre-trained spatiotemporal analysis model to perform multimodal fusion processing on the dynamic traffic feature set to generate a traffic state vector of the target intersection, wherein the traffic state vector is used to characterize the congestion level, vehicle traffic priority, and pedestrian safety risk index of the target intersection; Matching a candidate control strategy in a preset decision rule library based on the traffic state vector, and adjusting parameters of the candidate control strategy through a strategy optimization model to generate target traffic light control parameters; The target traffic light control parameters are sent to the traffic signal control terminal of the target intersection, and the traffic state change data of the target intersection is monitored in real time to update the weight parameters of the spatiotemporal analysis model.
2. The method according to claim 1, characterized in that: The step of extracting a dynamic traffic feature set from the real-time traffic video stream comprises: Performing frame processing on the real-time traffic video stream to obtain a continuous video frame sequence, and performing multi-target detection on each video frame in the video frame sequence to obtain a vehicle position coordinate set, a pedestrian position coordinate set and traffic sign boundary frame data; The vehicle position coordinate set is temporally and spatially correlated by a trajectory tracking algorithm to generate a driving trajectory sequence of each vehicle, and the vehicle speed change curve, acceleration distribution characteristics and lane occupancy frequency are calculated according to the driving trajectory sequence; Performing group behavior analysis on the pedestrian position coordinate set to extract pedestrian movement direction consistency characteristics, pedestrian density fluctuation characteristics, and pedestrian crossing stay time; The vehicle speed change curve, acceleration distribution characteristics, lane occupancy frequency, pedestrian movement direction consistency characteristics, pedestrian density fluctuation characteristics and pedestrian crossing stay time are normalized to generate the dynamic traffic feature set.
3. The method according to claim 2, characterized in that The calling of the pre-trained spatiotemporal analysis model to perform multimodal fusion processing on the dynamic traffic feature set to generate the traffic state vector of the target intersection includes: The dynamic traffic feature set is input into the time feature extraction branch of the spatiotemporal analysis model, and the periodic fluctuation characteristics and sudden abnormal event characteristics of traffic flow in the time dimension are extracted through a time convolution network; The dynamic traffic feature set is input into the spatial feature extraction branch of the spatiotemporal analysis model, the target intersection is divided into multiple sub-regions through a regional segmentation network, and the vehicle density gradient characteristics, pedestrian gathering hotspot characteristics and traffic light visibility distribution of each sub-region are extracted respectively; The cross-modal attention mechanism is used to weight the periodic fluctuation characteristics, sudden abnormal event characteristics, vehicle density gradient characteristics, pedestrian gathering hotspot characteristics and traffic light visibility distribution in the time dimension to generate a fused global spatiotemporal feature map; The global spatiotemporal feature map is mapped to a preset traffic state space, and the traffic state vector including the congestion level score, vehicle traffic priority weight and pedestrian safety risk index is output through a fully connected layer.
4. The method according to claim 3, characterized in that The step of adjusting the parameters of the candidate control strategy through the strategy optimization model to generate target signal light control parameters includes: Selecting at least two candidate control strategies with the highest matching degree with the traffic state vector from the decision rule library, the candidate control strategies including a green light duration baseline parameter, a yellow light transition interval threshold, and a pedestrian-specific phase trigger condition; Inputting each of the candidate control strategies and the traffic state vector into a strategy evaluation module in the strategy optimization model, and calculating the expected traffic efficiency gain and safety risk attenuation coefficient of each candidate control strategy; Dynamically weighting each candidate control strategy according to the expected traffic efficiency gain and the safety risk attenuation coefficient to generate an initial control parameter set; The execution effect of the initial control parameter set in historical traffic scenarios is simulated by a reinforcement learning algorithm, and control parameters that meet a preset efficiency threshold and whose safety risk is lower than a critical value are selected as the target traffic light control parameters.
5. The method according to claim 4, characterized in that The real-time monitoring of the traffic status change data of the target intersection to update the weight parameters of the spatiotemporal analysis model includes: After the target signal light control parameters are executed, the feedback video stream of the target intersection is continuously collected, and the vehicle passing rate change characteristics, pedestrian waiting queue length characteristics and signal light switching delay time in the feedback video stream are extracted; Perform difference analysis on the vehicle passing rate variation characteristics, pedestrian waiting queue length characteristics and signal light switching delay time and the traffic state vector to generate a model error indicator set; Adjusting the attention weight allocation parameters and the bias item of the fully connected layer in the spatiotemporal analysis model according to the model error indicator set, and fine-tuning the adjusted spatiotemporal analysis model using an incremental learning algorithm; When any indicator in the model error indicator set exceeds a preset update threshold, the full parameter retraining process of the spatiotemporal analysis model is triggered, and updated model weight parameters are generated based on the latest collected traffic video stream.
6. The method according to claim 1, characterized in that The pre-training process of the spatiotemporal analysis model includes: Collect historical traffic video data sets and corresponding annotated traffic status labels, including manually annotated congestion levels, vehicle traffic priority coefficients, and pedestrian accident records; Performing spatiotemporal slicing processing on the historical traffic video data set to generate a training sample sequence, and performing data enhancement processing on each training sample in the training sample sequence, wherein the data enhancement processing includes simulating illumination changes, adding occluders, and perturbing the video frame sampling rate; Constructing an initial spatiotemporal analysis model, wherein the initial spatiotemporal analysis model includes a parallel temporal feature extraction network and a spatial feature extraction network, and a cross-modal fusion layer; Iteratively training the initial spatiotemporal analysis model through a multi-task loss function, wherein the multi-task loss function includes a weighted sum of a congestion level prediction error, a vehicle traffic priority classification error, and a pedestrian safety risk regression error; When the convergence speed of the multi-task loss function on the validation set is lower than a preset threshold, the parameters of the temporal feature extraction network are frozen and only the weights of the spatial feature extraction network are updated.
7. The method according to claim 6, characterized in that The training process of the strategy optimization model includes: Extracting multiple historical traffic state vectors and corresponding actual execution control parameters from the historical traffic video data set, and calculating the real traffic efficiency and safety index of each actual execution control parameter in a subsequent time window; Constructing a neural network structure of a policy optimization model, the neural network structure comprising a policy generator and a discriminator, the policy generator is used to generate control parameters according to an input traffic state vector, and the discriminator is used to evaluate the distribution consistency between the generated control parameters and the real control parameters; Alternately optimizing the strategy generator and the discriminator through an adversarial training algorithm until the discriminator cannot distinguish the difference between the generated control parameters and the real control parameters; The trained strategy generator is integrated with the preset traffic rule constraints to generate the strategy optimization model that can simultaneously meet the requirements of efficiency optimization and rule compliance; Wherein, the alternately optimizing the strategy generator and the discriminator through the adversarial training algorithm includes: Introducing a traffic rule penalty term into the loss function of the strategy generator, the traffic rule penalty term is used to perform gradient penalties on control parameters that violate the minimum green light duration, the maximum red light waiting time, and the pedestrian priority principle; Adding a feature matching constraint to the loss function of the discriminator so that the generated control parameters maintain statistical distribution consistency with the real control parameters in the hidden layer feature space; The weight parameters of the strategy generator and the discriminator are updated through a sliding average mechanism, and the generated control parameters are tested in a simulated environment after each iteration to screen effective samples; When the strategy generator reaches a preset traffic efficiency improvement rate in a simulation environment and the number of violations is lower than a tolerance threshold, the adversarial training process is terminated.
8. The method according to claim 1, characterized in that The method further comprises the step of deploying a model lightweight module in the edge computing device: Perform channel pruning on the trained spatiotemporal analysis model to remove output channels in the convolutional layer whose absolute weight values are lower than the preset threshold. Perform quantization-aware training on the pruned model, convert floating-point weight parameters to a preset bit integer format, and optimize rounding errors during quantization. Convert the quantized model into an instruction set format supported by the hardware accelerator, and load the model parameters into the cache area of the edge computing device through memory mapping; Dynamic computational graph optimization is enabled during the model inference phase to automatically skip redundant convolution operations based on the resolution of the real-time video stream.
9. The method according to claim 8, characterized in that The deployment process of the model lightweight module also includes: Configuring a dual-model operation mode in the edge computing device, the dual-model operation mode comprising a high-precision mode and an energy-saving mode, the high-precision mode using an unpruned spatiotemporal analysis model, and the energy-saving mode using a pruned and quantized lightweight model; Monitor the computing resource occupancy rate and battery remaining power of the edge computing device in real time, and automatically switch to energy-saving mode when the resource occupancy rate exceeds a first threshold or the battery power is lower than a second threshold; The model output corrector is started in energy-saving mode, and the output deviation of the lightweight model is compensated by using the statistical distribution of the historical traffic state vector; When the complexity of the monitored traffic scene exceeds the processing capacity of the lightweight model, the cloud-based collaborative reasoning mechanism is triggered, part of the video stream data is uploaded to the cloud server, and the output of the cloud model is integrated to generate the final decision parameters.
10. An edge computing traffic light intelligent decision-making system, characterized in that: include: one or more processors; and one or more memories, wherein the memories store computer readable codes, and when the computer readable codes are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Vehicle-road cooperative road traffic system
CN112289059A
Traffic network traffic light control method and system based on edge calculation, and computer readable storage medium
CN113096418A
Intelligent dynamic traffic light control method and system based on visual identification
CN117711191A
Holographic intersection signal real-time optimization method and device based on edge calculation
CN117854299A
Efficiency and safety combined multi-target signal control method in network connection mixed traveling scene
CN117994992A
Cited By
Robot joint module based on harmonic reducer
CN120461425A
Intelligent intersection adaptive lighting and traffic signaling system integrating vehicle-road cooperation and visual perception
CN120612830A
An intelligent intersection adaptive lighting and traffic signaling system fusing car-road cooperation and visual perception
CN120612830B
Traffic signal lamp intelligent control method and system, electronic equipment and storage medium
CN120636179A
Traffic dynamic cooperative control method, system and device and storage medium
CN120708396A