Multi-modal traffic flow prediction method and system based on deep learning
Through the multimodal traffic flow prediction method of deep learning, the traffic flow prediction model and the reinforcement learning model of dual Q network structure are used to construct Markov decision-making model and optimize the traffic light control strategy, solving the problems of time-consuming calculation and low efficiency in the processing of sudden events in the existing technology, and achieving efficient traffic light control.
Patent Information
- Application Number
- CN202510428544.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing traffic flow prediction methods require frequent refitting parameters, and the calculation time consumed with the increase of data, so it is impossible to efficiently handle traffic mutations caused by sudden events.
Using a multimodal traffic flow prediction method based on deep learning, the traffic flow prediction model and the reinforcement learning model of dual Q network structure is used to build a Markov decision model, optimize the traffic light regulation strategy, integrate the spatial flow network layer, motion flow network layer, timing prediction layer, multimodal fusion layer and full connection layer, and realize it through PyTorch and TensorFlow.
It realizes that there is no need to recalculate the model parameters, reduces calculation time, can efficiently handle sudden events, optimize traffic light control strategies, and improves intersection traffic efficiency.
Smart Images

Figure CN120299266A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation, and particularly relates to a multi-modal traffic flow prediction method and system based on deep learning. Background Art
[0002] Traffic flow prediction is a process of estimating the traffic volume of vehicles or pedestrians on roads within a specific future time period by analyzing historical and real-time data and combining mathematical models and artificial intelligence technologies. Its core lies in using the fusion of multi-modal vehicle driving acquisition data and algorithm optimization to improve the prediction accuracy, so as to provide a scientific basis for traffic management, route planning, etc.
[0003] In the prior art, traffic flow prediction is mainly achieved based on traditional time series analysis methods. However, this method relies on the sliding window calculation of historical data and requires frequent re-fitting of parameters. Therefore, the model parameters need to be recalculated for each prediction, and the calculation time-consuming increases with the growth of the data volume; it is unable to efficiently handle sudden traffic volume changes caused by emergencies (such as traffic accidents). Summary of the Invention
[0004] In order to overcome the defects that the traffic flow prediction in the prior art requires frequent re-fitting of parameters, the model parameters need to be recalculated for each prediction, the calculation time-consuming increases with the growth of the data volume; and it is unable to efficiently handle sudden traffic volume changes caused by emergencies, the present invention provides a multi-modal traffic flow prediction method based on deep learning, including:
[0005] Obtaining multi-modal vehicle driving acquisition data of a specific intersection;
[0006] Based on the multi-modal vehicle driving acquisition data, using a traffic flow prediction model for prediction to obtain traffic flow prediction data of the specific intersection;
[0007] Inputting the traffic flow prediction data into a reinforcement learning model with a double Q-network structure to construct a Markov decision model with the goal of maximizing the intersection passing efficiency;
[0008] Solving the Markov decision model to obtain a traffic signal control strategy for the specific intersection;
[0009] Wherein, the traffic flow prediction model integrates a spatial flow network layer, a motion flow network layer, a time series prediction layer, a multi-modal fusion layer, and a fully connected layer; the reinforcement learning model with a double Q-network structure is implemented through PyTorch and TensorFlow.
[0010] Further, the step of based on the multi-modal vehicle driving acquisition data, using a traffic flow prediction model for prediction to obtain traffic flow prediction data of the specific intersection includes:
[0011] Based on the video frame data in the multi-modal vehicle driving acquisition data, use the spatial flow network layer to process the data to obtain vehicle distribution features;
[0012] Based on the vehicle trajectory data in the multi-modal vehicle driving acquisition data, use the motion flow network layer to analyze the optical flow field for vehicle motion trend analysis to obtain trajectory pattern features;
[0013] Based on the time-series data of the sensors in the multi-modal vehicle driving acquisition data, use the time-series prediction layer to weight key time nodes through the attention mechanism to obtain sensor time-series features;
[0014] Based on the vehicle distribution features, the trajectory pattern features, and the sensor time-series features, use the multi-modal fusion layer to perform tensor splicing to obtain the spliced target features;
[0015] Based on the target features, use the fully connected layer to perform feature interaction to obtain traffic flow prediction data.
[0016] Further, the training process of the traffic flow prediction model includes:
[0017] Obtain the multi-modal vehicle driving sample data and traffic flow real data of the specific intersection;
[0018] Based on the multi-modal vehicle driving sample data, use the traffic flow prediction model to make a prediction to obtain traffic flow sample prediction data;
[0019] Based on the traffic flow real data and the traffic flow sample prediction data, use the mean squared error (MSE) loss combined with the KL divergence loss to perform weighted summation to obtain a loss value;
[0020] Based on the loss value, use the Adaptive Moment Estimation (Adam) optimizer to update the parameters of the traffic flow prediction model to obtain a trained traffic flow prediction model.
[0021] Further, the training process of the reinforcement learning model with a double Q-network structure includes:
[0022] Obtain the traffic flow real data of the specific intersection;
[0023] Input the traffic flow real data into the reinforcement learning model with a double Q-network structure to construct a Markov decision model with the goal of maximizing the intersection passing efficiency; solve the Markov decision model to obtain a traffic signal sample control strategy;
[0024] Based on the traffic signal sample regulation strategy, road simulation is carried out by using road simulation technology to obtain vehicle passing data;
[0025] Calculate the reward value of the vehicle passing data by using the reward function;
[0026] Based on the reward value, iterative training is performed on the reinforcement learning model to obtain a trained reinforcement learning model with a double Q-network structure.
[0027] Furthermore, the reward function satisfies the following formula:
[0028]
[0029] where r t is the reward value of the vehicle passing data, w1, w2, and w3 are weights, v i is the speed of the i-th vehicle passing through the stop line in the vehicle passing data, N is the total number of vehicles passing through the stop line in the vehicle passing data, q j is the queue length of the j-th lane in the vehicle passing data, is the phase switching conflict information of the traffic signal.
[0030] Furthermore, after solving the Markov decision model to obtain the traffic signal regulation strategy for the specific intersection, it further includes:
[0031] Based on the traffic flow prediction data, a preset traffic flow mutation threshold, and the traffic signal adjustment range, an optimization algorithm is used to adjust the signal cycle in the traffic signal regulation strategy to obtain an updated traffic signal adjustment strategy.
[0032] Furthermore, after solving the Markov decision model to obtain the traffic signal regulation strategy for the specific intersection, it further includes:
[0033] Use the traffic signal regulation strategy to adjust the specific intersection to obtain traffic flow adjustment data;
[0034] Based on the traffic flow adjustment data, a congestion detection algorithm is used to perform congestion detection to obtain a congestion detection result;
[0035] If the congestion detection result shows a congestion trend, based on the traffic flow adjustment, fuzzy control rules are used to optimize the traffic signal control strategy to obtain an optimized traffic signal control strategy.
[0036] On the other hand, the present invention also provides a multi-modal traffic flow prediction system based on deep learning, including:
[0037] A data acquisition module for obtaining multi-modal vehicle driving acquisition data of a specific intersection;
[0038] A traffic flow prediction model module for predicting based on the multi-modal vehicle driving acquisition data using a traffic flow prediction model to obtain traffic flow prediction data of the specific intersection;
[0039] A dynamic signal optimization module for inputting the traffic flow prediction data into a reinforcement learning model with a double Q-network structure to construct a Markov decision model aiming at maximizing the intersection passing efficiency; solving the Markov decision model to obtain a traffic signal control strategy for the specific intersection; wherein, the traffic flow prediction model integrates a spatial flow network layer, a motion flow network layer, a time series prediction layer, a multi-modal fusion layer, and a fully connected layer; the traffic flow prediction model and the reinforcement learning model with the double Q-network structure are implemented through PyTorch / TensorFlow.
[0040] Furthermore, the traffic flow prediction model module is specifically used for processing data using the spatial flow network layer based on the video frame data in the multi-modal vehicle driving acquisition data to obtain vehicle distribution features;
[0041] Based on the vehicle trajectory data in the multi-modal vehicle driving acquisition data, analyzing the optical flow field using the motion flow network layer for vehicle motion trend analysis to obtain trajectory pattern features;
[0042] Based on the time series data of the sensors in the multi-modal vehicle driving acquisition data, using the time series prediction layer to weight key time nodes through an attention mechanism to obtain sensor time series features;
[0043] Based on the vehicle distribution features, the trajectory pattern features, and the sensor time series features, using the multi-modal fusion layer for tensor splicing to obtain the spliced target features;
[0044] Based on the target features, using the fully connected layer for feature interaction to obtain traffic flow prediction data.
[0045] Furthermore, the traffic flow prediction model module is also used to train the traffic flow prediction model in the following way:
[0046] Obtain multi-modal vehicle driving sample data and traffic flow real data of the specific intersection;
[0047] Based on the multi-modal vehicle driving sample data, use the traffic flow prediction model for prediction to obtain traffic flow sample prediction data;
[0048] Based on the real traffic flow data and the predicted traffic flow sample data, the mean squared error (MSE) loss and the KL divergence loss are weighted and summed to obtain a loss value;
[0049] Based on the loss value, the parameters of the traffic flow prediction model are updated using the Adaptive Moment Estimation (Adam) optimizer to obtain a trained traffic flow prediction model.
[0050] Furthermore, the dynamic signal optimization module is also used to train the reinforcement learning model with a double Q-network structure in the following way:
[0051] Obtain the real traffic flow data of the specific intersection;
[0052] Input the real traffic flow data into the reinforcement learning model with a double Q-network structure to construct a Markov decision model aiming at maximizing the intersection passing efficiency; solve the Markov decision model to obtain a sample traffic signal control strategy;
[0053] Furthermore, the dynamic signal optimization module is also used to perform road simulation based on the sample traffic signal control strategy using road simulation technology to obtain vehicle passing data;
[0054] Calculate the reward value of the vehicle passing data using a reward function;
[0055] Iteratively train the reinforcement learning model based on the reward value to obtain a trained reinforcement learning model with a double Q-network structure.
[0056] Based on the predicted traffic flow data, a preset traffic flow mutation threshold, and a traffic signal adjustment range, use an optimization algorithm to adjust the signal cycle in the traffic signal control strategy to obtain an updated traffic signal adjustment strategy.
[0057] Furthermore, the dynamic signal optimization module is also used to adjust the specific intersection using the traffic signal control strategy to obtain traffic flow adjustment data;
[0058] Based on the traffic flow adjustment data, use a congestion detection algorithm to perform congestion detection to obtain a congestion detection result;
[0059] If the congestion detection result indicates a congestion trend, then optimize the traffic signal control strategy using fuzzy control rules based on the traffic flow adjustment to obtain an optimized traffic signal control strategy.
[0060] On the other hand, the present invention also provides a computer device, which is characterized by including: one or more processors;
[0061] The processor is used to store one or more programs;
[0062] When the one or more programs are executed by the one or more processors, the multi-modal traffic flow prediction method based on deep learning described in any one of the above is implemented.
[0063] On the other hand, the present invention also provides a computer-readable storage medium, which is characterized in that a computer program is stored thereon, and when the computer program is executed, the multi-modal traffic flow prediction method based on deep learning described in any one of the above is implemented.
[0064] Compared with the prior art, the beneficial effects of the present invention are:
[0065] The present invention provides a multi-modal traffic flow prediction method and system based on deep learning. The method includes:
[0066] An electronic device acquires multi-modal vehicle driving acquisition data of a specific intersection; uses a traffic flow prediction model for prediction to obtain traffic flow prediction data of the specific intersection; inputs the traffic flow prediction data into a reinforcement learning model with a double Q-network structure to obtain a traffic signal control strategy for the specific intersection. In the embodiment of the present invention, traffic flow prediction data is obtained by using a traffic flow prediction model, and a Markov decision model with the goal of maximizing the intersection passing efficiency is constructed by using a reinforcement learning model with a double Q-network structure, and the Markov decision model is solved to obtain a traffic signal control strategy for the specific intersection. This method not only realizes the determination of the traffic signal control strategy, but also does not require recalculating model parameters, reduces the calculation time-consuming, and can handle sudden events. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a schematic flowchart of the multi-modal traffic flow prediction method based on deep learning of the present invention;
[0068] Figure 2 is a schematic structural diagram of the multi-modal traffic flow prediction system based on deep learning of the present invention;
[0069] Figure 3 is a schematic structural diagram of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] The following further details the specific embodiments of the present invention with reference to the drawings.
[0071] Embodiment 1:
[0072] A multi-modal traffic flow prediction method based on deep learning provided by the present invention, the schematic flowchart is as Figure 1 shown, and includes:
[0073] Step 101: Obtain multi-modal vehicle driving acquisition data of a specific intersection;
[0074] Step 102: Based on the multi-modal vehicle driving acquisition data, use a traffic flow prediction model to make a prediction, and obtain traffic flow prediction data of the specific intersection;
[0075] Step 103: Input the traffic flow prediction data into a reinforcement learning model with a double Q-network structure, and construct a Markov decision model with the goal of maximizing the intersection passing efficiency;
[0076] Step 104: Solve the Markov decision model to obtain a traffic signal control strategy for the specific intersection; wherein, the traffic flow prediction model integrates a spatial flow network layer, a motion flow network layer, a time series prediction layer, a multi-modal fusion layer, and a fully connected layer; the reinforcement learning model with a double Q-network structure is implemented through PyTorch and TensorFlow.
[0077] A multi-modal traffic flow prediction method based on deep learning provided by an embodiment of the present invention is applied to an electronic device, and the electronic device can be an intelligent device such as a personal computer (PC), a server, etc.
[0078] In order to accurately and effectively perform multi-modal traffic flow prediction based on deep learning, in the embodiment of the present invention, traffic flow prediction data is obtained by using a traffic flow prediction model, and a Markov decision model with the goal of maximizing the intersection passing efficiency is constructed by using a reinforcement learning model with a double Q-network structure. The Markov decision model is solved to obtain a traffic signal control strategy for the specific intersection. This method not only realizes the determination of the traffic signal control strategy, but also does not need to recalculate the model parameters, reduces the calculation time-consuming, and can realize the processing of sudden events.
[0079] In order to determine the control strategy of traffic lights at a specific intersection, in an embodiment of the present invention, an electronic device can obtain multimodal vehicle driving data at a specific intersection. Specifically, video frame data can be collected by a camera array deployed at a specific intersection, and transmitted to the electronic device using H.265 encoding compression; the Internet of Things devices composed of geomagnetic sensors and radar sensors are used to collect vehicle flow, vehicle speed, and vehicle model data in real time, and the sampling frequency is set to 1Hz, wherein the vehicle flow, vehicle speed, and vehicle model data can be referred to as sensor time series data; the Global Positioning System (GPS) trajectory data of the networked vehicle is obtained through the vehicle to X (V2X) communication protocol, that is, the vehicle trajectory data, and the longitude and latitude coordinates, driving direction, and instantaneous speed are extracted. The video frame data, vehicle flow, vehicle speed, vehicle model data, longitude and latitude coordinates, driving direction, and instantaneous speed are the multimodal driving data described in the embodiment of the present invention.
[0080] After acquiring the multimodal driving data, the electronic device can use a timestamp synchronization mechanism to unify the multimodal driving data from the camera array, sensors and V2X communication to the same time base to ensure the time consistency of the data. Through coordinate conversion technology, the vehicle's GPS trajectory data is mapped to the grid coordinate system of the electronic map to achieve accurate alignment of the vehicle position and the spatial structure of the intersection. The electronic device also establishes a data cleaning pipeline to compensate for missing values in the multimodal driving data using linear interpolation to ensure data continuity. Based on statistical methods, the 3σ criterion (triple standard deviation criterion) is used to identify and eliminate outliers to avoid interference of noise data on subsequent analysis.
[0081] In order to realize multimodal traffic flow prediction, in an embodiment of the present invention, the electronic device locally stores a traffic flow prediction model, and the electronic device can input the multimodal vehicle driving collection data into the traffic flow prediction model to obtain the traffic flow prediction data of the specific intersection output by the traffic flow prediction model. Among them, the traffic flow prediction model integrates the spatial flow network layer, the motion flow network layer, the time series prediction layer, the multimodal fusion layer, and the fully connected layer. In one example, the spatial flow network layer can extract the spatial characteristics of the intersection and capture the static spatial distribution of the traffic flow. The motion flow network layer can capture the motion characteristics of the vehicle and analyze the dynamic changes of the traffic flow. The time series prediction layer can analyze the time series characteristics of the traffic flow and predict future traffic changes. The multimodal fusion layer can deeply fuse the characteristics of the spatial flow, motion flow and time series prediction. The fully connected layer can integrate all features and output the final traffic flow prediction data.
[0082] The electronic device also locally stores a reinforcement learning model with a double Q-network structure. After obtaining traffic flow prediction data, the traffic flow prediction data is input into the reinforcement learning model with the double Q-network structure, and the reinforcement learning model with the double Q-network structure constructs a Markov decision model with the goal of maximizing the intersection passing efficiency.
[0083] In one example, the state space of the Markov decision model can include the current traffic flow, signal light status, intersection topology, etc. The action space can be the signal light switching strategy, such as the duration of red, green, and yellow lights. The reward function can be aimed at maximizing the intersection passing efficiency, and the reward function can include vehicle waiting time, traffic volume, queue length, etc. The double Q-network structure can use two independent Q-networks (Q1 and Q2) to estimate the action value function, reducing the overestimation problem. Regularly update the target network parameters to stabilize the training process. Store the historical state-action-reward-next state tuples, randomly sample for training, break the data correlation, effectively avoid the training oscillation problem, and improve the stability of training.
[0084] In the embodiments of the present invention, the reinforcement learning model with the double Q-network structure can be implemented using PyTorch and TensorFlow.
[0085] The electronic device can solve the Markov decision model to obtain the traffic signal control strategy for the specific intersection. In one example, the electronic device can alternately perform policy evaluation and policy improvement to gradually optimize the policy. Directly optimize the value function until convergence.
[0086] In order to obtain traffic flow prediction data, based on the above embodiments, in the embodiments of the present invention, based on multi-modal vehicle driving acquisition data, a traffic flow prediction model is used for prediction, and the traffic flow prediction data for a specific intersection includes:
[0087] Based on the video frame data in the multi-modal vehicle driving acquisition data, a spatial flow network layer is used for data processing to obtain vehicle distribution features;
[0088] Based on the vehicle trajectory data in the multi-modal vehicle driving acquisition data, a motion flow network layer analyzes the optical flow field for vehicle motion trend analysis to obtain trajectory pattern features;
[0089] Based on the sensor time series data in the multi-modal vehicle driving acquisition data, a time series prediction layer weights key time nodes through an attention mechanism to obtain sensor time series features;
[0090] Based on the vehicle distribution features, trajectory pattern features, and sensor time series features, a multi-modal fusion layer performs tensor splicing to obtain the spliced target features;
[0091] Based on the target features, use the fully connected layer for feature interaction to obtain traffic flow prediction data.
[0092] The electronic device can extract the static spatial features of the intersection from the video frame data in the multi-modal vehicle driving acquisition data, and generate vehicle distribution features reflecting the vehicle density and position distribution. This spatial flow network layer can capture key information such as the lane distribution and vehicle positions at the intersection, providing basic spatial data support for traffic flow prediction.
[0093] Based on the vehicle trajectory data, the motion flow network layer analyzes the vehicle motion trajectories and speed changes through the optical flow method or the Long Short-Term Memory Network (LSTM), and generates trajectory pattern features reflecting the vehicle driving directions and motion patterns. This motion flow network layer can identify the vehicle motion trends, acceleration changes, and lane-changing behaviors, providing an important basis for traffic flow dynamic analysis.
[0094] Based on the time-series data collected by sensors (such as traffic flow, vehicle speed, vehicle type, etc.), the time-series prediction layer weights the key time nodes through the attention mechanism, extracts the periodic changes and sudden event features of the traffic flow, and generates high-precision sensor time-series features. This time-series prediction layer can capture periodic rules such as morning and evening rush hours, and effectively predict abnormal situations such as sudden traffic jams.
[0095] The multi-modal fusion layer deeply fuses the vehicle distribution features, trajectory pattern features, and sensor time-series features, and dynamically adjusts the weights of each modal feature through tensor splicing and the attention mechanism to generate unified high-dimensional target features. This multi-modal fusion layer can make full use of the complementarity of multi-modal data, improving the richness and robustness of feature expression.
[0096] Based on the fused target features, the fully connected layer performs feature interaction and non-linear mapping through a multi-layer neural network, and outputs traffic flow prediction data. This fully connected layer maps the high-dimensional features to specific traffic flow values, providing scientific and reliable data support for the formulation of signal light control strategies. Through the collaborative work of each layer, the traffic flow prediction model can achieve high-precision and real-time traffic flow prediction, providing an intelligent solution for urban traffic management.
[0097] In order to obtain traffic flow prediction data, on the basis of the above embodiments, in the embodiments of the present invention, the training process of the traffic flow prediction model includes:
[0098] Obtain multi-modal vehicle driving sample data and traffic flow real data of a specific intersection;
[0099] Based on multi-modal vehicle driving sample data, use a traffic flow prediction model to make predictions and obtain traffic flow sample prediction data;
[0100] Based on the real traffic flow data and the traffic flow sample prediction data, use the MSE loss combined with the KL divergence loss for weighted summation to obtain a loss value;
[0101] Based on the loss value, use the Adam optimizer to update the parameters of the traffic flow prediction model to obtain a trained traffic flow prediction model.
[0102] To train the traffic flow prediction model, an electronic device can obtain multi-modal vehicle driving sample data at a specific intersection and the corresponding real traffic flow data. The multi-modal vehicle driving sample data includes video frame data, vehicle trajectories, sensor time-series data, etc., while the real traffic flow data is obtained through intersection monitoring devices or manual annotation and is used for model training and verification.
[0103] Based on the multi-modal vehicle driving sample data, use the traffic flow prediction model to make predictions and generate traffic flow sample prediction data. This model extracts multi-modal features and outputs traffic flow sample prediction data through the collaborative work of a spatial flow network layer, a motion flow network layer, a time-series prediction layer, a multi-modal fusion layer, and a fully connected layer.
[0104] To evaluate the accuracy of the prediction results, based on the real traffic flow data and the traffic flow sample prediction data, combine the MSE loss and the KL divergence loss for weighted summation to calculate the loss value of the model. The MSE loss is used to measure the difference between the predicted value and the real value, while the KL divergence loss is used to evaluate the similarity between the predicted distribution and the real distribution. The combination of the two can effectively improve the prediction accuracy and robustness of the model.
[0105] In one example, the electronic device can calculate the loss value through the following multi-task loss function formula:
[0106]
[0107] where L total is the loss value, α and β are weight coefficients respectively, and the sum of these two weight coefficients is 1, y i is the real traffic flow in the i-th time period in the real traffic flow data, is the predicted traffic flow in the i-th time period in the traffic flow sample prediction data, N is the number of time periods, C is the number of traffic state categories, and the traffic state categories include unobstructed, congested, and slow-moving. P c is the predicted probability of the c-th traffic state category predicted in the traffic flow sample prediction data, and Q c is the real probability of the c-th traffic category in the real traffic flow data.
[0108] Among them, the loss function can jointly optimize the traffic flow numerical accuracy and distribution consistency, and improve the description ability of the traffic flow prediction model for complex traffic patterns. In one example, during training, the initial weights α can be set to 0.7 and β to 0.3, and a dynamic adjustment strategy is adopted to balance the two loss magnitudes.
[0109] Based on the calculated loss value, the Adam optimizer is used to update the parameters of the traffic flow prediction model. The Adam optimizer combines the advantages of the momentum method and the adaptive learning rate, and can converge quickly during training and avoid falling into local optima. Through multiple iterative trainings, a finally trained traffic flow prediction model is obtained, which can accurately predict future traffic flows and provide reliable data support for the optimization of signal control strategies.
[0110] In order to train the reinforcement learning model with a double Q-network structure and improve the reinforcement result of the reinforcement learning model, based on the above embodiments, in the embodiments of the present invention, the training process of the reinforcement learning model with a double Q-network structure includes:
[0111] Obtain the real traffic flow data of a specific intersection;
[0112] Input the real traffic flow data into the reinforcement learning model with a double Q-network structure to construct a Markov decision model with the goal of maximizing the intersection passing efficiency; solve the Markov decision model to obtain a sample traffic signal control strategy;
[0113] Based on the sample traffic signal control strategy, use road simulation technology to simulate the road and obtain vehicle passing data;
[0114] Calculate the reward value of the vehicle passing data using the reward function;
[0115] Based on the reward value, perform iterative training on the reinforcement learning model to obtain a trained reinforcement learning model with a double Q-network structure.
[0116] In the embodiments of the present invention, an electronic device can obtain the real traffic flow data of a specific intersection, and these data can include key indicators such as traffic volume, vehicle speed, and queue length, which are usually collected through intersection monitoring devices or historical traffic databases.
[0117] Input the real traffic flow data into the reinforcement learning model with a double Q-network structure to construct a Markov decision model with the goal of maximizing the intersection passing efficiency. The Markov decision model formalizes the traffic signal control problem as a sequential decision-making problem by defining the state space, action space, and reward function. The double Q-network structure reduces the overestimation problem through two independent Q-networks, improving the stability and convergence speed of the model.
[0118] Solve the Markov decision model to generate a sample traffic signal control strategy. The solution process uses policy iteration or value iteration methods, combined with deep reinforcement learning algorithms, to gradually optimize the signal control strategy to maximize the intersection throughput efficiency.
[0119] Based on the generated sample traffic signal control strategy, use road simulation technology to perform high-fidelity road simulation. Road simulation technology constructs a virtual traffic environment, simulates the driving behavior of vehicles under the signal control strategy, and generates vehicle traffic data, including the time for vehicles to pass through the intersection, queue length, delay time, etc.
[0120] Use a reward function to evaluate the vehicle traffic data generated by the simulation and calculate the reward value. The reward function usually takes the intersection throughput efficiency as the core goal and combines multiple metrics (such as vehicle waiting time, traffic volume, queue length) for weighted calculation to comprehensively reflect the effect of the signal control strategy.
[0121] Based on the calculated reward value, iteratively train the reinforcement learning model with a double Q-network structure. By continuously adjusting the model parameters and optimizing the signal control strategy, finally obtain a trained reinforcement learning model with a double Q-network structure. This reinforcement learning model with a double Q-network structure can determine an efficient and adaptive signal control strategy in the actual traffic environment, significantly improving the intersection throughput efficiency and traffic management level.
[0122] In order to train the reinforcement learning model with a double Q-network structure, based on the above embodiments, in the embodiments of the present invention, the reward function satisfies the following formula:
[0123]
[0124] where r t is the reward value of the vehicle traffic data, w1, w2, w3 are weights, v i is the speed of the i-th vehicle passing through the stop line in the vehicle traffic data, N is the total number of vehicles passing through the stop line in the vehicle traffic data, q j is the queue length of the j-th lane in the vehicle traffic data, is the phase switching conflict information of the traffic signal.
[0125] In one example, w1, w2, w3 can be 0.5, 0.3, 0.2 respectively.
[0126] In the embodiments of the present invention, the five-tuple of the Markov decision model can be represented as <S, A, P, R, γ>, and the state space S is s t = [f t , d t , q t , τt ,s t is the state vector for the future t-th time period, f t is the predicted traffic flow for the future t-th time period, d t is the current signal phase duration for the future t-th time period, q t is the queue length of each lane for the future t-th time period, τ t is the time period feature for the future t-th time period, specifically the classification code for morning / evening rush hours / flat peak periods. The action space A is a t ∈{maintain the current phase, switch to the next phase in advance, extend the green light duration by △ seconds}, where a t is the action taken in the future t-th time period, and △ is an adjustable parameter, which can be 5 - 15s.
[0127] In order to accurately and effectively adjust the traffic lights and improve the traffic efficiency at intersections, based on the above embodiments, in the embodiments of the present invention, after solving the Markov decision model to obtain the traffic light control strategy for a specific intersection, it further includes:
[0128] Based on the traffic flow prediction data, a preset traffic flow mutation threshold, and the traffic light adjustment range, use an optimization algorithm to adjust the signal cycle in the traffic light control strategy to obtain an updated traffic light adjustment strategy.
[0129] During the optimization process of the traffic light control strategy, the traffic lights of a specific intersection can be solved first through the Markov decision model to obtain a preliminary traffic light control strategy. Subsequently, based on the traffic flow prediction data, a preset traffic flow mutation threshold, and the traffic light adjustment range, use an optimization algorithm to dynamically adjust the signal cycle.
[0130] Specifically, the adjustment formula for the signal cycle can be the following formula:
[0131]
[0132] where C new is the adjusted signal cycle, C base is the basic signal cycle preset according to historical data, clip is a mathematical function that limits the calculation result, i.e., C new between C min and C max The mathematical function, is the traffic flow prediction data, f max is the maximum value of the traffic flow data. In one example, it can be the predicted traffic flow in the next 5 minutes, f threshold is the preset traffic flow mutation threshold (for example, the year-on-year change rate ≥ 15%), C min$C$ is the minimum signal period within the adjustment range of the traffic signal, which can be 60s in one example. max The maximum signal period within the adjustment range of the traffic signal can be 180s in one example.
[0133] Through this adjustment method, the signal period can be dynamically adjusted according to real-time traffic flow prediction and sudden changes, thereby optimizing the control strategy of the traffic signal and improving the traffic efficiency at the intersection.
[0134] In order to accurately and effectively adjust the traffic signal to improve the traffic efficiency at the intersection, based on the above embodiments, in the embodiments of the present invention, after solving the Markov decision model to obtain the traffic signal control strategy for a specific intersection, it further includes:
[0135] Adjust the specific intersection using the traffic signal control strategy to obtain traffic flow adjustment data;
[0136] Based on the traffic flow adjustment data, use a congestion detection algorithm to detect congestion and obtain a congestion detection result;
[0137] If the congestion detection result indicates a congestion trend, optimize the traffic signal control strategy using fuzzy control rules based on the traffic flow adjustment to obtain an optimized traffic signal control strategy.
[0138] During the optimization process of the traffic signal control strategy, the electronic device can solve the traffic signal for a specific intersection through the Markov decision model to obtain a preliminary traffic signal control strategy. The traffic flow of the specific intersection can be adjusted using this traffic signal control strategy, and based on the traffic flow adjustment data obtained after the adjustment, congestion detection is performed through a congestion detection algorithm to obtain a congestion detection result.
[0139] If the congestion detection result indicates a congestion trend, the electronic device can further optimize the traffic signal control strategy using preset fuzzy control rules. For example, when the queue length at the east (the east, south, west, and north described here are the actual east, south, west, and north in the application scenario, and the same applies hereinafter, and will not be repeated) entrance is long and the traffic flow at the west entrance is low, the electronic device can extend the green light time in the east-west direction by 10%. If the waiting time differences among the phases are large, the electronic device will switch to the maximum pressure phase. These rules are processed through Mamdani fuzzy inference and defuzzified using the centroid method.
[0140] The specific formula for the electronic device to determine that the congestion detection result indicates a congestion trend is:
[0141]
[0142] Among them, K is the number of time periods, which can be 6, and y k is the actual traffic flow in the k-th time period, is the predicted traffic flow in the k-th time period, and ∈ is a preset traffic flow deviation threshold. When the mode switching condition represented by this formula is satisfied, the traffic signal control strategy can be optimized through the above-mentioned fuzzy control rules.
[0143] Through these mechanisms, the electronic device can dynamically adjust the signal control strategy according to the real-time traffic conditions, thereby effectively alleviating traffic congestion and improving the traffic efficiency at intersections.
[0144] In the embodiments of the present invention, a three-dimensional traffic situation display interface of a Web Graphics Library (WebGL) can also be developed, integrating and overlaying the display of an electronic map, a video surveillance screen, and a predicted traffic flow heat map; constructing a visual decision-making panel: displaying the comparison of historical / predicted traffic flows through a line chart and showing the vehicle turning distribution using a Sankey diagram; implementing a signal timing plan simulation function: supporting the deduction of traffic flow changes after parameter adjustment and providing comparison views of multiple optimization schemes; deploying an abnormal warning subsystem: when it is detected that the regional congestion index exceeds the threshold, automatically generating a detour suggestion and pushing it to the navigation platform.
[0145] Embodiment 2:
[0146] Based on the same inventive concept, the present invention also provides a multi-modal traffic flow prediction system based on deep learning, the structural schematic diagram of which is as Figure 2 shown, including:
[0147] A data acquisition module 201, which is used to acquire multi-modal vehicle driving acquisition data of a specific intersection;
[0148] A traffic flow prediction model module 202, which is used to predict based on the multi-modal vehicle driving acquisition data by using a traffic flow prediction model to obtain traffic flow prediction data of a specific intersection;
[0149] A dynamic signal optimization module 203, which is used to input the traffic flow prediction data into a reinforcement learning model with a double Q-network structure, construct a Markov decision model with the goal of maximizing the intersection traffic efficiency; solve the Markov decision model to obtain the traffic signal control strategy of a specific intersection; among them, the traffic flow prediction model integrates a spatial flow network layer, a motion flow network layer, a time series prediction layer, a multi-modal fusion layer, and a fully connected layer; the traffic flow prediction model and the reinforcement learning model with a double Q-network structure are implemented through PyTorch / TensorFlow.
[0150] In a specific implementation manner, the traffic flow prediction model module 202 is specifically used for:
[0151] Based on the video frame data in the multi-modal vehicle driving acquisition data, use the spatial flow network layer to process the data to obtain vehicle distribution features;
[0152] Based on the vehicle trajectory data in the multi-modal vehicle driving acquisition data, use the motion flow network layer to analyze the optical flow field for vehicle motion trend analysis to obtain trajectory pattern features;
[0153] Based on the time-series data of the sensors in the multi-modal vehicle driving acquisition data, use the time-series prediction layer to weight the key time nodes through the attention mechanism to obtain sensor time-series features;
[0154] Based on the vehicle distribution features, trajectory pattern features and sensor time-series features, use the multi-modal fusion layer to perform tensor splicing to obtain the spliced target features;
[0155] Based on the target features, use the fully connected layer to perform feature interaction to obtain traffic flow prediction data.
[0156] In a specific implementation, the traffic flow prediction model module 202 is also used for:
[0157] Train the traffic flow prediction model in the following way:
[0158] Obtain the multi-modal vehicle driving sample data and traffic flow real data of a specific intersection;
[0159] Based on the multi-modal vehicle driving sample data, use the traffic flow prediction model to make predictions to obtain traffic flow sample prediction data;
[0160] Based on the traffic flow real data and traffic flow sample prediction data, use the mean square error MSE loss combined with the KL divergence loss to perform weighted summation to obtain a loss value;
[0161] Based on the loss value, use the Adam optimizer with adaptive moment estimation to update the parameters of the traffic flow prediction model to obtain the trained traffic flow prediction model.
[0162] In a specific implementation, the dynamic signal optimization module 203 is also used for:
[0163] Train the reinforcement learning model with a double Q network structure in the following way:
[0164] Obtain the traffic flow real data of a specific intersection;
[0165] Input the traffic flow real data into the reinforcement learning model with a double Q network structure to construct a Markov decision model with the goal of maximizing the intersection passing efficiency; solve the Markov decision model to obtain the traffic signal sample control strategy;
[0166] In a specific implementation, the dynamic signal optimization module 203 is further configured to:
[0167] Based on the traffic signal sample control strategy, use road simulation technology to perform road simulation to obtain vehicle passing data;
[0168] Use the reward function to calculate the reward value of the vehicle passing data;
[0169] Based on the reward value, perform iterative training on the reinforcement learning model to obtain a trained reinforcement learning model with a dual Q-network structure.
[0170] Based on the traffic flow prediction data, the preset flow mutation threshold, and the traffic signal adjustment range, use an optimization algorithm to adjust the signal cycle in the traffic signal control strategy to obtain an updated traffic signal adjustment strategy.
[0171] In a specific implementation, the dynamic signal optimization module 203 is further configured to:
[0172] Use the traffic signal control strategy to adjust a specific intersection to obtain traffic flow adjustment data;
[0173] Based on the traffic flow adjustment data, use a congestion detection algorithm to perform congestion detection to obtain a congestion detection result;
[0174] If the congestion detection result indicates a congestion trend, then based on the traffic flow adjustment, use fuzzy control rules to optimize the traffic signal control strategy to obtain an optimized traffic signal control strategy.
[0175] Embodiment 3:
[0176] As Figure 3 shown, the present invention further provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and the data can be called and / or modified when the instructions are executed.
[0177] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a multi-modal traffic flow prediction method based on deep learning in the above embodiments.
[0178] Embodiment 4:
[0179] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device, and of course can also include the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of a multi-modal traffic flow prediction method based on deep learning in the above embodiments can be implemented.
[0180] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0181] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0182] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that after reading the present invention, various changes, modifications, or equivalent replacements can still be made to the specific implementation manners of the application. However, these changes, modifications, or equivalent replacements are all within the scope of protection of the claims pending for approval of the application.
Claims
1. A multi-modal traffic flow prediction method based on deep learning, characterized in that, The method includes: Obtaining multi-modal vehicle driving acquisition data of a specific intersection; Based on the multi-modal vehicle driving acquisition data, using a traffic flow prediction model for prediction to obtain traffic flow prediction data of the specific intersection; Inputting the traffic flow prediction data into a reinforcement learning model with a double Q-network structure to construct a Markov decision model with the goal of maximizing the intersection passing efficiency; Solving the Markov decision model to obtain a traffic signal control strategy for the specific intersection; Wherein, the traffic flow prediction model integrates a spatial flow network layer, a motion flow network layer, a time series prediction layer, a multi-modal fusion layer, and a fully connected layer; the reinforcement learning model with a double Q-network structure is implemented through PyTorch and TensorFlow.
2. The method according to claim 1, characterized in that, The process of using the traffic flow prediction model for prediction based on the multi-modal vehicle driving acquisition data to obtain traffic flow prediction data of the specific intersection includes: Based on the video frame data in the multi-modal vehicle driving acquisition data, using the spatial flow network layer for data processing to obtain vehicle distribution features; Based on the vehicle trajectory data in the multi-modal vehicle driving acquisition data, using the motion flow network layer to analyze the optical flow field for vehicle motion trend analysis to obtain trajectory pattern features; Based on the time series data of the sensors in the multi-modal vehicle driving acquisition data, using the time series prediction layer to weight key time nodes through an attention mechanism to obtain sensor time series features; Based on the vehicle distribution features, the trajectory pattern features, and the sensor time series features, using the multi-modal fusion layer for tensor splicing to obtain the spliced target features; Based on the target features, using the fully connected layer for feature interaction to obtain traffic flow prediction data.
3. The method according to claim 1 or 2, characterized in that, The training process of the traffic flow prediction model includes: Obtaining multi-modal vehicle driving sample data and traffic flow real data of the specific intersection; Based on the multi-modal vehicle driving sample data, using the traffic flow prediction model for prediction to obtain traffic flow sample prediction data; Based on the traffic flow real data and the traffic flow sample prediction data, using the mean square error MSE loss combined with the KL divergence loss for weighted summation to obtain a loss value; Based on the loss value, using the Adam optimizer with adaptive moment estimation to update the parameters of the traffic flow prediction model to obtain a trained traffic flow prediction model.
4. The method according to claim 1, wherein The training process of the reinforcement learning model with a double Q-network structure includes: Obtaining traffic flow real data of the specific intersection; Inputting the traffic flow real data into the reinforcement learning model with a double Q-network structure to construct a Markov decision model with the goal of maximizing the intersection passing efficiency; solving the Markov decision model to obtain a traffic signal sample control strategy; Based on the traffic signal sample control strategy, using road simulation technology for road simulation to obtain vehicle passing data; Calculating a reward value of the vehicle passing data using a reward function; Iteratively train the reinforcement learning model based on the reward value to obtain a trained reinforcement learning model with a double Q-network structure.
5. The method according to claim 4, characterized in that, The reward function satisfies the following formula: where r t is the reward value of the vehicle passing data, w1, w2, w3 are weights, and v i is the speed of the i-th vehicle passing the stop line in the vehicle passing data, N is the total number of vehicles passing the stop line in the vehicle passing data, q j is the queue length of the j-th lane in the vehicle passing data, is the phase switching conflict information of the traffic signal.
6. The method according to claim 1, characterized in that After solving the Markov decision model to obtain the traffic signal control strategy for the specific intersection, the method further includes: Based on the traffic flow prediction data, a preset traffic flow mutation threshold, and a traffic signal adjustment range, use an optimization algorithm to adjust the signal cycle in the traffic signal control strategy to obtain an updated traffic signal adjustment strategy.
7. The method according to claim 1, wherein After solving the Markov decision model to obtain the traffic signal control strategy for the specific intersection, the method further includes: Use the traffic signal control strategy to adjust the specific intersection to obtain traffic flow adjustment data; Based on the traffic flow adjustment data, use a congestion detection algorithm to perform congestion detection to obtain a congestion detection result; If the congestion detection result indicates a congestion trend, based on the traffic flow adjustment, use fuzzy control rules to optimize the traffic signal control strategy to obtain an optimized traffic signal control strategy.
8. A multi-modal traffic flow prediction system based on deep learning, characterized in that, Including: A data acquisition module for acquiring multi-modal vehicle driving acquisition data of a specific intersection; A traffic flow prediction model module for predicting based on the multi-modal vehicle driving acquisition data using a traffic flow prediction model to obtain traffic flow prediction data for the specific intersection; A dynamic signal optimization module for inputting the traffic flow prediction data into a reinforcement learning model with a double Q-network structure, constructing a Markov decision model with the goal of maximizing the intersection passing efficiency; solving the Markov decision model to obtain the traffic signal control strategy for the specific intersection; wherein, the traffic flow prediction model integrates a spatial flow network layer, a motion flow network layer, a time series prediction layer, a multi-modal fusion layer, and a fully connected layer; the traffic flow prediction model and the reinforcement learning model with a double Q-network structure are implemented through PyTorch / TensorFlow.
9. An electronic device, characterized in that, Including: At least one processor and a memory; The memory and the processor are connected by a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the multi-modal traffic flow prediction method based on deep learning according to any one of claims 1-7 is implemented.
10. A readable storage medium, characterized in that, There is an execution program stored thereon, and when the execution program is executed, the multi-modal traffic flow prediction method based on deep learning according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Traffic signal lamp strategy evaluation evolution method and system
CN118015860A
Real-time traffic flow prediction method and system based on deep learning
CN118736853A
Adaptive traffic signal control method based on reinforcement learning and self-attention mechanism
CN118942261A
Traffic flow prediction and signal adjustment method based on machine learning
CN119516780A
Traffic information prediction device, traffic information prediction method, and traffic information prediction program
JP2024176431A
Cited By
Multi-mode traffic accident response decision-making method and system
CN120690027A