Dynamic traffic flow distribution method based on multi-agent reinforcement learning
By applying multi-agent reinforcement learning and cloud-edge collaborative architecture in traffic flow distribution, dynamically optimize and update vehicle driving routes, the problems of path convergence, response delay and secondary congestion in the existing technology are solved, and efficient and real-time traffic flow distribution effect is achieved.
Patent Information
- Application Number
- CN202510479699.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-06-20
AI Technical Summary
The existing traffic flow distribution method is difficult to cope with the dynamic and complex urban road network environment, resulting in path convergence, response delay and secondary congestion problems.
The dynamic traffic flow distribution method based on multi-agent reinforcement learning is adopted, and a multi-agent traffic simulation environment is built through the cloud-edge collaborative architecture, and a parallel strategy optimization is used to optimize the multi-agent reinforcement strategy network and near-end strategy optimization algorithm, a multi-objective reward mechanism is designed, and the vehicle driving route is dynamically updated through game theory Nash equilibrium solution.
Dynamic optimization and real-time response of large-scale vehicle paths are realized, the contradiction between local optimization and global efficiency is solved, the strategy bias caused by a single indicator is avoided, and minute-level evacuation is achieved in the event of sudden congestion.
Smart Images

Figure CN120183198A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent traffic control, and in particular to a dynamic traffic flow allocation method based on multi-agent reinforcement learning. Background Art
[0002] The dynamic allocation of traffic flow is the core of studying traffic network flow. Traditional traffic flow allocation methods mainly rely on static models and centralized control, and it is difficult to cope with the dynamic and complex urban road network environment. Existing navigation systems (such as Google Maps, Amap) use algorithms such as Dijkstra and A* for path planning based on historical data, which is prone to the path convergence effect during peak hours. And there are significant technical bottlenecks in centralized decision-making systems (such as SCOOT, SCATS): at a vehicle scale of 100,000 or more, the processing delay of the central node can reach more than 800 ms, and centralized data collection involves sensitive information such as vehicle trajectories.
[0003] Secondly, there are also application limitations of reinforcement learning algorithms in traffic control: the dimension of the Q-table at urban road network intersections can reach the order of 10^6, resulting in the curse of dimensionality; the convergence time of independent Q-learning in scenarios with 100+ agents increases by 4.8 times compared with collaborative methods. Existing hyper-heuristic methods generate routing strategies through genetic programming, but the strategies fail when the road network structure changes from grid-shaped to radial-shaped, and new strategies need to be retrained. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a dynamic traffic flow allocation method based on multi-agent reinforcement learning. The present invention realizes the dynamic optimization and real-time response of large-scale vehicle paths through a cloud-edge collaborative architecture.
[0005] The technical solution of the present invention is: a dynamic traffic flow allocation method based on multi-agent reinforcement learning, which realizes the dynamic optimization and real-time response of large-scale vehicle paths through a cloud-edge collaborative architecture, and includes the following steps:
[0006] S1), constructing a multi-agent traffic simulation environment by a cloud server based on OpenStreetMap road network data, and generating agent vehicles with random starting points according to the real-time traffic flow distribution and through a Poisson distribution model;
[0007] S2), constructing a multi-dimensional observation vector including road topology coding, current road density, average neighborhood speed, target distance ratio, congestion coefficient;
[0008] S3), constructing a multi-agent reinforcement policy network, and using the proximal policy optimization algorithm to perform parallel policy optimization on the multi-agent reinforcement policy network in the Ray RLlib framework;
[0009] S4), Design a multi-objective reward mechanism based on path efficiency, congestion penalty, and progress reward;
[0010] S5), Generate steering decisions based on real-time observation status, and dynamically update the vehicle driving route through the Nash equilibrium solution of game theory;
[0011] S6), Trigger a collaborative emergency mechanism when sudden congestion occurs.
[0012] Preferably, in step S2), the construction of the multi-dimensional observation vector specifically includes the following steps:
[0013] S21), Map the current road ID to a 128-dimensional binary vector through the road identification hashing algorithm, and normalize and compress it to a unique code in the [0,1] interval for road topology coding;
[0014] S22), Generate the current road density according to the ratio of the number of vehicles on the current road to the maximum capacity;
[0015] S23), Take the average speed of at least n adjacent vehicles within a radius of A meters as the neighborhood speed average;
[0016] S24), Calculate the target distance ratio according to the proportion of the shortest path distance from the current road to the target road to the maximum distance of the road network;
[0017] S25), Calculate the real-time congestion coefficient according to the number of parked and waiting vehicles on the current road.
[0018] Preferably, in step S3), the proximal policy optimization algorithm is used to perform parallel policy optimization of the multi-agent reinforcement policy network in the Ray RLlib framework, specifically including the following steps:
[0019] S31), Embed an LSTM network in the policy network to predict future congestion trends; and dynamically adjust the memory weights by inputting the current multi-dimensional observation vector and historical data into the LSTM network;
[0020] S32), Build a distributed cluster based on a parameter server in the Ray RLlib framework, and share the multi-agent policy parameters; and the central server stores the global policy network parameters, and the edge nodes synchronously update the local models;
[0021] S33), Implement the proximal policy optimization algorithm, use a dynamic KL threshold to control the policy update amplitude, and adaptively adjust the policy update step size; calculate the average reward, path conflict rate, and decision delay after each round of training, and dynamically adjust the exploration rate. When the average reward volatility exceeds 5%, increase the exploration rate by 0.1; if the path conflict rate is lower than 2% for three consecutive rounds of training, then reduce the exploration rate by 0.05.
[0022] Preferably, in step S4), the expression of the multi-objective reward mechanism is:
[0023] R total = 0.5R route + 0.3R progress + R congestion + R penalty
[0024] In the formula, R total is the comprehensive reward; R route is the estimated travel time difference between the paths before and after turning;
[0025] R progress is the arrival progress reward; R congestion is the real-time vehicle density change; R penalty is the invalid turning penalty.
[0026] Preferably, in step S4), after the vehicle executes a turning action, the estimated travel time difference between the paths before and after turning is obtained based on the real-time road condition interface. The expression of the estimated travel time difference R route between the paths before and after turning is:
[0027] R route = (T prev - T new ) / T prev × η;
[0028] In the formula, T prev is the travel time of the remaining road sequence of the original path obtained through the road network interface; T new is the travel time of the new path obtained through the real-time traffic flow prediction interface; η is the reward coefficient.
[0029] Preferably, in step S4), the real-time vehicle density change R congestion is calculated according to the real-time vehicle density change of the road entered after turning, and its expression is:
[0030] R congest i on = -(N new - N prev ) × μ;
[0031] In the formula, N new is the vehicle density of the newly entered road obtained through the dynamic perception interface; N prev is the vehicle density of the originally scheduled next road obtained through the road network topology interface; μ is the penalty coefficient.
[0032] Preferably, in step S4), the arrival progress reward R progress is calculated based on the remaining distance attenuation function, and its calculation formula is:
[0033]
[0034] Wherein, D remaing is the remaining path length; κ is the reward base number.
[0035] Preferably, in step S4), when the turning causes the path length to increase by more than a threshold value, the invalid turning penalty R is triggered penalty , and its expression is:
[0036] R penalty = -min(0, (L new - L orignal ) / L orignal ) × ω;
[0037] Wherein, L new is the total length of the new path; L orignal is the total length of the obtained original path; ω is the penalty coefficient.
[0038] Preferably, in step S5), a turning decision is generated according to the real-time observation state, specifically:
[0039] Preferably, in step S5), the vehicle driving route is dynamically updated through the Nash equilibrium solution of game theory, specifically:
[0040] S51), Obtain all legal turning edges of the current road and filter out unreachable paths;
[0041] S52), Generate a new path including the selected turning edge according to the action selection result;
[0042] S53), Detect invalid paths and trigger path reset operations, and use the Dijkstra algorithm to verify path reachability;
[0043] S54), Based on the multi-agent game decision model, calculate the Nash equilibrium solution of path selection, and through iterative strategy update, find the equilibrium solution that maximizes the benefits of all participants, implement global load balancing evaluation, and select the optimal diversion path.
[0044] Preferably, in step S6), when sudden congestion is detected, perform the following emergency operations:
[0045] S61), Perform dynamic path replanning on vehicles within a certain range;
[0046] S62), Set the priority lane permission to allow specific types of vehicles to pass first;
[0047] S63), Link the traffic signal control system to implement diversion induction and dynamically adjust the signal phase duration.
[0048] Preferably, in step S63), shunt induction is implemented, which specifically includes the following steps:
[0049] S631), constructing a road network load matrix based on multi-agent collaborative perception data; where each element represents the remaining traffic capacity of the corresponding road; where:
[0050] Remaining traffic capacity = (section maximum capacity - current number of vehicles) × average vehicle speed / section length;
[0051] Among them, the section maximum capacity is dynamically calculated according to the number of lanes, and the capacity threshold for each lane is set to 30 vehicles / km; the current number of vehicles is obtained through the fusion of distributed perception data of agent vehicles;
[0052] S632), collaborative path planning decision-making, and each agent performs the following operations:
[0053] a) Receive the real-time path planning decisions of other agents within a radius of 300 meters;
[0054] b) Construct a path selection benefit matrix based on the game decision model and calculate the Nash equilibrium solution;
[0055] c) Dynamically adjust the threshold according to the road network load matrix. When the regional vehicle density exceeds 30 vehicles / km, generate at least 3 sets of alternative paths that meet the remaining traffic capacity;
[0056] d) Select the optimal detour path that minimizes the global road network load variance;
[0057] S633), large-scale path dynamic synchronization, and execute the following distributed coordination mechanism:
[0058] a) Broadcast path change instructions through a virtual communication network to ensure the consistency of adjacent agent decisions;
[0059] b) Adopt a time window mechanism to update paths in batches, and the number of updated vehicles in each batch does not exceed 20% of the total;
[0060] c) Implement priority arbitration for conflicting path decisions and process them in ascending order according to the remaining distance of the vehicle to the target point.
[0061] Preferably, the cloud-edge collaborative architecture includes:
[0062] a) The cloud executes offline training of the multi-agent reinforcement policy network and generates global optimization model parameters through a distributed reinforcement learning framework;
[0063] b) The edge computing node deploys a lightweight policy network, receives the model parameters sent by the cloud, and executes real-time inference and decision-making;
[0064] c) A location-aware communication protocol based on a geographical hash table is adopted between edge nodes to achieve collaborative state awareness within a range of 500 meters in radius.
[0065] The beneficial effects of the present invention are as follows:
[0066] 1. The present invention resolves the contradiction between local optimization and global efficiency through LSTM and distributed game decision-making;
[0067] 2. The present invention avoids strategy bias caused by a single indicator by integrating path efficiency, congestion penalty, and progress reward; separates offline training and online inference, taking into account both computational efficiency and model accuracy;
[0068] 3. Based on the remaining traffic capacity assessment and collaborative path arbitration, the present invention realizes the evacuation of sudden congestion within minutes; solves the problems of path convergence, response delay, and secondary congestion in traditional methods. Description of the Drawings
[0069] Figure 1 is a schematic flow diagram of the method of the present invention;
[0070] Figure 2 is a structural diagram of the policy network of the method of the present invention. Detailed Embodiments
[0071] The following further describes the detailed embodiments of the present invention with reference to the drawings:
[0072] As Figure 1 shown, this embodiment provides a dynamic traffic flow allocation method based on multi-agent reinforcement learning, which realizes the dynamic optimization and real-time response of large-scale vehicle paths through a cloud-edge collaborative architecture, including the following steps:
[0073] S1). Construct a multi-agent traffic simulation environment
[0074] The cloud server generates a topological structure based on the OpenStreetMap road network data; a generator is set according to the real-time traffic flow distribution, and agent vehicles with random starting points are generated through the Poisson distribution model; in this embodiment, the vehicle observation radius is set to 200 meters, supporting a steering angle limit of ±45 degrees.
[0075] S2). Construct a multi-dimensional observation vector including road topology coding, current road density, average neighborhood speed, target distance ratio, and congestion coefficient; specifically including the following steps:
[0076] S21). Map the current road ID to a 128-dimensional binary vector through a road identification hash algorithm, and normalize and compress it into a unique code in the [0, 1] interval for road topology coding;
[0077] S22), generate the current road density according to the ratio of the current number of road vehicles to the maximum capacity;
[0078] S23), take the average speed of at least 5 adjacent vehicles within a radius of 50 meters as the neighborhood speed average;
[0079] S24), calculate the target distance ratio according to the proportion of the shortest path distance from the current road to the target road to the maximum distance of the road network;
[0080] S25), calculate the real-time congestion coefficient congestion according to the number of parked and waiting vehicles on the current road; that is:
[0081] congestion = α * n stop + β * t wait
[0082] In the formula, n stop is the number of stationary vehicles; t wait is the average waiting time; α, β are adjustment factors.
[0083] S3), construct a multi-agent reinforcement policy network, such as Figure 2 , and use the proximal policy optimization algorithm to perform parallel policy optimization on the multi-agent reinforcement policy network under the Ray RLlib framework; specifically including the following steps:
[0084] S31), embed a 64-dimensional LSTM network in the policy network to predict the future congestion trend; and form a time series matrix by splicing the current multi-dimensional observation vector and historical data and input it into the LSTM network to dynamically adjust the memory weights;
[0085] h t = LSTM([ot, h t-1 )
[0086] where ot is the current observation vector; h t-1 is the historical hidden state; h t represents the decision feature;
[0087] Memory update: Automatically weaken irrelevant old memories and focus on strengthening sudden congestion features;
[0088] Prediction output: The LSTM network outputs a hidden state carrying spatio-temporal features, which is converted into a future road density value through the prediction layer, and then dynamically adjusted according to the prediction result. In this embodiment, the attenuation rate of the road traffic capacity within the next 3 minutes is predicted. For example, if the current road density is 0.7, the path avoidance strategy is triggered when the predicted value > 0.85. Concatenate the hidden state vector output by the LSTM with the real-time observation to form a 128-dimensional decision feature and input it into the fully connected layer of the policy network.
[0089] S32). Build a distributed cluster based on the parameter server under the Ray RLlib framework to share multi-agent policy parameters; the central server stores the global policy network parameters, and the edge nodes synchronously update the local models; adopt a dynamic batch scheduling strategy to balance the computing load. When a node fails, automatically migrate the unfinished tasks to the standby node and restore the training snapshot of the last 10 minutes.
[0090] S33). Implement the Proximal Policy Optimization (PPO) algorithm, adopt a dynamic KL divergence threshold control strategy to update the amplitude, set the KL divergence threshold to 0.02; adaptively adjust the policy update step size; calculate the average reward, path conflict rate, and decision delay after each round of training, and dynamically adjust the exploration rate. When the average reward volatility exceeds 5%, increase the exploration rate by 0.1; if the path conflict rate is lower than 2% for three consecutive rounds of training, then decrease the exploration rate by 0.05.
[0091] In this embodiment, the implementation details of the Proximal Policy Optimization algorithm are as follows under the Ray RLlib framework. The parallel optimization of the multi-agent policy network is achieved through the following steps:
[0092] S341). Experience sampling and storage: Each intelligent vehicle agent interacts with the traffic environment according to the current policy and records the following data in real time:
[0093] Observation state: road topology encoding, current road density, average neighborhood speed, target distance ratio, congestion coefficient;
[0094] Action selection: steering decision, such as turning left / go straight / turning right;
[0095] Immediate reward: calculated according to the multi-objective reward mechanism formula in step S4);
[0096] Prediction result: the decay rate of the road density in the next 3 minutes output by the LSTM;
[0097] Every time the edge node collects 512 pieces of experience data, it uploads them to the cloud parameter server. In case of node failure, it automatically switches to the standby node and restores the latest data snapshot within 10 minutes;
[0098] S342). Policy evaluation and advantage calculation
[0099] Value network evaluation: Calculate the value score of the current state through the value network to measure the quality of the policy.
[0100] Advantage value calculation: Combine the current reward and the value of the next state, and dynamically adjust the discount factor and the smoothing parameter to reduce the estimation bias. In this embodiment, the dynamically adjusted discount factor is set to 0.99; the smoothing parameter is set to 0.95.
[0101] S343). Policy optimization and update
[0102] Policy difference control: Compare the action selection probabilities of the new and old policies. If the difference exceeds the preset threshold, automatically reduce the parameter update amplitude to ensure training stability.
[0103] Loss function optimization, synchronously optimizing the following three items:
[0104] Policy loss: Prioritize actions with high advantage values, but limit the update amplitude to no more than ±20%;
[0105] Value loss: Reduce the prediction error of the value network;
[0106] Exploration incentive: Add a regularization term with a weight of 0.01 to the policy entropy to avoid premature convergence;
[0107] S344) Exploration rate adjustment:
[0108] When the average reward fluctuation exceeds 5%, increase the exploration rate by 0.1 to expand the search range;
[0109] If the path conflict rate is lower than 2% for three consecutive rounds, reduce the exploration rate by 0.05 to stabilize the policy;
[0110] Emergency response: When sudden congestion occurs, freeze the update of the regular policy and prioritize the execution of dynamic path replanning;
[0111] S345) Multi-agent collaborative decision-making:
[0112] Revenue matrix construction: Each vehicle calculates the expected revenue of different steering actions based on the real-time road conditions;
[0113] Nash equilibrium solution: Through iterative negotiation, each vehicle adjusts its path selection until no vehicle can obtain higher revenue by unilaterally changing its decision;
[0114] Global optimization: Finally, select the flow diversion plan that makes the road network load most balanced.
[0115] S4) Design a multi-objective reward mechanism based on path efficiency, congestion penalty, and progress reward, generate decision signals based on the multi-objective reward mechanism, and complete policy gradient update; specifically:
[0116] S41) Decision signal generation mechanism: Convert the multi-objective reward into a decision signal through the following steps:
[0117] Real-time reward calculation: After the vehicle executes a steering action, immediately obtain the following data through the road network interface:
[0118] The estimated travel time difference R of the path before and after steering route ;
[0119] The change in the real-time vehicle density R of the newly entered road congestion ;
[0120] Remaining path length;
[0121] Calculate the comprehensive reward of the multi-objective reward mechanism as:
[0122] R total = 0.5R route + 0.3R progress + R congestion + R penalty
[0123] In the formula, R total is the comprehensive reward; R route is the estimated travel time difference between the paths before and after turning;
[0124] R progress is the arrival progress reward; R congestion is the real-time vehicle density change; R penalty is the invalid turn penalty.
[0125] Reward normalization processing: Perform dynamic normalization on R total so that it falls within the interval [-1, 1]:
[0126] R normalized = (R totak - μ history ) / σ history ;
[0127] Among them, μ jistory and σ history are the moving average and standard deviation of the rewards in the last 100 times respectively;
[0128] R normalized is the normalized reward.
[0129] Decision signal output: Concatenate the normalized reward R normalized with the predicted future congestion trend by LSTM to form the final decision signal vector: [R normalized , congestion pred , action prod ;
[0130] Among them, congestion pred represents the prediction of the future congestion trend of the road; action prod is the action probability distribution output by the policy network.
[0131] S42), The policy gradient update process is based on the gradient update of the multi-objective reward and is executed according to the following steps:
[0132] Experience replay:
[0133] Store vehicle interaction data in the experience pool and trigger batch update when the number of experiences reaches 512;
[0134] Gradient calculation: Use the PPO algorithm to calculate the policy gradient:
[0135] Dynamically adjust the reward weight after each round of training;
[0136] If the path conflict rate > 10%, increase the R route weight by 0.05;
[0137] If the average congestion coefficient > 0.7, increase the R congestion weight by 0.03;
[0138] Parameter synchronization:
[0139] After the cloud completes the gradient update, send the latest policy parameters to the edge nodes to ensure real-time decision consistency;
[0140] After the vehicle executes a steering action, obtain the estimated travel time difference between the paths before and after steering based on the real-time traffic condition interface. The estimated travel time difference R route of the paths before and after steering is expressed as:
[0141] R route =(T prev -T new ) / T prev ×η;
[0142] In the formula, T prev is the travel time of the remaining road sequence of the original path obtained through the road network interface; T new is the travel time of the new path obtained through the real-time traffic flow prediction interface; η is the reward coefficient.
[0143] The described real-time vehicle density change R congestion is calculated based on the real-time vehicle density change of the road entered after steering, and its expression is:
[0144] R congestion =-(N new -N prev )×μ;
[0145] In the formula, N new is the vehicle density of the newly entered road obtained through the dynamic perception interface; N prev is the vehicle density of the originally scheduled next road obtained through the road network topology interface; μ is the penalty coefficient.
[0146] The described arrival progress reward R progress is calculated based on the remaining distance attenuation function, and its calculation formula is:
[0147]
[0148] In the formula, D remaing is the remaining path length; κ is the reward base number.
[0149] When the turning causes the path length to increase by more than the threshold, the invalid turning penalty R is triggered penalty , and its expression is:
[0150] R penalty = -min(0, (L new - L orignal ) / L orignal ) × ω;
[0151] In the formula, L new is the total length of the new path; L orignal is the total length of the original path obtained; ω is the penalty coefficient.
[0152] S5), Generate a turning decision according to the real-time observation state, and dynamically update the vehicle driving route through the Nash equilibrium solution of game theory; specifically including the following steps:
[0153] S51), Obtain all legal turning edges of the current road, filter out unreachable paths, for example; filter out roads closed due to construction or accidents.
[0154] S52), Generate a new path including the selected turning edge according to the action selection result; for example, if the current road is a three-lane intersection, only keep the valid options of "left turn", "go straight", and "right turn",
[0155] S53), Detect invalid paths and trigger path reset operations, and use the Dijkstra algorithm to verify path reachability;
[0156] S54), Based on the multi-agent game decision model, calculate the Nash equilibrium solution of path selection, and through iterative strategy updates, find the equilibrium solution that maximizes the benefits of all participants, implement global load balancing evaluation, and select the optimal diversion path.
[0157] Path selection optimization based on multi-agent game realizes Nash equilibrium solution and diversion path selection through the following steps:
[0158] S541), Cooperative decision initialization
[0159] Each intelligent agent vehicle broadcasts the following information through the virtual communication network:
[0160] The current position and the target point; the probability distribution of the turning action output by the current policy network; the expected benefits of each turning action calculated based on step S4);
[0161] Construct a local revenue matrix within a radius of 300 meters. The rows of the matrix represent vehicles, and the columns represent optional steering actions;
[0162] S542), Iterative strategy optimization:
[0163] Strategy proposal stage: Each vehicle selects the steering action that maximizes its own revenue R based on the current strategies of neighboring vehicles; synchronously update the strategy selection results through the Ray RLlib parameter server; total Maximize the steering action; synchronously update the strategy selection results through the Ray RLlib parameter server;
[0164] Dynamic convergence determination: When the strategy change amplitude of more than 90% of the vehicles is less than 5% in three consecutive iterations, it is determined that Nash equilibrium is reached; if convergence has not occurred after more than 10 iterations, force the selection of the current optimal solution;
[0165] S543), Global load assessment:
[0166] Calculate the predicted vehicle density of each road according to the equilibrium solution, and update the road network load matrix in combination with the steering selection results; screen 3 alternative paths that optimize the following indicators: the path capacity attenuation rate < 15%, and minimize the global load variance;
[0167] Prioritize path resource allocation: The vehicle with the shortest remaining distance preferentially selects the optimal path; other vehicles are sequentially allocated sub-optimal paths in ascending order of the remaining distance;
[0168] S544), Emergency strategy trigger:
[0169] Immediately start the emergency mechanism in step S6) when the following situations are detected:
[0170] The predicted density of a certain section exceeds 85% of the maximum capacity;
[0171] The path conflict rate exceeds 10%.
[0172] In this embodiment, the steering decision generation mechanism converts the observed state into a steering decision through the following steps:
[0173] Action space construction:
[0174] Generate a set of legal steering actions according to the current road topology, such as {left turn, go straight, right turn}
[0175] Filter out inaccessible paths closed due to construction / accidents;
[0176] Policy network inference:
[0177] Input 128-dimensional decision features into the fully connected layer of the policy network;
[0178] Output the probability distribution of each steering action, for example: left turn: 0.2, go straight: 0.6, right turn: 0.2.
[0179] Exploration - exploitation balance:
[0180] Perform ε - greedy selection according to the dynamic exploration rate (the initial exploration rate is set to 0.2) in step S33): Select the action with the highest probability with a probability of (1 - ε), and randomly select an action with a probability of ε;
[0181] S6) When sudden congestion occurs, trigger the collaborative emergency mechanism. When sudden congestion is detected, perform the following emergency operations:
[0182] S61) Implement dynamic path replanning for vehicles within 200m;
[0183] S62) Set the priority lane permission to allow specific types of vehicles to pass first;
[0184] S63) Link with the traffic signal control system to implement diversion induction and dynamically adjust the signal phase duration; Among them, implementing diversion induction specifically includes the following steps:
[0185] S631) Construct a road network load matrix based on multi - agent collaborative perception data; Each element represents the remaining traffic capacity of the corresponding road; Among them:
[0186] Remaining traffic capacity = (section maximum capacity - current number of vehicles) × average vehicle speed / section length;
[0187] Among them, the section maximum capacity is dynamically calculated according to the number of lanes, and the capacity threshold per lane is set to 30 vehicles / km; The current number of vehicles is obtained through the fusion of distributed perception data of intelligent vehicles;
[0188] S632) Collaborative path planning decision - making, each intelligent agent performs the following operations:
[0189] a) Receive the real - time path planning decisions of other intelligent agents within a radius of 300 meters;
[0190] b) Construct a path selection benefit matrix based on the game decision - making model and calculate the Nash equilibrium solution;
[0191] c) Dynamically adjust the threshold according to the road network load matrix. When the regional vehicle density exceeds 30 vehicles / km, generate at least 3 alternative path sets that meet the remaining traffic capacity;
[0192] d) Select the optimal detour path that minimizes the global road network load variance;
[0193] S633) Large - scale path dynamic synchronization, execute the following distributed coordination mechanism:
[0194] a) Broadcast path change instructions through a virtual communication network to ensure the consistency of adjacent intelligent agent decisions;
[0195] b) Adopt a time window mechanism to update the paths in batches, and the number of vehicles updated in each batch does not exceed 20% of the total;
[0196] c) Implement priority arbitration for conflict path decisions and process them in ascending order according to the remaining distance of the vehicle to the target point.
[0197] Preferably, the cloud-edge collaborative architecture includes:
[0198] a) The cloud performs offline training of the multi-agent reinforcement policy network and generates global optimized model parameters through a distributed reinforcement learning framework;
[0199] b) The edge computing node deploys a lightweight policy network, receives the model parameters sent by the cloud, and executes real-time inference decisions;
[0200] c) The edge nodes adopt a location-aware communication protocol based on a geographical hash table to achieve collaborative state awareness within a radius of 500 meters.
[0201] The above embodiments and the descriptions in the specification only illustrate the principles and the best embodiments of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A dynamic traffic flow allocation method based on multi-agent reinforcement learning, characterized in that: The method described herein realizes dynamic optimization and real-time response of large-scale vehicle paths through a cloud-edge collaborative architecture, and includes the following steps: S1), construct a multi-agent traffic simulation environment, generate intelligent vehicles with random starting points according to the real-time traffic flow distribution and through the Poisson distribution model; S2), constructing a multidimensional observation vector including road topology coding, current road density, neighborhood speed average, target distance ratio, and congestion coefficient; S3) Build a multi-agent reinforcement strategy network and use the proximal strategy optimization algorithm to optimize the parallel strategy of the multi-agent reinforcement strategy network under the Ray RLlib framework; S4) Design a multi-objective reward mechanism based on path efficiency, congestion penalty and progress reward; S5), generating a steering decision based on the real-time observation status, and dynamically updating the vehicle driving route through the game theory Nash equilibrium solution; S6) When sudden congestion occurs, the coordinated emergency mechanism is triggered.
2. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 1, characterized in that: The topology structure is generated based on the OpenStreetMap road network data through the cloud server, and the vehicle generator is set to generate random starting point intelligent vehicles using the Poisson distribution model. The observation radius is 200 meters and the maximum steering angle is ±45 degrees.
3. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 1, characterized in that: In step S2), the construction of the multidimensional observation vector specifically includes the following steps: S21), mapping the current road ID into a 128-dimensional binary vector through a road identification hash algorithm, and compressing it into a unique code in the interval [0,1] after normalization, so as to perform road topology coding; S22), generating the current road density according to the ratio of the current number of road vehicles to the maximum capacity; S23), taking the average of the speeds of at least n adjacent vehicles within a radius of A meters as the neighborhood speed average; S24), calculating the target distance ratio according to the ratio of the shortest path distance from the current road to the target road to the maximum distance of the road network; S25) Calculate the real-time congestion coefficient based on the number of parked and waiting vehicles on the current road.
4. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 1, characterized in that: In step S3), the proximal policy optimization algorithm is used to perform parallel policy optimization of the multi-agent reinforcement policy network under the Ray RLlib framework, which specifically includes the following steps: S31), embedding the LSTM network in the strategy network to predict future congestion trends; and dynamically adjusting the memory weight by splicing the current multi-dimensional observation vector and historical data and inputting them into the LSTM network; S32) Build a distributed cluster based on parameter servers under the Ray RLlib framework to share multi-agent policy parameters; the central server stores global policy network parameters, and edge nodes synchronously update local models; S33), implement the proximal strategy optimization algorithm, use the dynamic KL threshold to control the strategy update amplitude, and adaptively adjust the strategy update step size.
5. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 1, characterized in that: In step S4), the expression of the multi-objective reward mechanism is: <h2 style=";text-align:left;direction:ltr">R<h2 style=";text-align:left;direction:ltr"> total <h2 style=";text-align:left;direction:ltr"> <0.5R<h2 style=";text-align:left;direction:ltr"> route <h2 style=";text-align:left;direction:ltr"> +0.3R<h2 style=";text-align:left;direction:ltr"> progress <h2 style=";text-align:left;direction:ltr"> +R<h2 style=";text-align:left;direction:ltr"> congestion <h2 style=";text-align:left;direction:ltr"> +R<h2 style=";text-align:left;direction:ltr"> penalty In the formula, R total R is a comprehensive reward; route The estimated travel time difference between the paths before and after the turn; R progress Reward for reaching progress; R congestion is the real-time vehicle density change; R penalty Penalty for ineffective steering.
6. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 5, characterized in that: In step S4), after the vehicle performs a turning action, the estimated travel time difference of the path before and after the turning is obtained based on the real-time traffic interface, and the estimated travel time difference R before and after the turning is obtained. route The expression is: R route =(T prev -T new ) / T prev ×η; Where, T prev T is the travel time of the remaining road sequence of the original path obtained through the road network interface; new To obtain the travel time of new routes through the real-time traffic flow prediction interface; η is the reward coefficient; The real-time vehicle density change R congestion The calculation is based on the real-time vehicle density change entering the road after turning, and its expression is: R congest i on =-(N new -N prev )×μ; Where N new To obtain the density of new vehicles entering the road through the dynamic perception interface; N prev is to obtain the vehicle density of the next road originally planned through the road network topology interface; μ is the penalty coefficient; The progress reward R progress Based on the residual distance attenuation function calculation, its calculation formula is: Where D remaibg is the remaining path length; κ is the reward cardinality; When the turning causes the path length to increase by more than a threshold, the invalid turning penalty R is triggered. penalty , whose expression is: R penalty =-min(0,(L new -L orignal ) / L orignal )×ω; Where, L new is the total length of the new path; L orignal is to obtain the total length of the original path; ω is the penalty coefficient.
7. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 1, characterized in that: In step S5), the vehicle driving route is dynamically updated through the game theory Nash equilibrium solution, specifically: S51), obtain all legal turning edges of the current road and filter out unreachable paths; S52), generating a new path including the selected turning edge according to the action selection result; S53), detecting invalid paths and triggering path reset operations, and using Dijkstra algorithm to verify path reachability; S54) Based on the multi-agent game decision model, the Nash equilibrium solution of path selection is calculated, and through iterative strategy updates, the equilibrium solution that maximizes the benefits of all participants is found, global load balancing evaluation is implemented, and the optimal diversion path is selected.
8. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 1, characterized in that: In step S6), when sudden congestion is detected, the following emergency operations are performed: S61), implementing dynamic path replanning for vehicles within a certain range; S62) Setting priority lane permissions to allow certain types of vehicles to have priority; S63) The linked traffic signal control system implements diversion induction and dynamically adjusts the signal phase duration.
9. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 8, characterized in that: In step S63), shunt induction is performed, which specifically includes the following steps: S631) Construct a road network load matrix based on multi-agent collaborative perception data; each element represents the remaining capacity of the corresponding road; wherein: Remaining capacity = (maximum capacity of the road section - current number of vehicles) × average speed / length of the road section; The maximum capacity of the road section is dynamically calculated based on the number of lanes, and the capacity threshold of each lane is set to 30 vehicles / km; the current number of vehicles is obtained through the distributed perception data fusion of the intelligent vehicle; S632), collaborative path planning decision, each agent performs the following operations: a) Receive real-time path planning decisions from other agents within a 300-meter radius; b) Construct the path selection benefit matrix based on the game decision model and calculate the Nash equilibrium solution; c) Dynamically adjust the threshold according to the road network load matrix, and generate at least three alternative path sets that meet the remaining capacity when the regional vehicle density exceeds 30 vehicles / km; d) Selecting the optimal detour path that minimizes the global road network load variance; S633), large-scale path dynamic synchronization, execute the following distributed coordination mechanism: a) Broadcasting path change instructions through the virtual communication network to ensure the consistency of decisions made by adjacent agents; b) Use a time window mechanism to update routes in batches, with the number of vehicles updated in each batch not exceeding 20% of the total; c) Implement priority arbitration for conflicting path decisions and process them in ascending order based on the remaining distance from the vehicle to the target point.
10. The method for dynamic traffic flow allocation based on multi-agent reinforcement learning according to claim 1, characterized in that: The cloud-edge collaborative architecture includes: a) Perform offline training of the multi-agent reinforcement strategy network in the cloud and generate global optimization model parameters through a distributed reinforcement learning framework; b) Edge computing nodes deploy lightweight strategy networks, receive model parameters sent from the cloud, and perform real-time reasoning decisions; c) A location-aware communication protocol based on a geographic hash table is used between edge nodes to achieve collaborative status awareness within a radius of 500 meters.
Citation Information
Cited By
Bidirectional multi-lane expressway traffic flow dynamic control method and system based on traffic simulation
CN120356341A
Multimodal transport path selection method based on multi-agent reinforcement learning
CN120765154A
Disaster reduction and rescue path planning method and system based on Internet of Things
CN121031933A
Detour scheduling method based on large-scale grid map
CN121165797A
AI-enabled smart city traffic jam prediction and dispersion method
CN121527995A