Multi-target traffic path guidance method and system based on multiple agents
Through multi-target traffic path induction methods and systems based on multi-agents, the problem of unbalanced traffic resource utilization caused by single induction targets in the existing navigation system is solved, the balance between traffic efficiency and emissions is achieved, and the efficiency and responsiveness of urban traffic management are improved.
Patent Information
- Application Number
- CN202510217378.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The inducement of single targets in the existing navigation system leads to unbalanced utilization of transportation resources, unable to effectively balance traffic efficiency with greenhouse gas emissions, and difficult to meet the diversified needs of urban management.
Multi-target traffic path induction methods and systems based on multi-agents are adopted, and multiple target optimization of travel time and emissions are achieved through the combination of simulation modules, optimization control modules, iterative optimization modules, judgment modules and output modules, multiple target optimization modules, and differentiated traffic path induction, and flexibly adjust the induction strategy according to urban management needs.
It significantly improves the overall efficiency of the transportation network, reduces vehicle emissions, achieves dynamic balance of transportation resources, enhances the efficiency and responsiveness of urban traffic management, provides differentiated path induction, and improves travelers' satisfaction and compliance with path selection.
Smart Images

Figure CN120014831A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic management, and in particular to a multi-agent-based multi-target traffic path induction method and system. Background Art
[0002] With the popularization of mobile communication technology, route guidance - providing drivers with real-time traffic information and optimal traffic routes through information systems (such as smartphone applications, variable message signs, and vehicle applications) - plays an increasingly important role in urban transportation. Studies have shown that route guidance can significantly affect the efficiency and greenhouse gas emissions of the transportation system. In the case of dense urban road traffic networks, it has become a feasible idea to improve the efficiency of the transportation system and reduce traffic carbon emissions through route guidance methods.
[0003] However, traditional route guidance methods often focus on a single goal (such as the shortest travel time), while ignoring the balance between traffic efficiency and greenhouse gas emissions, as well as the diverse needs of urban management or the diversified preferences of travelers. Therefore, how to establish an effective multi-intelligence reinforcement learning method to provide travelers with differentiated traffic route guidance, and be able to flexibly switch between different guidance goals to better meet the traffic guidance needs in different situations, and grasp and enhance the ability of multi-target traffic guidance to regulate the entire urban traffic system, is a key issue that needs to be solved urgently. Summary of the invention
[0004] The present invention aims to solve the problem that the existing navigation system has a single induction target, which leads to unbalanced utilization of traffic resources. A multi-target traffic path induction method and system based on multi-agents are proposed. The present invention takes into account multiple targets such as travel time and emissions, and provides differentiated induction paths for travelers with different start-end pairs (OD pairs) and different travelers with the same OD pair. While improving the efficiency of the traffic system, it reduces system emissions and improves the overall efficiency of the traffic system. At the same time, the present invention can also flexibly adjust the induction strategy according to the needs of urban management, thereby improving the efficiency and responsiveness of urban traffic management. The specific technical scheme is as follows:
[0005] In one aspect of the present invention, a multi-agent-based multi-target traffic path guidance system is provided, the system comprising:
[0006] A simulation module for calibrating simulated traffic demand;
[0007] An optimization control module to maximize the cumulative reward of the entire system over time;
[0008] Iterative optimization module, used to implement iterative optimization of strategies in information interaction;
[0009] A judgment module is used to judge whether the system revenue tends to be stable through the simulation module information;
[0010] Output module, used to output control solutions.
[0011] Specifically, in the simulation module, SUMO is used as a micro-simulation platform to calibrate the traffic demand, and 3-4 alternative paths are pre-screened for each OD pair based on feasibility.
[0012] Specifically, in the optimization control module, a multi-agent Markov decision process model is included, and its optimization goal is to maximize the cumulative reward of the entire system over time by improving the strategy of the agent.
[0013] Specifically, the agent’s decision variables include state, action, and reward;
[0014] Among them, the state is a measure of the traffic situation, represented by a vector S t =(b t,i ,v t,i ,l t,i ) indicates that, where b t,i represents the number of vehicles on each edge at a given time step, v t,i represents the average vehicle speed, l t,i represents the average driving distance, and the action of each agent j is an induced diversion rate, which is proportionally distributed among the pre-determined alternative paths;
[0015] The action vector is represented as where m(j) represents the number of alternative paths. For example, for a particular OD pair d, action represents the ratio of vehicles assigned to alternative routes 1, 2, and 3;
[0016] The reward is defined as a comprehensive indicator of travel time (TT) reflecting traffic efficiency and carbon dioxide emissions (CE) reflecting environmental impact, expressed as:
[0017] r t =-αTT t -(1-α)*γ*CE t
[0018] Among them, TT t =∑TT t,e is the time it takes for vehicle e to complete the entire journey under the current road conditions, CE t =∑CE t,i is the CO2 emission of road section i;
[0019] The parameter γ is used to unify the order of magnitude of the two values. The specific value is estimated based on the total driving time and total emissions in the simulation scenario, as well as the total driving time and total emissions of a single objective optimization;
[0020] The optimization goal of the model is adjusted by the weight α: when α = 0, the model goal is to reduce emissions; when α = 1, the model prioritizes travel time and efficiency;
[0021] When α = 0.5, the model takes reducing emissions and improving traffic efficiency as common goals, aiming to find an induced optimization solution that can achieve a balance between the two goals.
[0022] Specifically, in the optimization control module, it also includes the use of an initialized reinforcement learning algorithm to optimize the traffic efficiency and emissions of the entire network;
[0023] The initial reinforcement learning algorithm is the VDA2C algorithm, which uses the transitions obtained from the multi-agent Markov decision process as input to learn the agent's state-action mapping and returns the reward action according to the updated strategy.
[0024] Specifically, in the optimization control module, it also includes setting traveler preference attributes.
[0025] Specifically, the implementation of strategy iteration optimization in information interaction includes: selecting OD pairs and determining corresponding alternative routes based on traffic demand analysis and simulation scenarios;
[0026] An information interaction mechanism is established between the simulation module and the optimization control module to collect the reward information obtained by the agent when taking various actions in different states through the simulation process;
[0027] The dynamic adjustment mechanism in reinforcement learning is introduced, and the policy gradient method is used to optimize the strategy.
[0028] Specifically, the standard for judging whether the system benefit tends to be stable through the simulation module information is: when the reward value output by the simulation module tends to be stable, or the algorithm reaches the maximum number of iterations, the optimization process is considered to be over.
[0029] Specifically, the output control plan includes verifying the effectiveness of the optimization strategy by comparing the efficiency and emissions of the transportation system before and after optimization after the strategy optimization is completed.
[0030] Another aspect of the present invention provides a multi-agent based multi-target traffic path guidance method, comprising the following steps:
[0031] S1: Initialize the simulation module to calibrate the simulated traffic demand;
[0032] S2: Initialize the optimization control module to maximize the cumulative reward of the entire system over time;
[0033] S4: Use the iterative optimization module to achieve iterative optimization of strategies in information interaction;
[0034] S5: Using the information from the simulation module to determine whether the system revenue is stable through the judgment module;
[0035] S6: Output the control plan through the output module.
[0036] Beneficial effects:
[0037] 1. The present invention provides differentiated path guidance information to avoid traffic congestion caused by homogeneous navigation information, achieve dynamic balance of traffic flow, significantly improve the overall efficiency of the traffic network and reduce vehicle emissions, and provide strong support for the development of green transportation. Tests based on traffic data from Liushi Expressway and surrounding areas in Hangzhou have proven the effect of optimizing the guidance strategy on improving system performance under multiple objectives. Among them, under the ecological goal, the overall efficiency of the road network increased by 13.90%, and carbon dioxide emissions decreased by 16.09%, achieving a significant win-win effect in achieving the two major goals of improving efficiency and reducing emissions.
[0038] 2. This paper proposes a multi-agent Markov decision framework suitable for real-time traffic induction, which can flexibly adapt to different induction targets and respond to diverse urban traffic management needs. This framework improves the accuracy and adaptability of traffic management and provides a solid theoretical and practical foundation for efficient induction in complex traffic scenarios.
[0039] 3. The present invention fully considers the individual preferences of travelers and proposes a differentiated and customized route induction strategy, which can deeply explore the potential laws of travelers' needs and behavior patterns, and significantly improve travelers' satisfaction and compliance with route selection. Under the optimized induction strategy, the allocated routes show significant tendencies for travelers with different preference attributes, opening up a new direction for personalized traffic management. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a structural schematic diagram of the induced optimization method of the present invention;
[0041] Figure 2 It is a schematic diagram of the execution process of the simulation module of the present invention;
[0042] Figure 3 It is a schematic diagram of the execution process of the optimization control module of the present invention;
[0043] Figure 4 It is a schematic diagram of the framework of the multi-agent Markov decision process of the present invention;
[0044] Figure 5It is a schematic diagram of the calculation process of the algorithm involved in the present invention. DETAILED DESCRIPTION
[0045] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. The embodiments of the present invention and the technical features in the embodiments may be combined with each other unless there is a conflict.
[0046] The term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0047] Embodiment 1, a multi-agent-based multi-objective traffic path guidance system, the system comprising: a simulation module, for calibrating simulated traffic demand;
[0048] An optimization control module to maximize the cumulative reward of the entire system over time;
[0049] Iterative optimization module, used to implement iterative optimization of strategies in information interaction;
[0050] A judgment module is used to judge whether the system revenue tends to be stable through the simulation module information;
[0051] Output module, used to output control solutions.
[0052] Specifically, in actual application, the simulation module is first initialized to calibrate the simulated traffic demand, that is, SUMO (an open source micro traffic simulation tool) is selected as the micro simulation platform, and the traffic demand is calibrated, that is, the simulated traffic demand is calibrated according to the road flow, speed and other data of the actual traffic environment. In order to improve the calculation efficiency and ensure the driver's acceptance of the induction, 3-4 alternative paths are preset for each OD pair according to the feasibility (that is, the acceptable degree of detour) and compiled into the routing file. Write a Python script that communicates with the TraCI interface of the platform, and use commands such as getLastStepMeanSpeed to record specific indicators (such as road section travel time, carbon emissions, etc.); according to the induction strategy, that is, the vehicle diversion rate between several alternative paths obtained by the control module, use the setRouteID command to adjust the vehicle path.
[0053] Next, the optimization control module is initialized. The optimization control module contains a multi-agent Markov decision process model, whose optimization goal is to maximize the cumulative reward of the entire system over time by improving the strategy of the agent. The decision variables of the agent include the current state and the selected action, and the strategy is the mapping rule from state to action. Considering the complexity of traffic problems, each OD pair is responsible for decision-making and optimization by an agent. According to the induced objectives (such as shortest travel time or minimum emissions) and the corresponding alternative paths, the state, action and reward of each agent in the Markov state transfer equation are defined. Compared with the single-agent model, the use of multi-agent design can significantly reduce the action dimension of each agent (the diversion rate of all alternative paths), thereby improving the learning efficiency of the model.
[0054] More specifically, in the present invention, the states, actions, and rewards are defined as follows:
[0055] The state is a measure of the traffic situation, represented by a vector S t =(b t,i ,v t,i ,l t,i ) indicates that, where b t,i represents the number of vehicles on each edge at a given time step, v t,i represents the average vehicle speed, l t,i represents the average driving distance. The action of each agent j is an induced diversion rate, which is proportionally distributed among the pre-determined alternative paths. The action vector is represented as where m(j) represents the number of alternative paths. For example, for a particular OD pair d, action Indicates the proportion of vehicles assigned to alternative routes 1, 2, and 3.
[0056] The reward is defined as a comprehensive indicator of travel time (TT) reflecting traffic efficiency and carbon dioxide emissions (CE) reflecting environmental impact, expressed as:
[0057] r t =-αTT t -(1-α)*γ*CE t
[0058] TT t =∑TT t,e is the time it takes for vehicle e to complete the entire journey under the current road conditions, CE t =∑CE t,iis the CO2 emission of road section i. The parameter γ is used to unify the order of magnitude of the two values. The specific value is estimated based on the total travel time and total emissions in the simulation scenario, as well as the total travel time and total emissions of a single objective optimization. The optimization goal of the model is adjusted by the weight α: when α = 0, the model goal is to reduce emissions; when α = 1, the model prioritizes travel time and efficiency; when α = 0.5, the model takes reducing emissions and improving travel efficiency as common goals, aiming to find an induced optimization solution that can achieve a balance between the two goals.
[0059] The initialization of the reinforcement learning algorithm includes adopting the VDA2C algorithm to enhance the collaboration between multiple OD pairs to optimize the traffic efficiency and emissions of the entire network. The state-action mapping (i.e., strategy) of the agent is learned using the transitions (i.e., state, action, and reward) obtained from the multi-agent Markov decision process as input, and the reward action is returned according to the updated strategy. The value function V is used to evaluate the expected cumulative reward of the strategy π after the agent takes action in a given state. The value function V is calculated by the formula Calculate, where represents the local value function of the ith agent in the next state s', r i (s) is the immediate reward of agent i in state s, φ i are the parameters of agent i. The global value function is calculated by mixing the local value functions of multiple agents: Where V tot is the global value function, which represents the comprehensive evaluation of the joint actions of all agents, g ψ is a hybrid network for combining local values, ψ is the parameter of the hybrid network, and u is the joint action of all agents. Within a given time frame, VDA2C continuously interacts and iterates with the environment, and updates the value function through the following steps:
[0060] ① Sampling phase: The agent interacts with the environment and collects data (i.e. state, action, reward, etc.), which is called the sampling phase.
[0061] ② Feedback stage: Agents calculate local value functions based on their own experience and share information. The local value functions of all agents are combined through a hybrid network to generate a global value function.
[0062] ③ Update phase: Calculate the advantage function A based on the global value function and immediate reward t , and update the strategy. The advantage function is calculated by the following formula:
[0063] A(s')=r(s')+λV tot (s')-V tot (s)
[0064] Among them, r(s') is the immediate reward of the next state s', V tot (s') and V tot (s) is the value of the global value function at the next state s' and the current state s, and λ is the discount factor used to balance the immediate reward and future reward.
[0065] Based on the advantage function A(s'), the policy gradient method is used to calculate the policy gradient and update the policy parameter θ. The specific calculation method is as follows:
[0066]
[0067] Among them, σ is the learning rate, which controls the parameter update step.
[0068] Through the above steps, the intelligent agent gradually improves its decision-making ability, thereby improving the traffic efficiency and emission reduction effects of the entire network.
[0069] When setting the preference attributes of travelers, when travelers have different preferences, each agent corresponds to a group of travelers with preference characteristics. Set different preference attributes for different agents, such as time-sensitive, environmentally friendly, etc. Adjust the weights in the reward function according to the preference attributes to reflect the needs of different travelers. For example, time-sensitive travelers pay more attention to reducing travel time, while environmentally friendly travelers pay more attention to reducing carbon dioxide emissions.
[0070] In the process of realizing policy iteration optimization in information interaction, policy iteration is a crucial link in the optimization process based on SUMO traffic simulation environment. In order to ensure the effectiveness and adaptability of the strategy, a series of information interaction steps are taken.
[0071] According to the traffic demand analysis and simulation scenarios, OD pairs are selected and corresponding alternative paths are determined. In the process of path selection, the current traffic conditions and road characteristics need to be comprehensively considered to ensure the rationality and effectiveness of the selected path in practical applications. An information interaction mechanism is established between the simulation module and the optimization control module, and the reward information obtained by the agent for taking various actions in different states is collected through the simulation process. These reward information provides key data support for subsequent strategy optimization. The agent uses the collected reward information to update the strategy, thereby improving its decision-making quality in specific situations. Through continuous strategy optimization, the agent gradually adapts to the complex traffic environment and makes more effective decisions. In order to further improve the learning ability of the agent, the dynamic adjustment mechanism in reinforcement learning is introduced, and the policy gradient method is used to optimize the strategy. The optimized strategy is passed back to the simulation module through the information interaction mechanism to generate new state and reward information, thereby further updating the strategy.
[0072] Among them, when judging whether the system benefit tends to be stable through the simulation module information, when the reward value output by the simulation module tends to be stable, or the algorithm reaches the maximum number of iterations, the optimization process can be considered to be over.
[0073] When outputting the control plan, that is, after the strategy optimization is completed, the effectiveness of the optimization strategy is verified by comparing the efficiency and emissions of the traffic system before and after optimization. The impact of different strategies on different routes and different OD vehicles is analyzed to ensure that the proposed route guidance strategy can improve the overall traffic management effect in practical applications.
[0074] Embodiment 2, a multi-target traffic path guidance method based on multi-agent, comprises the following steps:
[0075] S1: Initialize the simulation module to calibrate the simulated traffic demand;
[0076] S2: Initialize the optimization control module to maximize the cumulative reward of the entire system over time;
[0077] S4: Use the iterative optimization module to achieve iterative optimization of strategies in information interaction;
[0078] S5: Using the information from the simulation module to determine whether the system revenue is stable through the judgment module;
[0079] S6: Output the control plan through the output module.
[0080] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A multi-agent based multi-target traffic path guidance system, characterized by: The system includes: A simulation module for calibrating simulated traffic demand; An optimization control module to maximize the cumulative reward of the entire system over time; Iterative optimization module, used to implement iterative optimization of strategies in information interaction; A judgment module is used to judge whether the system revenue tends to be stable through the simulation module information; Output module, used to output control solutions.
2. A multi-agent based multi-target traffic path guidance system as claimed in claim 1, characterized in that: In the simulation module, SUMO is used as a micro-simulation platform to calibrate the traffic demand, and 3-4 alternative paths are pre-screened for each OD pair based on feasibility.
3. The multi-agent-based multi-target traffic path guidance system according to claim 1, characterized in that: In the optimization control module, a multi-agent Markov decision process model is included, and its optimization goal is to maximize the cumulative reward of the entire system over time by improving the strategies of the agents.
4. The multi-agent-based multi-target traffic path guidance system according to claim 1, characterized in that: The agent’s decision variables include state, action, and reward; Among them, the state is a measure of the traffic situation, represented by a vector S t =(b t,i ,v t,i ,l t,i ) indicates that, where b t,i represents the number of vehicles on each edge at a given time step, v t,i represents the average vehicle speed, l t,i represents the average driving distance, and the action of each agent j is an induced diversion rate, which is proportionally distributed among the pre-determined alternative paths; The action vector is represented as Where m(j) represents the number of alternative paths; The reward is defined as a comprehensive indicator of travel time (TT) reflecting traffic efficiency and carbon dioxide emissions (CE) reflecting environmental impact, expressed as: r t =-αTT t -(1-a)*c*CE t Among them, TT t =∑TT t,e is the time it takes for vehicle e to complete the entire journey under the current road conditions, CE t =∑CE t,i is the CO2 emission of road section i; The parameter γ is used to unify the order of magnitude of the two values. The specific value is estimated based on the total driving time and total emissions in the simulation scenario, as well as the total driving time and total emissions of a single objective optimization; The optimization goal of the model is adjusted by the weight α: when α = 0, the model goal is to reduce emissions; when α = 1, the model prioritizes travel time and efficiency; When α = 0.5, the model takes reducing emissions and improving traffic efficiency as common goals, aiming to find an induced optimization solution that can achieve a balance between the two goals.
5. The multi-agent-based multi-target traffic path guidance system according to claim 3, characterized in that: In the optimization control module, it also includes the use of an initialized reinforcement learning algorithm to optimize traffic efficiency and emissions across the entire network; The initial reinforcement learning algorithm is the VDA2C algorithm, which uses the transitions obtained from the multi-agent Markov decision process as input to learn the agent's state-action mapping and returns the reward action according to the updated strategy.
6. A multi-agent based multi-target traffic path guidance system as claimed in claim 3, characterized in that: In the optimization control module, it also includes setting traveler preference attributes.
7. The multi-agent-based multi-target traffic path guidance system according to claim 1, characterized in that: The implementation of strategy iteration optimization in information interaction includes: selecting OD pairs and determining corresponding alternative routes based on traffic demand analysis and simulation scenarios; An information interaction mechanism is established between the simulation module and the optimization control module to collect the reward information obtained by the agent when taking various actions in different states through the simulation process; The dynamic adjustment mechanism in reinforcement learning is introduced, and the policy gradient method is used to optimize the strategy.
8. The multi-agent-based multi-target traffic path guidance system according to claim 1, characterized in that: The criterion for judging whether the system benefit tends to be stable through the simulation module information is: when the reward value output by the simulation module tends to be stable, or the algorithm reaches the maximum number of iterations, the optimization process is considered to be over.
9. The multi-agent-based multi-target traffic path guidance system according to claim 3, characterized in that: The output control plan includes verifying the effectiveness of the optimization strategy by comparing the efficiency and emissions of the transportation system before and after optimization after the strategy optimization is completed.
10. A multi-agent based multi-target traffic path guidance method, characterized in that: The steps include: S1: Initialize the simulation module to calibrate the simulated traffic demand; S2: Initialize the optimization control module to maximize the cumulative reward of the entire system over time; S4: Use the iterative optimization module to achieve iterative optimization of strategies in information interaction; S5: Using the information from the simulation module to determine whether the system revenue is stable, the judgment module is used; S6: Output the control plan through the output module.
Citation Information
Patent Citations
Real-time traffic guidance system and method considering long-term compliance behavior change of traveler
CN113506445A
Emergency signal priority and social vehicle dynamic path induction collaborative optimization method
CN113724510A
Operating vehicle-oriented road side real-time accurate induction method for road confluence area
CN114863708A
Multi-agent road traffic signal control method based on deep reinforcement learning algorithm
CN116863729A
Traffic flow distribution method and system suitable for intelligent traffic operation vehicle-road cooperation
CN117173911A
Cited By
Multi-mode traffic simulation method based on multi-agent sequential decision
CN121600715A
A traffic demand regulation system and method based on behavioral experiments
CN122617083A