Trunk line intersection cooperative control method considering robustness
Through the multi-agent collaborative control method of deep reinforcement learning and robust optimization, the robustness and coordination problems of trunk intersections in the dynamic traffic environment in the prior art are solved, and efficient traffic flow management and stability improvement are achieved.
Patent Information
- Application Number
- CN202510380116.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-28
AI Technical Summary
When faced with complex changes in dynamic traffic environments, existing traffic control methods are difficult to deal with uncertain factors such as sudden changes in traffic flow and bad weather, resulting in instability of control strategies. The coordinated control of multiple agents does not fully integrate the dynamic correlation characteristics of adjacent intersections, which is easy to fall into local optimality and affect the efficiency of trunk traffic.
Deep reinforcement learning, robust optimization and multi-agent collaboration mechanisms are adopted to build a robust trunk intersection collaborative control method. By modularly configuring each intersection as an agent, combining perception, decision-making, execution and communication modules, a robust optimization strategy and deep reinforcement learning algorithm are introduced to establish a robust multi-agent network framework to realize signal light collaborative control under dynamic traffic flow.
It improves the traffic efficiency and system stability of the trunk line intersection, can effectively respond to traffic flow fluctuations and emergencies, improves the system's robustness and adaptability in uncertain environments, and reduces traffic delays.
Smart Images

Figure CN120279737A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent traffic control, and in particular, to a coordinated control method for arterial intersections considering robustness. Background Art
[0002] With the acceleration of urbanization, arterial intersections are facing challenges of surging and fluctuating traffic flows. Traditional traffic control methods rely on fixed signal cycles and static traffic flow models, making it difficult to cope with complex changes in dynamic traffic environments. During peak hours or in the event of emergencies, fixed timing plans are prone to problems such as queue spills and sudden increases in vehicle delays, severely restricting the traffic efficiency of arterial roads. Although coordinated control methods based on multi-agent systems have improved flexibility through distributed decision-making, existing technologies still have significant limitations: they lack robust modeling for uncertain factors such as sudden traffic flow changes and bad weather, and control strategies are prone to instability under perturbations; multi-agent coordination mechanisms mostly rely on local information interaction and do not fully integrate the dynamic correlation characteristics of adjacent intersections, easily falling into local optima and sacrificing the overall coordination of the arterial road; in addition, traditional reinforcement learning algorithms converge slowly in dynamic scenarios and do not introduce robust constraints, resulting in lagging strategy adjustments in the face of emergencies and exacerbating traffic imbalance. For example, although existing patent solutions attempt to dynamically adjust signal timing, they do not incorporate road conditions and environmental perception parameters into real-time decision-making, leading to a significant decline in control effectiveness under complex perturbations. Therefore, in the context of increasing traffic flow uncertainty, how to construct a coordinated control method for arterial intersections with robust response, global coordination, and online adaptability has become a key problem in improving urban traffic operation efficiency. Summary of the Invention
[0003] The object of the present invention is to overcome the deficiencies in the prior art and provide a coordinated control method for arterial intersections considering robustness. By integrating deep reinforcement learning, robustness optimization, and multi-agent cooperation mechanisms, it realizes the coordinated control of traffic lights under dynamic traffic flows to cope with traffic flow fluctuations, emergencies, and environmental interferences, and improves the traffic efficiency and system stability of arterial intersections.
[0004] To achieve the above object, the present invention provides a coordinated control method for arterial intersections considering robustness, including the following steps:
[0005] Step S1, Construction of Multi-Agent Arterial Intersection System: Modular configuration, regarding each intersection on the traffic artery as an agent. Each agent has independent decision-making ability and information processing ability. Each agent is an independent module with good scalability, and can be conveniently added, deleted or replaced as the city develops. Configure a perception module, a decision-making module, an execution module and a communication module. Agents exchange information through a communication network to achieve cooperative control. The perception module obtains real-time traffic data in real time, and performs data cleaning and feature extraction to provide input for subsequent modeling of the uncertainty fluctuation model.
[0006] Step S2, Robustness Configuration of Arterial Intersections: Considering the uncertainty of arterial traffic flow, such as fluctuations in vehicle arrival rates, changes in road conditions, etc., introduce a robust optimization strategy. By proposing a robustness modeling strategy based on deep reinforcement learning and introducing a negative binomial moving average randomization algorithm, enhance the robustness of the system by dynamically adjusting the model parameters combined with real-time environmental perception information, and propose a modeling of the uncertainty fluctuation model of arterial traffic flow. Propose a robustness objective setting function and robustness constraint conditions, set the minimization of vehicle delay and traffic congestion degree as the robustness objective, and dynamically adjust the parameters of the objective function according to real-time perception data to cope with the changes in the environment for signal coordination control of arterial intersections and improve the robustness of the control system.
[0007] Step S3, Robust Multi-Agent Model Configuration: Build a Robustness-Mainline-IntelligentAgent-Actor-Critic (RMIAC) network framework through the traffic flow uncertainty fluctuation model, robustness objective setting and robustness constraint conditions in Step S2, integrating the robust traffic joint traffic state space, action control parameters and robust reward function, so that each agent can not only autonomously optimize its own decision-making, but also comprehensively consider the dynamic changes of adjacent intersections to ensure the optimization of the global traffic flow. Fully consider the traffic flow dynamics at intersections and their relevance to adjacent intersections, propose the definition of the robust joint traffic state space, the definition of action control parameters and introduce a multi-level feedback mechanism to propose the definition of the robust reward function.
[0008] Step S4, Trunk Intersection Cooperative Control Strategy Configuration: Use the deep reinforcement learning algorithm to train the agents, and propose a robust agent trunk intersection signal cooperative control algorithm (CR-MADDPG), enabling it to select appropriate control strategies according to the current traffic state, so that each agent, while pursuing its own maximum benefit, considers the impact on the overall traffic flow of the trunk line; through information exchange and cooperation among the agents, realize the cooperative control of the trunk intersections; introduce a robustness parameter to enhance the system's ability to respond to emergencies (such as traffic accidents, weather changes, etc.), and propose a strategy update formula considering robustness; combine the robustness parameter to propose an agent objective function, which not only considers the reward of traffic flow, but also adds a robustness adjustment factor based on environmental perception, thereby improving the system's adaptability in an uncertain environment; propose a loss function considering robustness, which, in addition to considering the standard error, also adds a robustness adjustment term, comprehensively considering the different challenges faced by the agents in cooperation and competition, and ensuring the stable operation of the system in an uncertain environment;
[0009] Step S5, Online Learning and Adaptive Adjustment: During the actual operation process, use the multi-agent deep deterministic policy gradient algorithm considering robustness in Step S4 as the initial policy for online learning, collect experiences according to real-time traffic data, and randomly select a batch of samples to update the network parameters of the agents for fine-tuning the online learning strategy, that is, the agents can collect experiences, conduct online learning, randomly select small batches of samples to update the network parameters according to real-time traffic data, and continuously optimize the control strategy; at the same time, the agents can adaptively adjust the control parameters according to the changes in traffic flow, improving the adaptability and robustness of the control system.
[0010] Beneficial Effects: By introducing a robustness modeling strategy based on deep reinforcement learning and a negative binomial moving average randomization algorithm, dynamically adjusting the model parameters in combination with environmental perception information, effectively coping with uncertainty factors such as vehicle arrival rate fluctuations, road condition changes, and weather conditions. The setting of the robustness objective function and constraints, combined with the robust reward function, ensures that the multi-agent trunk intersection system can still operate stably under emergencies such as accidents and extreme weather. Compared with the traditional fixed-cycle method, the traffic delay is significantly reduced.
[0011] Furthermore, the construction of the multi-agent trunk intersection system described in Step S1 specifically includes:
[0012] S11. Construction of a single intersection agent: Each intersection is regarded as an independent agent with independent decision-making and information processing capabilities, responsible for monitoring and managing the traffic conditions at its intersection, including vehicle arrivals, departures, lane occupancy, signal states, etc. It can make appropriate control decisions based on its own state and the surrounding environment, and exchange information with other agents to achieve coordinated control between arterial intersections;
[0013] S12. Configuration of the agent software structure: It includes a perception module, a decision-making module, an execution module, and a communication module. The perception module is responsible for collecting traffic data; the decision-making module makes decisions based on the proposed algorithm; the execution module is responsible for controlling traffic facilities such as signal lights; the communication module is responsible for exchanging information with other agents;
[0014] S13. Establishment of a coordinated control mechanism: The goal of coordinated control is to enable all intersections on the arterial to achieve the optimal traffic state on the basis of achieving robustness. The proposed deep learning algorithm is used to train the agents so that they can select appropriate coordinated control strategies according to the current traffic state;
[0015] S14. Configuration of system scalability: Modular configuration is adopted. Each agent is an independent module with good scalability and can be easily added, deleted, or replaced as the city develops.
[0016] Furthermore, the robust configuration of the arterial intersections described in step S2 specifically includes the following steps:
[0017] S21. Modeling of the uncertainty fluctuations of arterial traffic flow: To effectively cope with the uncertainty fluctuations of arterial traffic flow, the impacts of real-time environmental factors such as vehicle arrival rate, road conditions, and weather conditions are considered; a robustness modeling strategy based on deep reinforcement learning is proposed, and the negative binomial moving average randomization algorithm is introduced By dynamically adjusting the model parameters combined with real-time environmental perception information, the robustness of the system is enhanced; the specific traffic flow modeling formula is:
[0018]
[0019] In the formula: Q m is the maximum arterial traffic flow, k is the number of vehicles arriving per unit time, p k is the probability of k vehicles arriving per unit time, β is the difference between the maximum arterial traffic flow Q m and the number of vehicles arriving per unit time, δ B is the vehicle arrival fluctuation parameter, with a value range of [0, 2], δ W is the weather condition parameter, with a value range of [0, 1], δ Cis the road condition change parameter, and its value range is [0,1], δ Env is the environmental perception information parameter, which represents the impact of real-time environmental changes including road conditions, accidents, etc., and its value range is [0,1].
[0020] S22. Robustness objective setting: To ensure the stability of the system under dynamic traffic flow and real-time environmental changes, the minimization of vehicle delay and traffic congestion degree is set as the robustness objective; according to the real-time perception data, the parameters of the objective function are dynamically adjusted to cope with environmental changes; the updated robustness objective function is as follows:
[0021]
[0022] In the formula: Q is the uncertain fluctuating traffic flow, Q m is the maximum traffic flow of the arterial road, R is the red light time at the intersection, a is the robustness parameter, ε R,d is the vehicle delay robustness parameter, and its value range is [0,a], ε R,j is the traffic congestion degree robustness parameter, and its value range is [0,a], δ B is the vehicle arrival fluctuation parameter, V Q is the average speed of the arterial road under the uncertain fluctuating traffic flow;
[0023] S23. Robustness constraint conditions: To limit the behavior of the system under uncertain factors and ensure the stability and safety of the system, the red light time R is set to fluctuate within to avoid traffic chaos caused by frequent large-scale switching of the signal light cycle.
[0024] Furthermore, the proposed robustness multi-agent model configuration in step S3 specifically includes the following steps:
[0025] S31. To further improve the robustness of the arterial intersection collaborative control system, a robustness arterial multi-agent (Robustness-Mainline-IntelligentAgent-Actor-Critic, RMIAC) network framework is configured and built; this framework integrates the robust traffic joint traffic state space, action control parameters and robust reward function, enabling each agent to not only autonomously optimize its own decisions, but also comprehensively consider the dynamic changes of adjacent intersections to ensure the optimization of the global traffic flow; the network training adopts the mode of "centralized learning, distributed execution": centralized learning enables the global state and reward information to be centrally processed, while distributed execution ensures that the agents at each intersection can operate independently and respond quickly to emergencies; specifically, in the test stage, each intersection agent will output the optimal control action through its independent Actor network to ensure fast response and effectively cope with traffic flow fluctuations and environmental changes;
[0026] S32. Robust Joint Traffic State Space Definition: To fully consider the traffic flow dynamics at intersections and their correlation with adjacent intersections, the definition of a robust joint traffic state space is proposed. This state space not only includes the traffic characteristics of the local intersection but also incorporates traffic data and environmental perception information from adjacent intersections. Specifically, the state vector of each intersection is composed not only of basic traffic data such as the maximum queue length and vehicle arrival flow rate but also combines real-time environmental perception information such as road conditions, sudden accidents, and weather to form a robust joint traffic state space. Its definition is as follows:
[0027]
[0028] In the formula, represents the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the number of vehicles corresponding to the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the arrival flow rate per unit time of the i-th phase of agent m in the previous control cycle; represents the maximum queue length of the i-th phase of the adjacent agent n of agent m in the previous control cycle, represents the number of vehicles corresponding to the maximum queue length of the i-th phase of the adjacent agent n of agent m in the previous control cycle, represents the arrival flow rate per unit time of the i-th phase of the adjacent agent n of agent m in the previous control cycle; a is a robustness parameter, and ε R,d is the vehicle delay robustness parameter, and its value range is [0, a], and ε R,j is the traffic congestion degree robustness parameter, and its value range is [0, a];
[0029] S33. Robust Action Control Parameter Definition: In terms of action control, by incorporating robustness considerations, a dynamic control parameter definition scheme based on the maximum cycle duration is proposed. Each intersection agent selects the optimal control action (i.e., the green signal ratio of each phase) according to the input joint traffic state vector. This control strategy not only optimizes the local traffic flow but also takes into account the impact of adjacent intersections on the global traffic flow. In the configuration of control parameters, the robust action control parameter is proposed.
[0030]
[0031] In the formula, represents the ratio of the green light duration of the i-th phase of agent m to the cycle duration; represents the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the number of vehicles corresponding to the maximum queue length of agent m in the i-th phase during the previous control period. represents the arrival flow rate per unit time of agent m in the i-th phase during the previous control period; a is the robustness parameter.
[0032] S34. Definition of the robust reward function: In the configuration of the reward function, we introduce a multi-level feedback mechanism, which not only considers the control effect of local intersection agents but also incorporates the state changes of adjacent intersections; this configuration ensures that while optimizing the traffic flow of a single intersection, the system can also comprehensively consider the global traffic conditions and avoid local optimal solutions; the specific robust reward function is defined as follows:
[0033]
[0034] In the formula, represents the arrival traffic volume of agent m in the i-th phase during the previous control period, represents the average vehicle delay of agent m in the i-th phase during the previous control period; represents the arrival traffic volume of agent n in the i-th phase of agent m during the previous control period, represents the average vehicle delay of agent n in the i-th phase of agent m during the previous control period; represents the maximum queue length of agent m in the i-th phase during the previous control period, represents the number of vehicles corresponding to the maximum queue length of agent m in the i-th phase during the previous control period, represents the maximum queue length of the adjacent agent n of agent m in the i-th phase during the previous control period, represents the number of vehicles corresponding to the maximum queue length of the adjacent agent n of agent m in the i-th phase during the previous control period; a is the robustness parameter; S R is the robustness objective function; this reward function considers the comprehensive influence of vehicle delay, traffic flow, and the system robustness objective, and while maximizing the traffic flow, it ensures that the system can still maintain efficient and stable operation in the face of traffic fluctuations and external disturbances.
[0035] Beneficial effects: By adopting a modular multi-agent architecture and the "centralized learning - distributed execution" mode, each intersection agent makes independent decisions while integrating the dynamic data of adjacent intersections through the joint traffic state space, achieving trunk-level global optimization. The vehicle delay and flow data of local and adjacent intersections are integrated into the reward function, avoiding trunk congestion caused by local optimization, and improving the overall traffic efficiency.
[0036] Furthermore, the proposed robust multi-agent deep deterministic policy gradient algorithm (CR-MADDPG) in step S4 specifically includes:
[0037] S41. Propose a policy update formula considering robustness: To achieve stronger robustness, this step proposes a policy update formula based on robustness enhancement, which can optimize the policy selection of the agent considering traffic flow and environmental disturbances. The update formula not only optimizes the policy through traditional gradient information but also introduces a robustness parameter to enhance the system's ability to handle emergencies (such as traffic accidents, weather changes, etc.). The specific update formula is as follows:
[0038]
[0039] In the formula: μ i (s i ∣θ i ) is the policy of agent i in state s i ; α is the learning rate; a is the robustness parameter; is the gradient with respect to the policy network parameter θ i ; ∈ i is the adversarial noise, aiming to enhance the decision-making robustness of the agent in complex and dynamic environments. Through this update method, the agent can optimize the traffic flow while ensuring the stability and robustness of the policy in the face of uncertain factors;
[0040] S42. Propose an agent objective function considering robustness: To better measure the performance of the agent's behavior in terms of robustness, this step introduces a new objective function, which combines the robustness parameter on the basis of the traditional objective function to optimize the agent's decision-making process. Specifically, the objective function not only considers the reward of traffic flow but also adds a robustness adjustment factor based on environmental perception, thereby improving the adaptability of the system in uncertain environments. The objective function is defined as follows:
[0041]
[0042] In the formula: J i (θ i ) is the objective function of agent i; is the expectation of the trajectory τ i ; T is the time termination step; γ is the discount factor; is the reward for taking action in the state; S R is the robustness objective function; a is the robustness parameter. This objective function improves the stability and anti-interference ability of the system by optimizing the long-term reward and considering robustness;
[0043] S43. Propose a loss function considering robustness: To further improve the learning efficiency of multi-agent systems in non-ideal environments, a loss function combined with robustness is proposed. In addition to considering the standard Q-learning error, this loss function also adds a robustness adjustment term, comprehensively considering the different challenges faced by agents in cooperation and competition, and ensuring that the system can operate stably in an uncertain environment. The specific loss function is as follows:
[0044]
[0045] Where:
[0046] In the formula: L i (φ i ) is the loss function of agent i; is the experience replay pool; Q i (s,a∣φ i ) is the value function, where φ i are the network parameters of the value function; y i is the target value; are the parameters of the target network; γ is the discount factor; S j ,a j is a mini-batch of experience data extracted from the experience pool; a is the robustness parameter. By comprehensively considering the cooperation and competition of multiple agents, this loss function improves the robustness and stability of the system in complex and dynamic traffic environments.
[0047] Furthermore, the proposed online learning and adaptive adjustment in step S5 specifically includes the following steps:
[0048] S51. Initialize based on the multi-agent deep deterministic policy gradient algorithm considering robustness (CR-MADDPG) proposed in step S4: Initialize the policy network parameters θ i and the value function network parameters φ i of all agents, and set the target network parameters to be the same as φ i . This initialization step lays the foundation for the learning process of the system and determines that the agents can perform adaptive learning in a dynamic environment. The specific steps are as follows:
[0049] θ i ,φ i Initialization:
[0050] S52. Collect experiences based on the multi-agent deep deterministic policy gradient algorithm considering robustness (CR-MADDPG) proposed in step S4: At each time step t, each agent selects an action according to the current policy This action selection not only considers the current state but also incorporates a robustness parameter to ensure that the agent can cope with traffic flow fluctuations. The experience experienced by each agent is stored in the experience replay pool to provide data support for subsequent learning and parameter updates. The specific steps are as follows:
[0051] Experience collection: Stored in the experience pool
[0052] S53. Based on the multi-agent deep deterministic policy gradient algorithm with robustness consideration (CR-MADDPG) proposed in step S4, randomly sample a small batch of samples from the experience pool for network parameter update: Randomly sample a small batch of experience data from the experience replay pool and update the value function network parameters φ i of the agent, the policy network parameters θ i and the target network parameters During the update process, in addition to traditional policy updates and value function optimization, a robustness objective function S R is introduced as a pre-tuning item to ensure that the agent can maintain high-quality decision-making when facing complex and dynamic traffic flows. This step will continue until the policy network converges, ensuring that the system can adapt to various traffic changes and optimize signal control. The specific steps are as follows:
[0053] Small batch sampling: Update the network parameters φ i , θ i ,
[0054] Beneficial effects: By collecting traffic data in real time through an online learning mechanism and updating network parameters, combined with a robust policy update formula with adversarial noise injection, it can adapt to traffic flow mutations.
[0055] Compared with the prior art, the present invention realizes significant optimization of signal control at arterial intersections through the combination of a multi-agent collaborative architecture and a robustness design. The beneficial effects achieved by the present invention are:
[0056] The construction of the distributed system based on modular agents in the present invention, through the independent decision-making and information interaction mechanism, while ensuring the rapid response of each intersection, enhances the global cooperation ability. At the same time, the present invention innovatively configures robustness, through dynamic traffic flow modeling, an objective function with parameter self-adaptation, and red light time constraints, effectively coping with uncertainties such as vehicle arrival rate fluctuations, weather, and road condition changes. The robustness multi-agent model configured by the present invention integrates traffic state data of adjacent intersections and a multi-level feedback mechanism, fuses local and environmental perception information in the joint traffic state space definition, realizes the optimal balance of global traffic flow, and further adopts a deep reinforcement learning algorithm, introduces robustness parameters and adversarial noise, and improves the decision-making robustness in complex environments by enhancing the stability of policy updates and the anti-interference ability of the loss function. It also supports real-time data collection and network parameter updates through an online learning mechanism, enabling the control strategy to continuously adapt to the dynamically changing traffic environment, improving the trunk line passing efficiency while significantly enhancing the system's adaptability to emergencies and environmental disturbances. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings forming a part of the specification depict embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.
[0058] Referring to the drawings, the present invention can be more clearly understood from the following detailed description, where:
[0059] Figure 1 is a flowchart of the coordinated control method for trunk line intersections considering robustness provided by an embodiment of the present invention;
[0060] Figure 2 is a schematic diagram of the construction of a multi-agent trunk line intersection system provided by an embodiment of the present invention;
[0061] Figure 3 is a schematic diagram of the network framework of the robustness multi-agent model provided by an embodiment of the present invention;
[0062] Figure 4 is a schematic diagram of the optimal trajectory of trunk line vehicles under the coordinated control strategy provided by an embodiment of the present invention;
[0063] Figure 5 is a graph showing the continuous iteration and convergence of the coordinated control strategy provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other. The following embodiments are only used to illustrate the technical solution of the present invention more clearly, and cannot be used to limit the protection scope of the present invention.
[0065] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects.
[0066] As Figures 1 - 5 shown, the present invention provides a coordinated control method for arterial intersections considering robustness, which specifically includes the following steps:
[0067] Step S1, construction of a multi-agent arterial intersection system: Each intersection on the traffic artery is regarded as an agent, and each agent has independent decision-making ability and information processing ability. Information exchange is carried out through a communication network to achieve coordinated control. The agent includes a sensing module, a decision-making module, an execution module, and a communication module; among them, the sensing module collects the vehicle arrival flow rate, lane queue length, and signal light status data of the intersection in real time, performs data cleaning and feature extraction, and provides input for subsequent modeling of the uncertainty fluctuation model; the communication module realizes the interaction of traffic state information between adjacent agents; the execution module is used to adjust the signal light timing, and the decision-making module is used to generate control instructions;
[0068] Step S2, robustness configuration of arterial intersections: Considering the uncertainty of arterial traffic flow, such as fluctuations in vehicle arrival rates and changes in road conditions, a robustness optimization strategy is introduced to model the uncertainty fluctuation model of arterial traffic flow, set robustness objectives, and construct robustness constraint conditions to improve the robustness of the multi-agent arterial intersection system; preferably, a traffic flow model is constructed based on vehicle arrival rate fluctuation parameters, road capacity, road condition parameters, and weather condition parameters, where the road capacity represents the maximum number of vehicles that can pass through the intersection per unit time; set a robustness objective function with the goal of minimizing vehicle delay time and lane queue length, and dynamically adjust the value range of the vehicle delay robustness parameter and the congestion robustness parameter in the objective function; constrain the red light time at the intersection to fluctuate within a preset interval;
[0069] Step S3. Robust multi-agent model configuration: Build a Robustness-Mainline-IntelligentAgent-Actor-Critic (RMIAC) network framework through the traffic flow uncertainty fluctuation model, robust objective setting, and robust constraint conditions in Step S2, and respectively define and integrate the robust joint traffic state space, action control parameters, and robust reward function; preferably, define the joint traffic state space, including the following parameters of the local intersection and adjacent intersections in the previous control cycle: the ratio of the maximum queue length of each phase to the corresponding number of vehicles, the ratio of the arrival flow rate per unit time to the road capacity, and generate a robust joint state vector by combining the vehicle delay robust parameter and congestion robust parameter; define the action control parameter as the ratio of the green light duration of each phase to the cycle duration, which is calculated from the joint state vector and the robust parameter; configure the reward function, and comprehensively consider the vehicle delay, arrival flow, and queue characteristics of the local and adjacent intersections, and introduce the robust objective function as an adjustment factor;
[0070] Step S4. Trunk intersection collaborative control strategy configuration: Use the robust joint traffic state space and action control parameters as the input of the deep reinforcement learning algorithm to design the observation space and action selection strategy of the intelligent agent, train the intelligent agent using the deep reinforcement learning algorithm, and construct a multi-agent deep deterministic policy gradient algorithm considering robustness (CR-MADDPG) to select an appropriate control strategy according to the current traffic state, so that each intelligent agent, while pursuing its own maximum benefit, considers the impact on the overall traffic flow of the trunk line. Through information exchange and cooperation between intelligent agents, realize the collaborative control of trunk intersections, and vehicles drive according to the optimal trajectory; preferably, update the policy network parameters of the intelligent agent in a centralized learning manner and output control actions in a distributed execution manner; introduce adversarial noise and robust parameters in the policy update to generate a policy gradient; when defining the objective function, fuse the reward function and the robust objective function, and add a robust adjustment term to the loss function to define the intelligent agent objective function considering robustness, so as to drive the intelligent agent to learn the robust strategy and achieve global robustness optimization. The decision-making module can be used to execute the control instructions generated by the multi-agent deep deterministic policy gradient algorithm considering robustness (CR-MADDPG);
[0071] Step S5, Online learning and adaptive adjustment: The agent collects experience based on real-time traffic data, conducts online learning, randomly selects a batch of samples to update network parameters, continuously optimizes the control strategy, can adaptively adjust control parameters according to the changes in traffic flow, and continuously improves the adaptability and robustness of the control system. Preferably, traffic state data is collected in real time and stored in the experience replay pool, and samples are randomly selected to batch update network parameters; the robustness parameter and adversarial noise intensity are dynamically adjusted according to traffic flow fluctuations until the policy network converges.
[0072] In summary, through the proposed robustness configuration for arterial intersections, the robustness-mainline-intelligent-agent-actor-critic (RMIAC) network framework and the multi-agent deep deterministic policy gradient algorithm considering robustness provide a reliable and practical means for the coordinated control of arterial intersections when facing large fluctuations in traffic flow.
[0073] Further, the construction of the multi-agent arterial intersection system in step S1 specifically includes:
[0074] Step S11, Construction of a single intersection agent: Each intersection is regarded as an independent agent, and each agent has independent decision-making ability and information processing ability;
[0075] Step S12, Configuration of the agent software structure: including a perception module, a decision-making module, an execution module, and a communication module; configure the perception module to collect traffic data, configure the decision-making module to generate control instructions based on the multi-agent deep deterministic policy gradient algorithm considering robustness, configure the execution module to adjust the signal light timing, and configure the communication module to transmit the status information of adjacent intersections;
[0076] Step S13, Establishment of a coordinated control mechanism: Use deep learning algorithms to train agents so that they can select appropriate coordinated control strategies according to the current traffic state;
[0077] Step S14, Configuration of system scalability: Adopt modular configuration, and each agent is used as an independent module for easy expansion.
[0078] Beneficial effects: Adopting a modular agent design supports the rapid addition or replacement of intersection nodes, and for newly added intersections, only the local agent network needs to be initialized.
[0079] Further, the robustness configuration of arterial intersections in step S2 specifically includes:
[0080] Step S21: Modeling the Uncertain Fluctuations of Trunk Traffic Flow: Considering the impacts of vehicle arrival rate, road conditions, weather conditions, and real-time environmental factors, a robust modeling strategy based on deep reinforcement learning is constructed, and the negative binomial moving average randomization algorithm is introduced. Model the uncertain fluctuations of trunk traffic flow through the following formula:
[0081]
[0082] Among them, Q m is the maximum trunk traffic flow; k is the number of vehicles arriving per unit time; p k is the probability of k vehicles arriving per unit time; β is the difference between the maximum trunk traffic flow Q m and the number of vehicles arriving per unit time; δ B is the vehicle arrival fluctuation parameter; δ W is the weather condition parameter; δ C is the road condition change parameter; δ Env is the environmental perception information parameter;
[0083] Step S22: Setting the Robustness Objective. Set minimizing vehicle delay and traffic congestion level as the robustness objective, and dynamically adjust the parameters of the objective function according to real-time perception data. The robustness objective function is:
[0084]
[0085] Among them, Q is the uncertain fluctuating traffic flow, Q m is the maximum trunk traffic flow, R is the red light time at the intersection, a is the robustness parameter, ε R,d is the vehicle delay robustness parameter, ε R,j is the traffic congestion level robustness parameter, V Q is the average trunk speed under the uncertain fluctuating traffic flow;
[0086] Step S23: Robustness Constraint Conditions. Set the red light time R to fluctuate within the range of .
[0087] Furthermore, the configuration of the robust multi-agent model in Step S3 specifically includes:
[0088] Step S31: Build a robust trunk multi-agent network framework, update the policy network parameters of the agents in a centralized learning manner, and output control actions through a distributed execution method;
[0089] Step S32: Definition of the Robust Joint Traffic State Space. The robust joint traffic state space is defined as:
[0090]
[0091] Among them, represents the maximum queue length of agent m in the i-th phase within the previous control period t; represents the number of vehicles corresponding to the maximum queue length of agent m in the i-th phase within the previous control period t; represents the arrival flow rate per unit time of agent m in the i-th phase within the previous control period t, and C is the road capacity; represents the maximum queue length of the adjacent agent n of agent m in the i-th phase within the previous control period t; represents the number of vehicles corresponding to the maximum queue length of the adjacent agent n of agent m in the i-th phase within the previous control period t; represents the arrival flow rate per unit time of the adjacent agent n of agent m in the i-th phase within the previous control period t; a is a robustness parameter; ε R,d is the robustness parameter of vehicle delay; ε R,j is the robustness parameter of traffic congestion degree;
[0092] Step S33, Definition of robust action control parameter, the robust action control parameter is defined as:
[0093]
[0094] Among them, represents the ratio of the green light duration to the cycle duration of the i-th phase of agent m, and a is a robustness parameter;
[0095] Step S34, Definition of robust reward function, the robust reward function is defined as:
[0096]
[0097] Among them, represents the arrival flow of agent m in the i-th phase within the previous control period t; represents the average vehicle delay of agent m in the i-th phase within the previous control period t; represents the arrival flow of agent n in the i-th phase within the previous control period t; represents the average vehicle delay of agent n in the i-th phase within the previous control period t; S R is the robustness objective function.
[0098] Beneficial effects: When designing the robust action control parameter, a dynamic control scheme for the maximum cycle duration is introduced, combined with the red light time fluctuation constraint in step S23 effectively avoids frequent large-scale switching of the signal cycle.
[0099] Furthermore, the robust multi-agent deep deterministic policy gradient algorithm constructed in step S4 specifically includes:
[0100] Step S41: Construct a policy update formula considering robustness:
[0101]
[0102] where μ i (s i ∣θ i ) is the policy of agent i in state s i ; α is the learning rate; a is the robustness parameter, is the gradient with respect to the policy network parameter θ i ; ∈ i is the adversarial noise;
[0103] Step S42: Construct an agent objective function considering robustness:
[0104]
[0105] where J i (θ i ) is the objective function of agent i; is the expectation of the trajectory τ i ; T is the time termination step, and γ is the discount factor; is the reward for taking action in state ; S R is the robustness objective function;
[0106] Step S43: Construct a loss function considering robustness:
[0107]
[0108] where L i (φ i ) is the loss function of agent i; is the experience replay pool, Q i (s,a∣φ i ) is the value function; φ i is the value function network parameter; y i is the target value; is the parameter of the target network; γ is the discount factor; s j and a j are the batch of experience data extracted from the experience replay pool.
[0109] Furthermore, the online learning and adaptive adjustment in step S5 specifically include:
[0110] Step S51: Initialize the policy network parameters θ of all agents i and the value function network parameters φ i , and set the target network parameters to be the same as φ i ;
[0111] Step S52: Each agent selects an action according to the current policy and stores the experience experienced by each agent into the experience replay pool ;
[0112] Step S53: Randomly sample a small batch of experience data from the experience replay pool and update the value function network parameters φ i , the policy network parameters θ i and the target network parameters of the agent based on these experiences until the policy network converges.
[0113] Example 1
[0114] Figure 1 is a flowchart of a coordinated control method for arterial intersections considering robustness in Example 1 of the present invention. This flowchart only shows the logical sequence of the method described in this example. On the premise of non - conflict, in other possible embodiments of the present invention, the steps shown or described can be completed in a different Figure 1 order than that shown.
[0115] This example is a typical implementation mode of the present invention, providing a coordinated control method for arterial intersections considering robustness. This method can be applied to terminals and can be executed by an electronic terminal. The electronic terminal can be implemented in software and / or hardware and can be integrated into the terminal. For example: any smart phone, tablet computer or computer device with communication functions. As Figure 1 shown, the method of this example specifically includes the following steps:
[0116] Step S1: Construction of a multi - agent arterial intersection system; Each intersection on the arterial is regarded as an agent, and each agent has independent decision - making ability and information - processing ability; Agents exchange information through a communication network to achieve coordinated control, specifically as follows:
[0117] (11) Construction of single intersection agent: Each intersection is regarded as an independent agent with independent decision-making ability and information processing ability, responsible for monitoring and managing the traffic conditions at its own intersection, including vehicle arrivals, departures, lane occupancy, signal states, etc., capable of making appropriate control decisions based on its own state and the surrounding environment, and exchanging information with other agents to achieve coordinated control between arterial intersections;
[0118] (12) Agent software structure configuration: It includes a perception module, a decision-making module, an execution module, and a communication module; The perception module is responsible for collecting traffic data; The decision-making module makes decisions based on the proposed algorithm; The execution module is responsible for controlling traffic facilities such as signal lights; The communication module is responsible for exchanging information with other agents;
[0119] (13) Establishment of coordinated control mechanism: The goal of coordinated control is to enable all intersections on the arterial road to achieve the optimal traffic state on the basis of achieving robustness; The proposed deep learning algorithm is used to train the agents so that they can select appropriate coordinated control strategies according to the current traffic state;
[0120] (14) System scalability configuration: Adopt modular configuration, each agent is an independent module, with good scalability, and can be easily added, deleted, or replaced as the city develops;
[0121] Step S2, Robustness configuration of arterial intersections: Considering the uncertainties of arterial traffic flow, such as fluctuations in vehicle arrival rates, changes in road conditions, etc., introduce a robustness optimization strategy; By proposing the modeling of uncertainties and fluctuations in arterial traffic flow, setting the robustness objectives of signal coordinated control at arterial intersections, and robustness constraint conditions, improve the robustness of the control system, as follows:
[0122] (21) Modeling of uncertainties and fluctuations in arterial traffic flow; To effectively cope with the uncertainties and fluctuations in arterial traffic flow, consider the impacts of real-time environmental factors such as vehicle arrival rates, road conditions, weather conditions, etc.; Propose a robustness modeling strategy based on deep reinforcement learning and introduce a negative binomial moving average randomization algorithm By dynamically adjusting the model parameters combined with real-time environmental perception information, enhance the robustness of the system; The specific traffic flow modeling formula is:
[0123]
[0124] In the formula: Q m is the maximum arterial traffic flow, k is the number of vehicles arriving per unit time, p k is the probability of k vehicles arriving per unit time, β is the difference between the maximum arterial traffic flow Q m and the number of vehicles arriving per unit time, δ Bis the vehicle arrival fluctuation parameter, with a value range of [0, 2], δ W is the weather condition parameter, with a value range of [0, 1], δ C is the road condition change parameter, with a value range of [0, 1], δ Env is the environmental perception information parameter, indicating the impact of real-time environmental changes including road surface conditions, accidents, etc., with a value range of [0, 1];
[0125] (22) Robustness objective setting: Set minimizing vehicle delay as the main robustness objective and minimizing traffic congestion as the secondary objective, and propose the robustness objective function S R , to ensure that the performance of the system control strategy can still be maintained within an acceptable range in the presence of uncertain factors:
[0126]
[0127] In the formula: Q is the uncertain fluctuating traffic flow, Q m is the maximum traffic flow of the main line, R is the red light time at the intersection, a is the robustness parameter, ε R,d is the vehicle delay robustness parameter, with a value range of [0, a], ε R,j is the traffic congestion degree robustness parameter, with a value range of [0, a], δ B is the vehicle arrival fluctuation parameter, V Q is the average speed of the main line under the uncertain fluctuating traffic flow;
[0128] (23) Robustness constraint conditions: To limit the behavior of the system under uncertain factors and ensure the stability and safety of the system, set the red light time R to fluctuate within , to avoid traffic chaos caused by frequent large-scale switching of the signal light cycle;
[0129] Step S3, Robust multi-agent model configuration; Build the network framework of the Robustness-Mainline-IntelligentAgent-Actor-Critic (RMIAC), and propose the definition of the robust joint traffic state space, the definition of action control parameters, and the definition of the robust reward function, as follows:
[0130] (31) Construction of the Robustness-Mainline-Intelligent Agent-Actor-Critic (RMIAC) network framework; construct the Robustness-Mainline-Intelligent Agent-Actor-Critic (RMIAC) network framework, including the definition of the robust traffic joint traffic state space, the definition of action control parameters, and the definition of the robust reward function; during training, based on the proposed algorithm, the "centralized learning, distributed execution" method is used for agent training; during testing, the intersection agent outputs the optimal control action through its independent Actor network to ensure fast response;
[0131] (32) Definition of the robust joint traffic state space; the ratio of the maximum queue length of each phase lane group in the previous signal cycle at the intersection to the number of vehicles in that lane group is used as the normalized queue feature, and the arrival flow rate of each phase lane group at the intersection per unit time in the previous signal cycle is used as the normalized flow rate feature. At the same time, the state features of the local intersection agent m and the adjacent intersection agent n are considered, and the joint feature is used as the state representation of the local agent. Considering robustness, the joint traffic state
[0132] In the formula, represents the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the number of vehicles corresponding to the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the arrival flow rate per unit time of the i-th phase of agent m in the previous control cycle; represents the maximum queue length of the i-th phase of the adjacent agent n of agent m in the previous control cycle, represents the number of vehicles corresponding to the maximum queue length of the i-th phase of the adjacent agent n of agent m in the previous control cycle, represents the arrival flow rate per unit time of the i-th phase of the adjacent agent n of agent m in the previous control cycle; a is the robustness parameter, ε R,d is the vehicle delay robustness parameter, and its value range is [0, a], ε R,j is the traffic congestion degree robustness parameter, and its value range is [0, a];
[0133] (33) Definition of the robust action control parameters; select the maximum cycle length of the key intersection as the cycle length of the mainline system. The local agent selects the green signal ratio of each phase of the local intersection according to the input joint traffic state vector. Considering robustness, the robust action control parameters
[0134]
[0135] In the formula, represents the ratio of the green light duration of the i-th phase of agent m to the cycle length; represents the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the number of vehicles corresponding to the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the arrival flow rate per unit time of the i-th phase of agent m in the previous control cycle; a is a robustness parameter;
[0136] (34) Definition of the robust reward function; Considering the reward value of the local intersection agent m, and at the same time considering the reward value of the adjacent intersection agent n, considering robustness, a robust reward function is proposed
[0137]
[0138] In the formula, represents the arrival flow of the i-th phase of agent m in the previous control cycle, represents the average vehicle delay of the i-th phase of agent m in the previous control cycle; represents the arrival flow of the i-th phase of agent n in the previous control cycle, represents the average vehicle delay of the i-th phase of agent n in the previous control cycle; represents the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the number of vehicles corresponding to the maximum queue length of the i-th phase of agent m in the previous control cycle, represents the maximum queue length of the i-th phase of the adjacent agent n of agent m in the previous control cycle, represents the number of vehicles corresponding to the maximum queue length of the i-th phase of the adjacent agent n of agent m in the previous control cycle; a is a robustness parameter; S R is the robustness objective function;
[0139] Step S4, Trunk intersection collaborative control strategy configuration; Use the deep reinforcement learning algorithm to train the agents, and propose a multi-agent deep deterministic policy gradient algorithm considering robustness (CR-MADDPG), so that it can select appropriate control strategies according to the current traffic state, so that each agent, while pursuing its own maximum interests, considers the impact on the overall trunk traffic flow; Through information exchange and cooperation between agents, the collaborative control of trunk intersections is realized, specifically as follows:
[0140] (41) A strategy update formula considering robustness is proposed:
[0141] To improve the decision-making ability of agents in complex traffic environments, a robustness strategy update formula is proposed; based on the traditional policy gradient algorithm, this formula adds a robustness factor and adversarial noise to enhance the adaptability and stability of the strategy in uncertain and fluctuating traffic flows; the specific formula is as follows:
[0142]
[0143] In the formula: μ i (s i |θ i ) is the strategy of agent i in state s i ; α is the learning rate; a is the robustness parameter; is the gradient with respect to the policy network parameter θ i ; ∈ i is the adversarial noise, used to counter external uncertainties and environmental disturbances to ensure the robustness of the policy update;
[0144] (42) A goal function for agents considering robustness is proposed:
[0145] To effectively guide the decision-making of agents in complex and dynamic traffic environments, a goal function for agents with enhanced robustness is proposed; this goal function comprehensively considers the long-term rewards of agents, the stability and robustness of traffic flows, aiming to maintain better performance in the face of uncertain factors such as traffic flow fluctuations and equipment failures; the specific formula is as follows:
[0146]
[0147] In the formula: J i (θ i ) is the goal function of agent i; is the expectation of the trajectory τ i ; T is the time termination step; γ is the discount factor; is the reward for taking action in the state; S R is the robustness goal function; a is the robustness parameter; this goal function improves the stability and anti-interference ability of the system by optimizing long-term rewards and considering robustness;
[0148] (43) A loss function considering robustness is proposed:
[0149] To further improve the learning efficiency of multi-agent systems in non-ideal environments, a loss function combined with robustness is proposed. In addition to considering the standard Q-learning error, this loss function also adds a robustness adjustment term, comprehensively considering the different challenges faced by agents in cooperation and competition, and ensuring the stable operation of the system in uncertain environments. The specific loss function is as follows:
[0150]
[0151] Where:
[0152] In the formula: L i (φ i ) is the loss function of agent i; is the experience replay pool; Q i (s,a∣φ i ) is the value function, where φ i are the parameters of the value function network; y i is the target value; are the parameters of the target network; γ is the discount factor; S j ,a j is the mini-batch of experience data extracted from the experience pool; a is the robustness parameter. By comprehensively considering the cooperation and competition of multiple agents, this loss function improves the robustness and stability of the system in complex and dynamic traffic environments;
[0153] Step S5, Online learning and adaptive adjustment; During the actual operation, agents can perform online learning based on real-time traffic data and continuously optimize the control strategy. At the same time, agents can adaptively adjust the control parameters according to the changes in traffic flow to improve the adaptability and robustness of the control system, as follows:
[0154] (51) Initialize based on the robust multi-agent deep deterministic policy gradient algorithm (CR-MADDPG) proposed in step S4:
[0155] Initialize the policy network parameters θ i and the value function network parameters φ i , and set the target network parameters to be the same as φ i ; This initialization step lays the foundation for the learning process of the system and determines that agents can perform adaptive learning in dynamic environments. The specific steps are as follows:
[0156] θ i ,φ i Initialization;
[0157] (52) Experience collection is carried out based on the robust multi-agent deep deterministic policy gradient algorithm (CR-MADDPG) proposed in step S4:
[0158] At each time step t, each agent selects an action according to the current policy This action selection not only considers the current state but also incorporates a robustness parameter to ensure that the agent can cope with traffic flow fluctuations; the experience experienced by each agent is stored in the experience replay pool to provide data support for subsequent learning and parameter updates; the specific steps are as follows:
[0159] Experience collection: Stored in the experience pool
[0160] (53) Based on the robust multi-agent deep deterministic policy gradient algorithm (CR-MADDPG) proposed in step S4, randomly sample a small batch of samples from the experience pool for network parameter update:
[0161] Randomly sample a small batch of experience data from the experience replay pool and update the value function network parameters φ i of the agent, the policy network parameters θ i and the target network parameters During the update process, in addition to traditional policy updates and value function optimization, a robustness objective function S R is introduced as a pre-tuning item to ensure that the agent can maintain high-quality decision-making when facing complex dynamic traffic flows; this step will continue until the policy network converges to ensure that the system can adapt to various traffic changes and optimize signal control; the specific steps are as follows:
[0162] Small batch sampling: Update the network parameters φ_i, θ_i, φ_i^.
[0163] In summary, the present invention constructs a distributed system based on modular agents. Through an independent decision-making and information interaction mechanism, while ensuring the rapid response of each intersection, it improves the global coordination ability, innovates the robustness configuration, and effectively deals with uncertainty factors such as vehicle arrival rate fluctuations, weather, and road condition changes through dynamic traffic flow modeling, a parameter-adaptive objective function, and red light time constraints;
[0164] Furthermore, the robust multi-agent model configured by the present invention integrates the traffic state data of adjacent intersections and a multi-level feedback mechanism, fuses local and environmental perception information in the definition of the joint traffic state space, realizes the optimal balance of global traffic flow, adopts a deep reinforcement learning algorithm, introduces robustness parameters and adversarial noise, and improves the decision-making robustness in complex environments by enhancing the stability of policy updates and the anti-interference ability of the loss function;
[0165] In addition, the online learning mechanism of the present invention supports real-time data collection and network parameter update, enables the control strategy to continuously adapt to the dynamically changing traffic environment, and significantly enhances the adaptability of the system to emergencies and environmental disturbances while improving the trunk line traffic efficiency.
[0166] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A multi-agent arterial intersection signal cooperative control method considering robustness, characterized in that The steps are as follows: Step S1, Construction of multi-agent arterial intersection system: Each intersection on the traffic artery is regarded as an agent. The agent obtains real-time traffic data in real time, and performs data cleaning and feature extraction to provide input for subsequent modeling of the uncertainty fluctuation model; Step S2, Robustness configuration of arterial intersections: Model the uncertainty fluctuation model of arterial traffic flow, set the robustness objective, and construct the robustness constraint conditions; Step S3, Robust multi-agent model configuration: Build a robust arterial multi-agent network framework through the traffic flow uncertainty fluctuation model, robustness objective setting, and robustness constraint conditions in Step S2, and define the robust joint traffic state space, action control parameters, and robust reward function; Step S4, Collaborative control strategy configuration of arterial intersections: Use the robust joint traffic state space and action control parameters as the input of the deep reinforcement learning algorithm to design the observation space and action selection strategy of the agent. Use the deep reinforcement learning algorithm to perform multi-agent collaborative training on the agent, construct a multi-agent deep deterministic policy gradient algorithm CR-MADDPG considering robustness, and embed the robust reward function to drive the agent to learn the robust strategy and achieve global robustness optimization; Step S5, Online learning and adaptive adjustment: Use the CR-MADDPG algorithm in Step S4 as the initial strategy for online learning, collect experience according to real-time traffic data, and randomly extract a batch of samples to update the network parameters of the agent for fine-tuning the strategy of online learning.
2. The multi-agent trunk intersection signal cooperative control method considering robustness according to claim 1, characterized in that The construction of the multi-agent arterial intersection system in Step S1 specifically includes: Step S11, Construction of a single intersection agent: Each intersection is regarded as an independent agent, and each agent has independent decision-making ability and information processing ability; Step S12, Agent software structure configuration: The agent includes a perception module, a decision-making module, an execution module, and a communication module; Configure the perception module to collect traffic data, configure the decision-making module to generate control instructions based on the multi-agent deep deterministic policy gradient algorithm considering robustness, configure the execution module to adjust the signal light timing, and configure the communication module to transmit the state information of adjacent intersections; Step S13, Establishment of collaborative control mechanism: Use the deep learning algorithm to train the agent so that the agent can select an appropriate collaborative control strategy according to the current traffic state; Step S14, System scalability configuration: Adopt modular configuration, and each agent is used as an independent module for easy expansion.
3. The multi-agent trunk intersection signal cooperative control method considering robustness according to claim 1, characterized in that, The robustness configuration of arterial intersections in Step S2 specifically includes: Step S21, Modeling of Uncertainty Fluctuations in Trunk Traffic Flow: Considering the impacts of vehicle arrival rate, road conditions, weather conditions, and real-time environmental factors, a robust modeling strategy based on deep reinforcement learning is constructed, and the negative binomial moving average randomization algorithm is introduced. Model the uncertainty fluctuations in trunk traffic flow through the following formula: Among them, Q m is the maximum traffic flow on the main line; k is the number of vehicles arriving per unit time; p k is the probability of k vehicles arriving per unit time; β is the difference between the maximum traffic flow Q m on the main line and the vehicles arriving per unit time; δ B is the vehicle arrival fluctuation parameter; δ W is the weather condition parameter; δ C is the road condition change parameter; δ Env is the environmental perception information parameter; Step S22, Robustness objective setting: Set minimizing vehicle delay and traffic congestion degree as the robustness objective, and dynamically adjust the parameters of the objective function according to real-time perception data. The robustness objective function is: Among them, Q is the uncertain fluctuating traffic flow, Q m is the maximum traffic flow of the arterial road, R is the red light time at the intersection, a is the robustness parameter, ε R,d is the robust parameter of vehicle delay, ε R,j is the robust parameter of traffic congestion degree, V Q is the average speed of the arterial road under the uncertain fluctuating traffic flow; Step S23, robustness constraint condition, set the red light time R to fluctuate within the range.
4. The multi-agent trunk intersection signal cooperative control method considering robustness according to claim 1, characterized in that The robust multi-agent model configuration in Step S3 specifically includes: Step S31: Build a robust trunk multi-agent network framework, update the policy network parameters of the agent in a centralized learning manner, and output control actions through a distributed execution manner; Step S32, Robust Joint Traffic State Space Definition, where the robust joint traffic state space is defined as: Among them, represents the maximum queue length of agent m in the i-th phase during the previous control period t; represents the number of vehicles corresponding to the maximum queue length of agent m in the i-th phase during the previous control period t; represents the arrival flow rate per unit time of agent m in the i-th phase during the previous control period t, where C is the road capacity; represents the maximum queue length of the adjacent agent n of agent m in the i-th phase during the previous control period t; represents the number of vehicles corresponding to the maximum queue length of the adjacent agent n of agent m in the i-th phase during the previous control period t; represents the arrival flow rate per unit time of the adjacent agent n of agent m in the i-th phase during the previous control period t; a is a robustness parameter; ε R,d is the vehicle delay robustness parameter; ε R,j is the traffic congestion degree robustness parameter; Step S33, definition of robust motion control parameters, where the robust motion control parameters are defined as: Among them, represents the ratio of the green light duration of the i-th phase of the agent m to the cycle duration, and a is the robustness parameter; Step S34, Robust reward function definition, the robust reward function is defined as: Among them, represents the arrival flow of agent m at the i-th phase in the previous control period t; represents the average vehicle delay of agent m at the i-th phase in the previous control period t; represents the arrival flow of agent n at the i-th phase in the previous control period t; represents the average vehicle delay of agent n at the i-th phase in the previous control period t; S R is the robustness objective function.
5. The multi-agent trunk intersection signal cooperative control method considering robustness according to claim 1, characterized in that The robust multi-agent deep deterministic policy gradient algorithm constructed in step S4 specifically includes: Step S41: Construct a policy update formula considering robustness: where μ i (s i ∣θ i ) is the policy of agent i in state s i ; α is the learning rate; a is the robustness parameter, is the gradient with respect to the policy network parameter θ i ; ∈ i is the adversarial noise; Step S42: Construct an agent objective function considering robustness: Among them, J i (θ i ) is the objective function of agent i; is the expectation of the trajectory τ i ; T is the time termination step, and γ is the discount factor; is the reward for taking the action in the state ; S R is the robustness objective function; Step S43: Construct a loss function considering robustness: Among them, L i (φ i ) is the loss function of agent i; is the experience replay pool, Q i (s,a∣φ i ) is the value function; φ i are the parameters of the value function network; y i is the target value; are the parameters of the target network; γ is the discount factor; s j and a j are the batch of experience data extracted from the experience replay pool.
6. The multi-agent trunk intersection signal collaborative control method considering robustness according to claim 1, characterized in that The online learning and adaptive adjustment in step S5 specifically include: Step S51: Initialize the policy network parameters θ of all agents i and the value function network parameters φ i , and set the target network parameters to be the same as φ i ; Step S52: Each agent selects an action according to the current policy Select an action Store the experience experienced by each agent into the experience replay pool ; Step S53: Randomly sample a small batch of experience data from the experience replay pool and update the value function network parameters φ i of the agent, the policy network parameters θ i and the target network parameters based on these experiences until the policy network converges.
Citation Information
Patent Citations
Multi-agent reinforcement learning rolling scheduling method and device, equipment and storage medium
CN115310775A
Active power distribution network fault first-aid repair method and system considering road-source uncertainty
CN116933940A
Multi-intersection signal cooperative control method based on deep reinforcement learning
CN119252051A
Interval and bounded probability mixed uncertainty-based mechanical arm robustness optimization design method
WO2020228310A1