A Distributed Multi-Agent Cooperative Traffic Signal Control Method
Through the distributed multi-agent reinforcement learning framework and adaptive entropy regularization mechanism, the problems of insufficient exploration efficiency, collaboration capabilities and scalability in the existing technology are solved, and efficient signal control and effective collaboration in large-scale road networks are realized.
Patent Information
- Application Number
- CN202510353897.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing multi-agent reinforcement learning methods have shortcomings in exploring efficiency, collaboration capabilities and scalability, which makes it difficult to find the global optimal signal control strategy, it is difficult to achieve effective collaboration in complex traffic environments, and the computing complexity in large-scale road networks has increased sharply.
The distributed multi-agent reinforcement learning framework is adopted to treat each traffic light-controlled intersection in the urban road network as an independent agent. Through local information sharing and adaptive individual entropy and joint entropy regularization mechanisms, exploration efficiency and collaboration capabilities are improved, and computational complexity and communication overhead are reduced.
By introducing adaptive individual entropy and joint entropy regularization mechanisms, exploration efficiency and collaboration capabilities are improved, efficient signal control in large-scale road networks is realized, and computing complexity and communication overhead are reduced.
Smart Images

Figure CN119889066B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent traffic control, and particularly relates to a distributed multi-agent traffic signal cooperative control method. Background Art
[0002] Urban traffic congestion has become a severe challenge faced by major cities around the world. Traditional traffic signal control methods, such as fixed-time control and inductive control, are unable to cope with complex and changing traffic flows. Fixed-time control relies on a pre-set fixed timing plan and cannot be dynamically adjusted according to real-time traffic conditions, resulting in low intersection passing efficiency, especially in the case of large fluctuations in traffic demand. Although inductive control can make certain adjustments based on the information of vehicle detectors, its adjustment range is limited, and it lacks a macroscopic grasp of the overall traffic conditions of the road network.
[0003] With the rise of artificial intelligence and reinforcement learning technologies, more and more research has been devoted to applying reinforcement learning to traffic signal control. The single-agent reinforcement learning method regards a single intersection as an independent agent and optimizes the signal timing of this intersection through learning. However, the urban traffic network is a complex system, and there are close interdependencies between intersections. The single-agent method is difficult to capture this global correlation and is prone to the phenomenon of "treating the head when the head aches and treating the foot when the foot aches", that is, optimizing a single intersection may deteriorate the traffic conditions of adjacent intersections.
[0004] To overcome the limitations of the single-agent method, the multi-agent reinforcement learning (MARL) method has emerged. The multi-agent reinforcement learning (MARL) method models each intersection as an independent agent and optimizes the traffic performance of the entire road network through the cooperation and competition between agents. Currently, some multi-agent reinforcement learning (MARL) methods have been applied to traffic signal control, such as:
[0005] Independent Learning: Each agent independently learns its own strategy, completely ignoring the existence of other agents. The advantage of this method is simple implementation, but due to the lack of coordination between agents, the overall performance is usually poor.
[0006] Centralized Learning: A central controller is used to learn the joint strategy of all agents. This method can theoretically fully consider the mutual influence between agents, but it has serious problems of high computational complexity and poor scalability, and it is difficult to be applied to large-scale road networks.
[0007] Value Decomposition Networks (VDN): Decompose the global Q-value function into the sum of individual agent Q-value functions. This method alleviates the computational pressure of centralized learning to a certain extent, but the way of decomposing the value function limits the expressive power of the model.
[0008] Multi-Agent Attention (MAA): Selectively focus on the state and action information of other agents through the attention mechanism. This method can improve the cooperation efficiency among agents, but the design and training of the attention mechanism are relatively complex, and the computational cost is large.
[0009] Although the above multi-agent reinforcement learning (MARL) methods have improved the traffic signal control effect to a certain extent, there are still the following problems to be solved urgently:
[0010] Insufficient exploration: Reinforcement learning needs to balance exploration (trying new actions) and exploitation (selecting the known best actions). Traditional multi-agent reinforcement learning (MARL) methods usually adopt ε-greedy or Boltzmann exploration strategies, and the exploration efficiency of these strategies is relatively low, which easily leads to the algorithm converging to a local optimal solution prematurely and being unable to find the global optimal signal control strategy.
[0011] Difficult cooperation: Effective cooperation among agents is the key to excellent performance of multi-agent reinforcement learning (MARL). However, in a complex traffic environment, it is difficult for agents to spontaneously establish an effective cooperation mechanism, especially in the absence of explicit communication. How to design a mechanism to promote cooperation among agents is an important research direction in the field of multi-agent reinforcement learning (MARL).
[0012] Poor scalability: The computational complexity of many existing multi-agent reinforcement learning (MARL) methods increases sharply when the number of agents increases, making it difficult to apply to large-scale and high-density urban road networks. How to reduce the computational complexity and communication overhead while ensuring the algorithm performance is the key for multi-agent reinforcement learning (MARL) methods to move towards practical applications. Summary of the Invention
[0013] In view of the above deficiencies in the prior art, the present invention provides a distributed multi-agent traffic signal cooperative control method to solve the following problems: Exploration deficiency problem: The exploration strategies (such as ε-greedy) of existing reinforcement learning methods are less efficient, easily leading to the algorithm falling into local optima and being unable to find the global optimal signal control strategy; Collaboration difficulty problem: In a multi-agent environment, it is difficult for agents to spontaneously establish an effective collaboration mechanism, resulting in the overall traffic efficiency not being maximally improved; Scalability problem: For many existing multi-agent reinforcement learning methods, the computational complexity increases sharply when the number of agents increases, making it difficult to apply to large-scale road networks.
[0014] To achieve the above objectives, the technical solution adopted by the present invention is: A distributed multi-agent traffic signal cooperative control method, comprising the following steps:
[0015] S1. Initialization stage: Initialize the urban road network environment, construct an adjacency list according to the topological structure of the urban road network, and regard each traffic signal control intersection in the urban road network as an independent agent, and initialize each agent. Among them, the agents share local information through the adjacency list.
[0016] S2. Environment interaction stage: Independently obtain the local observation state of each agent at its corresponding intersection.
[0017] S3. Decision-making stage: Based on the local observation states obtained by each agent, obtain the action probability distribution, and the agent selects an action through sampling according to the action probability distribution. Among them, the action probability distribution represents the probability that the agent selects each available traffic signal phase.
[0018] S4. Execution stage: Each agent sends the selected action to the traffic simulator to update the state of the traffic signal and simulate the movement of vehicles, and the traffic simulator calculates and returns the agent information at the next time step. Among them, the agent information includes the reward of each agent and the next local observation state of each agent.
[0019] S5. Information sharing stage: Each agent sends the action probability distribution to all its neighbor agents. Among them, only the action probability distribution is shared between each agent, and the neighbor agents are determined by the adjacency list.
[0020] S6. Learning stage: Update the parameters according to the action probability distribution obtained in S3, the reward of the traffic simulator and the next local observation state obtained in S4, and the action probability distribution received from neighboring agents obtained in S5. Update the Actor network, Critic network, and the adaptive temperature coefficients alpha and beta by calculating the individual entropy and joint entropy. Here, the individual entropy is used to measure the randomness of the action selection of a single agent, and the joint entropy is used to measure the joint distribution of actions between adjacent agents;
[0021] S7. Loop iteration: Repeat S2 to S6 until the preset number of training rounds is reached. The agent learns the optimal signal control decisions under different traffic states and collaborates with all its neighboring agents to complete the cooperative control of traffic signals.
[0022] Further, the agent in S1 includes:
[0023] The Actor network is used to generate an action probability distribution according to the local observation state observed by the agent. The local observation state includes the lane queue lengths, vehicle average speeds, and vehicle numbers of each approach at the current intersection, as well as the phase state of the current traffic signal and the duration of the phase;
[0024] The Critic network is used to evaluate the value of the current state according to the local observation state observed by the agent, that is, to predict the long-term cumulative return of the current state;
[0025] The optimizer is used to update the parameters of the Actor network, Critic network, and the adaptive temperature coefficients alpha and beta;
[0026] The entropy regularization module is used to calculate the individual entropy according to the action probability distribution output by the Actor network, and calculate the joint entropy based on each agent sending the action probability distribution to all neighboring agents determined by the adjacency list at each time step. The intensity of the individual entropy regularization and joint entropy regularization is adjusted respectively by using the adaptive temperature coefficients alpha and beta;
[0027] The agent identification and adjacency relationship module is used to set a unique identifier for each agent and set the identifiers of all its neighboring agents for each agent according to the adjacency list.
[0028] Furthermore, the expression of the individual entropy is as follows:
[0029] H(π(·|s))=-Σπ(a|s)*log(π(a|s))
[0030] Among them, \(H(\pi(\cdot|s))\) represents the individual entropy, \(H(\cdot)\) represents the operation of calculating entropy, \(\pi(a|s)\) represents the probability of taking action \(a\) in state \(s\), and \(\sum\) represents the summation over all actions \(a\).
[0031] Furthermore, the calculation process of the joint entropy is as follows:
[0032] A1. Obtain the action probability distribution of the current agent output by the Actor network, and the action probability distributions of all neighbor agents of the current agent according to the adjacency list;
[0033] A2. Perform an outer product operation on the action probability distributions obtained in A1;
[0034] A3. Reshape the tensor obtained in A2, and calculate the reshaped tensor to complete the calculation of the joint entropy. Among them, the reshaped tensor becomes a two-dimensional tensor with a shape of \((batch\_size, A)\), , where each row contains the probabilities of all joint actions in this batch, \(batch\_size\) represents the batch size, \(A\) represents the number of all joint action combinations, and \(action\_dim\_N\) represents the size of the action space of the \(N\)th agent.
[0035] Furthermore, the expression of the joint entropy is as follows:
[0036] \(H_{joint}=-\sum P*\log(P)\)
[0037] where \(H_{joint}\) represents the joint entropy and \(P\) represents the joint probability distribution.
[0038] Furthermore, the specific outer product operation is as follows:
[0039] Expand the action probability distributions of the agents in sequence;
[0040] Multiply the expanded tensors to obtain a tensor with a shape of , and complete the outer product operation. Among them, each element in the obtained tensor represents the probability of a group of joint actions, and \(action\_dim\_N\) represents the size of the action space of the \(N\)th agent.
[0041] Furthermore, the expression for updating the adaptive temperature coefficient \(\alpha\) is as follows:
[0042] \(L_{\alpha}=-\log(\alpha)*(H(\pi(\cdot|s)+target\_entropy\_individual)\)
[0043] Among them, \(L_{\alpha}\) represents the loss function value of the adaptive temperature coefficient \(\alpha\), \(\alpha\) represents the adaptive temperature coefficient used to adjust the individual entropy regularization strength, \(H(\pi(\cdot|s))\) represents the individual entropy, \(target\_entropy\_individual\) represents the preset target individual entropy, which is set to the negative size of the action space, and \(\log(\alpha)\) represents the logarithm of the adaptive temperature coefficient \(\alpha\);
[0044] The update expression of the adaptive temperature coefficient \(\beta\) is as follows:
[0045] \(L_{\beta}=-\log(\beta)*(H(\pi_{joint}) + target\_entropy_{joint})\)
[0046] Among them, \(L_{\beta}\) represents the loss function value of the adaptive temperature coefficient \(\beta\), \(\beta\) represents the adaptive temperature coefficient used to adjust the joint entropy regularization strength, \(\log(\beta)\) represents the logarithm of the adaptive temperature coefficient \(\beta\), \(H(\pi_{joint})\) represents the joint entropy, and \(target\_entropy_{joint}\) represents the preset target joint entropy, which is set to the negative size of the action space * the number of joint agents.
[0047] Furthermore, the expression of the loss function of the Critic network is as follows:
[0048] \(L_{critic}=(R + \gamma*V(s') - V(s))^2\)
[0049] Among them, \(L_{critic}\) represents the loss function of the Critic network, \(R\) represents the immediate reward, \(\gamma\) represents the discount factor, and \(V(s)\) and \(V(s')\) respectively represent the value estimates of the current state and the next state;
[0050] The expression of the loss function of the Actor network is as follows:
[0051] \(L_{actor}=-(\log(\pi(a|s))*A(s, a)+\alpha*H(\pi(\cdot|s))+\beta*H(\pi_{joint}))\)
[0052] Among them, \(L_{actor}\) represents the loss function of the Actor network, \(\pi(a|s)\) represents the probability of taking action \(a\) in state \(s\), \(A(s, a)\) represents the advantage function, \(H(\pi(\cdot|s))\) represents the individual entropy, \(H(\pi_{joint})\) represents the joint entropy, and \(\alpha\) and \(\beta\) respectively represent the adaptive temperature coefficients of the individual entropy and the joint entropy.
[0053] The beneficial effects of the present invention are:
[0054] (1) The present invention adopts a distributed multi-agent reinforcement learning framework, regarding each traffic signal control intersection in the urban road network as an independent agent. The agents share limited local information (only sharing the action probability distribution) through a predefined adjacency list. Each agent learns based on the A2C (Advantage Actor-Critic) algorithm framework and innovatively introduces an adaptive individual entropy regularization and an adaptive joint entropy regularization mechanism to improve the exploration efficiency and promote the collaborative cooperation among agents.
[0055] (2) Higher exploration efficiency: By introducing individual entropy regularization, the present invention encourages the agents to explore diverse action strategies, avoiding the premature convergence of the algorithm to suboptimal solutions. The adaptive temperature coefficient alpha can dynamically adjust the regularization strength according to the entropy value of the current policy, enabling the agents to achieve a better balance between exploration and exploitation, and thus more effectively explore the state-action space.
[0056] (3) Stronger collaboration ability: The most prominent advantage of the present invention is the introduction of joint entropy regularization. Joint entropy regularization encourages adjacent agents to explore diverse joint actions, that is, it encourages them to adopt different action combinations, which is quite different from the traditional method that encourages individual agents to explore independently. By maximizing the joint entropy, the agents not only need to explore their respective action spaces but also consider the action choices of each other, thus avoiding situations such as "all agents choose the same action" or "adjacent agents take conflicting actions". The adaptive temperature coefficient beta can dynamically adjust the regularization strength according to the entropy value of the current joint policy, making the collaboration among agents more effective and flexible. The joint entropy calculation method (based on the outer product) proposed by the present invention can accurately measure the entropy of the joint distribution of the actions of adjacent agents, providing key technical support for achieving effective collaboration.
[0057] (4) Better scalability: The present invention adopts a fully distributed learning framework. Each agent only needs to learn based on its own local observation information and the information from neighboring agents, without the need for global information or a central controller. This distributed architecture greatly reduces the computational complexity and communication overhead, enabling the system to easily scale to large-scale and high-density urban road networks.
[0058] (5) Unification of theoretical guarantee and actual performance: Many existing multi-agent reinforcement learning methods lack theoretical convergence guarantees. The present invention is based on the A2C (Advantage Actor-Critic) algorithm framework and introduces entropy regularization, which can theoretically guarantee the convergence of the algorithm (or at least local convergence). At the same time, the results of a large number of simulation experiments show that the present invention also achieves excellent performance in actual traffic signal control tasks, significantly outperforming traditional traffic signal control methods and other existing multi-agent reinforcement learning methods. Description of the Drawings
[0059] Figure 1 It is a flowchart of the method of the present invention. Specific Embodiments
[0060] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0061] Embodiment
[0062] As Figure 1 shown, the present invention provides a distributed multi-agent traffic signal cooperative control method, and its implementation method is as follows:
[0063] S1. Initialization stage: Perform initialization operations on the urban road network environment, construct an adjacency list according to the topological structure of the urban road network, and regard each traffic signal control intersection in the urban road network as an independent agent, and initialize each agent, where local information sharing is performed between agents through the adjacency list;
[0064] In this embodiment, the urban road network environment (traffic simulator) is initialized, an adjacency list (defining the connection relationship between agents) is constructed according to the topological structure of the urban road network, and each traffic signal control intersection is regarded as an independent agent. Each agent is initialized, including initializing its Actor network, Critic network, optimizer, target entropy and other parameters. Local information sharing is performed between agents through the adjacency list.
[0065] In this embodiment, the agents include:
[0066] The Actor network is used to generate an action probability distribution based on the local observation state observed by the agent. The local observation state includes the lane queue lengths, vehicle average speeds, and vehicle numbers of each approach at the current intersection, as well as the phase state and the duration of the phase of the current traffic signal.
[0067] The Critic network is used to evaluate the value of the current state based on the local observation state observed by the agent, that is, to predict the long-term cumulative return of the current state. In this embodiment, the value output by the Critic network is applied in the learning stage of S6: calculating the TD error; serving as the basis for updating the parameters of the Critic network (the goal is to reduce the TD error); calculating the advantage function, which indirectly affects the update of the Actor network. In this embodiment, the TD error = TD target value - current value, that is: TD error = TD_target - V(state).
[0068] The optimizer is used to update the parameters of the Actor network, the Critic network, and the adaptive temperature coefficients alpha and beta.
[0069] The entropy regularization module is used to calculate the individual entropy based on the action probability distribution output by the Actor network, and to calculate the joint entropy based on each agent sending the action probability distribution to all neighbor agents determined by the adjacency list at each time step. Among them, the adaptive temperature coefficients alpha and beta are used to adjust the intensities of the individual entropy regularization and the joint entropy regularization respectively. The calculation process of the joint entropy is as follows:
[0070] A1. Obtain the action probability distribution of the current agent (here, the current does not refer to the current in the time sense, because it is a multi-agent environment, so it refers to the individual) output by the Actor network, and obtain the action probability distributions of all neighbor agents of the current agent according to the adjacency list;
[0071] A2. Perform an outer product operation on the action probability distributions obtained in A1. The specific outer product operation is as follows:
[0072] Expand the action probability distributions of the agents in sequence;
[0073] Multiply the expanded tensors to obtain a tensor with a shape of to complete the outer product operation. Each element in the obtained tensor represents the probability of a group of joint actions, and action_dim_N represents the size of the action space of the Nth agent;
[0074] A3. Reshape the tensor obtained in A2 and calculate the reshaped tensor to complete the calculation of the joint entropy. After reshaping, the tensor becomes a two-dimensional tensor with a shape of (batch_size, A). Each row contains the probabilities of all joint actions in the batch. batch_size represents the batch size, which usually refers to the number of data samples processed in one training iteration in deep learning and reinforcement learning. A represents the number of all joint action combinations, which defines the dimension of the joint action space. The joint action space refers to the space composed of all possible action combinations of all agents. action_dim_N represents the size of the action space of the Nth agent.
[0075] Agent identification and adjacency relationship module, used to set a unique identifier for each agent and, according to the adjacency list, set the identifiers of all its neighbor agents for each agent.
[0076] In this embodiment, the present invention adopts a fully distributed architecture. An agent is deployed at each intersection, and the agents communicate through message passing. There is no central controller, and each agent is responsible for controlling the traffic lights at one intersection.
[0077] In this embodiment, initialize the urban road network environment (traffic simulator) and the adjacency list; initialize the Actor network, Critic network, target individual entropy, target joint entropy, adaptive temperature coefficients alpha and beta, and optimizer (Adam), etc. of each agent.
[0078] In this embodiment, the main components of the agent include:
[0079] (1) Actor network (PolicyNet): Its input is the local observation state of the agent. The local observation state can include: the lane queue lengths of each approach at the current intersection; the average vehicle speeds of each approach at the current intersection; the number of vehicles on each approach at the current intersection; the phase state of the current traffic light (e.g., red, green, yellow); the duration of the current phase. Other relevant state features can also be added according to the actual application scenario.
[0080] The output of the Actor network (PolicyNet) is: the probability distribution of the agent on each optional action. In the present invention, the action is the selection of discrete signal phases. For example, phases such as "east-west straight", "east-west left turn", "north-south straight", "north-south left turn" can be selected.
[0081] Network Structure of the Actor Network: The multi-layer perceptron (MLP) structure is adopted, which includes an input layer, multiple hidden layers, and an output layer. The activation function uses ReLU (Rectified Linear Unit). The output layer uses the Softmax function to ensure that the output is a legal probability distribution.
[0082] (2) Critic Network (ValueNet): The input of the Critic network is: the local observation state of the agent. The local observation state can include: the lane queue lengths of each approach at the current intersection; the average vehicle speeds of each approach at the current intersection; the number of vehicles of each approach at the current intersection; the phase state of the current traffic signal (e.g., red, green, yellow); the duration of the current phase; other relevant state features can also be added according to the actual application scenario.
[0083] The output of the Critic network is: the value estimation of this state, that is, the prediction of the long-term cumulative reward for the current state.
[0084] Network Structure of the Critic Network: The MLP structure is adopted, which includes an input layer, multiple hidden layers, and an output layer, and the activation function uses ReLU.
[0085] (3) The entropy regularization module includes:
[0086] 1. Calculate the individual entropy: According to the action probability distribution output by the Actor network, calculate the individual entropy of the current policy. The individual entropy measures the randomness of the action selection of a single agent. The formula is as follows:
[0087] H(π(·|s)) = -Σπ(a|s) * log(π(a|s))
[0088] Among them, H(π(·|s)) represents the individual entropy, H() represents the operation of calculating entropy, π(a|s) represents the probability of taking action a in state s, and Σ represents the summation over all actions a.
[0089] More specifically, H() is a function whose input is an action probability distribution (in this example, the policy π(·|s)), and the output is the entropy of this action probability distribution (a numerical value), mapping an action probability distribution to a numerical value (entropy) representing its degree of uncertainty.
[0090] The purpose of calculating the individual entropy is: individual entropy regularization encourages the agent to explore different actions and avoid prematurely converging to a sub-optimal policy.
[0091] 2. Calculate the joint entropy: The joint entropy measures the entropy of the joint distribution of actions between adjacent agents. Different from individual entropy that encourages a single agent to explore different actions, joint entropy encourages adjacent agents to explore different combined actions, thereby promoting their cooperation and avoiding situations like "everyone chooses the same action" or "conflicting actions". The detailed calculation method of joint entropy is as follows:
[0092] Probability tensor construction: For each agent, first obtain its own action probability distribution (output by the Actor network, with shape (batch_size, action_dim)). Then, according to the adjacency list, obtain the action probability distributions of all its neighbor agents. batch_size represents the batch size, and action_dim represents the dimension of the action space, that is, the number of different actions that each agent can take.
[0093] Outer Product calculation: Perform an outer product operation on the action probability distributions. The outer product operation is a key step in calculating the joint entropy of the present invention. Specifically, assume that the total number of the current agent and its neighbor agents is N, and perform the following operations: Expand the action probability distribution of the first agent (with shape (batch_size, action_dim_1)) to (batch_size, action_dim_1, 1, 1,...) (N - 1 ones); Expand the action probability distribution of the second agent (with shape (batch_size, action_dim_2)) to (batch_size, 1, action_dim_2, 1,...); and so on, expand the action probability distribution of the Nth agent to (batch_size, 1, 1,..., action_dim_N); Then, multiply these expanded tensors. Due to the broadcasting mechanism of PyTorch, this multiplication operation actually calculates the outer product of these probability distributions, resulting in a tensor with shape (batch_size, action_dim_1, action_dim_2,..., action_dim_N). Each element in this tensor represents the probability of a corresponding set of joint actions (i.e., a combination where each agent takes a specific action), and action_dim_N represents the size of the action space of the Nth agent.
[0094] Reshape: Reshape the tensor obtained in the previous step into (batch_size, action_dim_1 * action_dim_2 *,..., * action_dim_N). In this way, each row represents the probability of a joint action of all relevant agents (the current agent and its neighbors) in a batch of data.
[0095] Entropy calculation: Finally, calculate the entropy of the reshaped tensor using the following formula:
[0096] H_joint = -ΣP * log(P)
[0097] Where H_joint represents the joint entropy, P represents the joint probability distribution, and represents the probability values of different joint action combinations. This formula calculates the weighted sum of the probabilities of all joint actions and their logarithms, and then takes the negative value to obtain the joint entropy.
[0098] The calculation process of the joint entropy can be regarded as combining the action probability distributions of multiple agents into a joint probability distribution, and then calculating the entropy of this joint distribution. The outer product operation is the key to realizing this combination. It "unfolds" multiple independent probability distributions into a high-dimensional tensor, where each element represents the probability of a possible joint action. The reshape operation "flattens" this high-dimensional tensor into a two-dimensional tensor for convenient subsequent entropy calculation. Finally, according to the entropy formula, calculate the entropy of this joint probability distribution as an index to measure the diversity of action cooperation among neighboring agents.
[0099] Calculating the adaptive temperature coefficients: The individual entropy and the joint entropy respectively correspond to two learnable adaptive temperature coefficients alpha and beta, which are used to adjust the strength of entropy regularization. These two temperature coefficients are another key innovation of this invention. They are not fixed hyperparameters but are adaptively adjusted by the gradient descent method.
[0100] Update of the adaptive temperature coefficient alpha: The update objective of the adaptive temperature coefficient alpha is to make the individual entropy close to a preset target individual entropy value (target_entropy_individual), and the formula is as follows:
[0101] L_alpha = -log(alpha) * (H(π(·|s) + target_entropy_individual)
[0102] Among them, $L_{\alpha}$ represents the loss function value of the adaptive temperature coefficient $\alpha$, $\alpha$ represents the adaptive temperature coefficient used to adjust the individual entropy regularization strength, $H(\pi(\cdot|s))$ represents the individual entropy, $target\_entropy\_individual$ represents the preset target individual entropy, which is set to the negative of the size of the action space, and $\log(\alpha)$ represents the logarithm of the adaptive temperature coefficient $\alpha$.
[0103] Update process of the adaptive temperature coefficient $\alpha$: Clear the old gradient: Use the optimizer to zero the calculated gradient of the adaptive temperature coefficient $\alpha$ parameter; Calculate the new gradient: Automatically calculate the direction and magnitude (gradient) that the adaptive temperature coefficient $\alpha$ parameter needs to be adjusted according to the loss of the adaptive temperature coefficient $\alpha$; Update the adaptive temperature coefficient $\alpha$: Use the Adam optimizer to automatically adjust the adaptive temperature coefficient $\alpha$ parameter according to the calculated gradient (actually, it is the logarithm of the adaptive temperature coefficient $\alpha$ that is adjusted).
[0104] Update of the adaptive temperature coefficient $\beta$: The update goal of the adaptive temperature coefficient $\beta$ is to make the joint entropy close to a preset target joint entropy value ($target\_entropy\_joint$). The formula is as follows:
[0105] $L_{\beta}=-\log(\beta)*(H(\pi_{joint}) + target\_entropy\_joint)$
[0106] Among them, $L_{\beta}$ represents the loss function value of the adaptive temperature coefficient $\beta$, $\beta$ represents the adaptive temperature coefficient used to adjust the joint entropy regularization strength, $\log(\beta)$ represents the logarithm of the adaptive temperature coefficient $\beta$, $H(\pi_{joint})$ represents the joint entropy, and $target\_entropy\_joint$ represents the preset target joint entropy, which is set to the negative of the size of the action space * the number of joint agents.
[0107] Update of the adaptive temperature coefficient $\beta$: Clear the old gradient: Use the optimizer to zero the calculated gradient of the adaptive temperature coefficient $\beta$ parameter; Calculate the new gradient: Automatically calculate the direction and magnitude (gradient) that the adaptive temperature coefficient $\beta$ parameter needs to be adjusted according to the loss of the adaptive temperature coefficient $\beta$; Update the adaptive temperature coefficient $\beta$: Use the Adam optimizer to automatically adjust the adaptive temperature coefficient $\beta$ parameter according to the calculated gradient (actually, it is the logarithm of the adaptive temperature coefficient $\beta$ that is adjusted).
[0108] (4)Optimizer: The Adam optimizer is used to update the parameters of the Actor network, the Critic network, and the two adaptive temperature coefficients (alpha and beta) respectively. Adam is a commonly used stochastic gradient descent optimization algorithm with the advantage of an adaptive learning rate.
[0109] In this embodiment, the adjacency list is used to define the communication topology structure between agents. It is a dictionary where the key is the name (or ID) of the agent, and the value is a list that contains the names (or IDs) of all the agents adjacent to this agent. The adjacency list stipulates which agents can exchange information. In the present invention, the adjacency list is predefined according to the actual topology structure of the road network, that is, the agents corresponding to adjacent intersections are neighbors of each other.
[0110] S2. Environment interaction stage: For each agent, independently obtain the local observation state of the intersection where it is located. For example, the lane queue length, the average vehicle speed, etc.
[0111] In this embodiment, for each agent, obtain the current local observation state. According to the action probability distribution output by the Actor network, sample an action (signal phase), apply the action to the environment (traffic simulator), and obtain information such as the next state, reward, and whether it is terminated.
[0112] In this embodiment, at the beginning of each simulation time step, each agent independently observes the local state of the intersection where it is located. The local state information is the basis for the agent to make decisions and usually includes: the queue length of each approach lane; the average vehicle speed of each approach lane; the number of vehicles in each approach lane; the state of the current signal phase; the duration of the current phase; and other information related to the traffic condition.
[0113] S3. Decision-making stage: Based on the local observation state obtained by each agent, obtain the action probability distribution, and the agent makes an action selection by sampling according to the action probability distribution, where the action probability distribution represents the probability that the agent selects each available traffic signal phase.
[0114] In this embodiment, each agent obtains an action probability distribution through the Actor network based on its local observation state. This action probability distribution represents the probability that the agent selects each available traffic signal phase (for example, straight ahead in the east-west direction, left turn in the north-south direction, etc.). Then, the agent samples according to this action probability distribution and selects a specific action (i.e., signal phase). That is, each agent inputs the collected local observation state into its Actor network. After processing, the Actor network outputs an action probability distribution, which represents the probability that the agent selects each available signal phase. The agent selects a specific action (signal phase) through sampling according to the action probability distribution, and the agent records the action probability distribution of this time.
[0115] S4. Execution stage: Each agent sends the selected action to the traffic simulator, updates the status of the traffic lights and simulates the movement of vehicles, and the traffic simulator calculates and returns the agent information at the next time step, where the agent information includes the reward of each agent and the next local observation state of each agent;
[0116] In this embodiment, each agent sends the selected action (signal phase) to the traffic simulator. The simulator updates the status of the traffic lights according to the action and simulates the movement of vehicles. The traffic simulator calculates and returns the agent information at the next time step, including the reward (reward, for example, negative waiting time) of each agent and the next local observation state (next_state).
[0117] In this embodiment, in the S4 execution stage, the results obtained after the interaction between the agent and the traffic simulator (the reward reward and the next local observation state next_state of each agent) are applied to the learning stage in S6. The specific application method is as follows:
[0118] Using the reward reward and the next local observation state next_state of the agent, and a discount factor (gamma), calculate the "target value" (TD target), which represents an estimate of the future cumulative reward. Subtract the value of the current state from the TD target calculated using the reward reward to obtain the advantage function and the TD error. The advantage function is used to update the Actor network, and the TD error is used to update the Critic network.
[0119] S5. Information sharing stage: Each agent sends the action probability distribution to all its neighbor agents, where only the action probability distribution is shared between each agent, and the neighbor agents are determined by the adjacency list;
[0120] In this embodiment, each agent sends its action probability distribution (i.e., the output of the Actor network in the S3 stage) to all its neighbor agents (determined according to the adjacency list constructed in the S1 stage). Only the action probability distribution is shared among agents, and other information such as the observation state is not shared.
[0121] In this embodiment, the result of information sharing in S5 (the action probability distributions of neighbor agents) is used to calculate the joint entropy between neighbor agents in the S6 learning stage. As a regularization means, the joint entropy is incorporated into the loss function of the Actor network to encourage cooperation between adjacent agents, prevent them from all taking the same action or conflicting actions, and ultimately improve the performance of the entire system.
[0122] In this embodiment, information sharing between agents is carried out through message passing. Specifically, each agent sends the action probability distribution output by its Actor network to all its neighbor agents (determined according to the adjacency list) at each time step. After receiving this information, the neighbor agents use it to calculate the joint entropy. The present invention does not share the observation state of agents, only shares the action probability distribution, which helps to reduce communication overhead and protect the local privacy of agents.
[0123] S6. Learning stage: Update the parameters according to the action probability distribution obtained in S3, the reward from the traffic simulator and the next local observation state obtained in S4, and the action probability distributions received from neighbor agents obtained in S5, and update the Actor network, Critic network, and the adaptive temperature coefficients alpha and beta by calculating the individual entropy and joint entropy, where the individual entropy is used to measure the randomness of the action selection of a single agent, and the joint entropy is used to measure the joint distribution of actions between adjacent agents;
[0124] In this embodiment, each agent receives: the reward from the traffic simulator and the next local observation state, as well as the action probability distributions from all its neighbor agents.
[0125] Each agent uses the received information to update its internal Actor network, Critic network, and the adaptive temperature coefficients alpha and beta. The update process is as follows:
[0126] First, data preparation: Prepare the received data (which means converting the information into tensors); estimate the value V(state) of the current state, estimate the value V(next_state) of the next state; calculate the TD target value (TD_target); calculate the TD error (Temporal Difference Error) and the advantage function Advantage (which represents the advantage of taking a specific action relative to the average level). The advantage function and the TD error are numerically the same here, that is, both are equal to TD_target - V(state), but their usage and meanings are different, so two terms are used for distinction.
[0127] The expression of the advantage function Advantage is as follows:
[0128] Advantage = TD_target - V(state)
[0129] Among them, V(state) represents the estimate of the value of the current state state by the Critic network, and TD_target represents an estimated value of the future cumulative reward, that is, combining the immediate reward reward and the estimated value of the future expected reward V(next_state) to improve the evaluation of the value of the current state.
[0130] The expression of the target value TD_target is as follows:
[0131] TD_target = reward + gamma * V(next_state) * (1 - done)
[0132] Among them, reward represents the immediate reward obtained at the current time step, gamma represents the discount factor, which represents the importance of future rewards (0 < gamma <= 1), V(next_state) represents the estimate of the value of the next state by the Critic network, and done represents a boolean value indicating whether the current episode has ended (1 means ended, 0 means not ended);
[0133] Then, calculate the individual entropy. The calculation process is as follows: Calculate the individual entropy of the current policy according to the action probability distribution output by the current agent Actor network. The individual entropy measures the randomness (or uncertainty) of the action selection of a single agent.
[0134] Next, calculate the joint entropy. The calculation process is as follows: If the current agent has neighbor agents, calculate the joint entropy; obtain the action probability distributions of all neighbor agents; add the action probability distribution of the current agent itself; calculate the joint entropy of these probability distributions. The joint entropy measures the diversity (or uncertainty) of the joint actions of adjacent agents; update the target joint entropy.
[0135] Next, calculate the loss function. The calculation process is as follows:
[0136] Loss of the Critic network: Measures the accuracy of the Critic network's estimation of the state value;
[0137] Loss of the Actor network: Consists of three main parts: (1) Encourage the agent to take actions with higher advantages (i.e., the advantage function term); (2) Individual entropy regularization term: Encourage the agent to explore diverse actions. By maximizing the individual entropy, the agent will not only take a certain action prematurely but is more likely to try various actions; (3) Joint entropy regularization term: Encourage adjacent agents to explore diverse joint actions, promote their cooperation, and avoid all agents taking the same action or taking conflicting actions. Among them, the adaptive temperature coefficient alpha loss is used to adjust the adaptive temperature coefficient alpha to make the individual entropy close to the preset target individual entropy. Adaptive temperature coefficient beta loss: Used to adjust the adaptive temperature coefficient beta to make the joint entropy close to the preset target joint entropy.
[0138] Finally, update the parameters: Update the parameters of the Actor network, Critic network, adaptive temperature coefficient alpha, and beta respectively. The update process includes the following three steps: Clear the gradients: Clear the previous gradient information; Backpropagation: Calculate the gradients of each parameter according to the loss function; Update the parameters: Adjust the values of the parameters according to the gradients.
[0139] In this embodiment, calculate the loss (Loss Calculation):
[0140] For each agent: Calculate the loss of the Critic network: Use the square of the TD error (Temporal Difference Error) as the loss function. The formula is as follows:
[0141] L_critic=(R+γ*V(s')-V(s))^2
[0142] Where, L_critic represents the loss function of the Critic network, R represents the immediate reward, γ represents the discount factor, and V(s) and V(s') represent the value estimations of the current state and the next state respectively.
[0143] Calculate the loss of the Actor network: It consists of three parts: the Advantage Function, the individual entropy regularization term, and the joint entropy regularization term. The formula is as follows:
[0144] L_actor = -(log(π(a|s)) * A(s, a) + α * H(π(·|s)) + β * H(π_joint))
[0145] Where, L_actor represents the loss function of the Actor network, π(a|s) represents the probability of taking action a in state s, A(s, a) represents the Advantage Function, H(π(·|s)) represents the individual entropy, H(π_joint) represents the joint entropy, and α and β respectively represent the adaptive temperature coefficients of the individual entropy and the joint entropy.
[0146] Calculate the loss of the adaptive temperature coefficient alpha:
[0147] L_alpha = -log(alpha) * (H(π(·|s) + target_entropy_individual)
[0148] Where, L_alpha represents the value of the loss function of the adaptive temperature coefficient alpha, alpha represents the adaptive temperature coefficient used to adjust the individual entropy regularization strength, H(π(·|s)) represents the individual entropy, target_entropy_individual represents the preset target individual entropy, which is set to the negative of the size of the action space, and log(alpha) represents the logarithm of the adaptive temperature coefficient alpha.
[0149] Calculate the loss of the adaptive temperature coefficient beta:
[0150] L_beta = -log(beta) * (H(π_joint) + target_entropy_joint)
[0151] Where, L_beta represents the value of the loss function of the adaptive temperature coefficient beta, beta represents the adaptive temperature coefficient used to adjust the joint entropy regularization strength, log(beta) represents the logarithm of the adaptive temperature coefficient beta, H(π_joint) represents the joint entropy, target_entropy_joint represents the preset target joint entropy, which is set to the negative of the size of the action space * the number of joint agents.
[0152] In this embodiment, the purpose of individual entropy regularization is to encourage each agent to independently explore more possible actions and avoid premature convergence to sub-optimal policies. Its mechanism is as follows: By adding individual entropy to the Actor network loss function, while maximizing the cumulative reward, the agent also tends to maximize the randomness of its action selection.
[0153] The adaptive temperature coefficient alpha is used to control the intensity of individual entropy regularization. The larger the adaptive temperature coefficient alpha, the greater the encouragement for individual entropy, and the more the agent tends to explore; the smaller the adaptive temperature coefficient alpha, the less the encouragement for individual entropy, and the more the agent tends to utilize the learned knowledge. The adaptive temperature coefficient alpha will be automatically adjusted.
[0154] In this embodiment, the purpose of joint entropy regularization is to encourage adjacent agents to explore different joint action combinations, promote effective cooperation between them, and avoid situations such as "everyone chooses the same action" or "actions conflict with each other". Its mechanism is to add joint entropy to the Actor network loss function, so that while maximizing the cumulative reward, the agent also tends to maximize the diversity of its joint actions with neighbor agents.
[0155] The role of the adaptive temperature coefficient beta is to control the intensity of joint entropy regularization. The larger the adaptive temperature coefficient beta, the greater the encouragement for joint entropy, and the more inclined the agents are to adopt different action combinations; the smaller the adaptive temperature coefficient, the less the encouragement for joint entropy. The adaptive temperature coefficient will be automatically adjusted.
[0156] In this embodiment, the adaptive temperature coefficients alpha and beta are not fixed but will be automatically adjusted. Its advantages lie in dynamically balancing exploration and exploitation: The adaptive adjustment of the adaptive temperature coefficient alpha enables the agent to conduct more exploration in the initial stage of training and gradually reduce exploration in the later stage of training. Dynamically balancing individual and cooperation: The adaptive adjustment of the adaptive temperature coefficient beta enables the agent to adjust the degree of cooperation according to the current learning state.
[0157] S7. Loop iteration: Repeat S2 to S6 until the preset number of training rounds is reached, and the cooperative control of traffic signals is completed.
[0158] In this embodiment, the process of repeating S2 to S6 continuously conducts environment interaction, decision-making, execution, information sharing, and learning until the preset number of training rounds (or other stop conditions are met). Through this loop iteration process, the agent gradually learns how to make optimal signal control decisions under different traffic conditions and cooperate with other agents, thereby achieving the cooperative control of the entire traffic network.
[0159] In summary, the beneficial effects of the present invention are as follows:
[0160] Improve exploration efficiency: By introducing an adaptive dual regularization mechanism of individual entropy and joint entropy, encourage the agent to explore diverse behavioral strategies, and improve the exploration efficiency and global optimization ability of the algorithm;
[0161] Enhance agent collaboration: Through joint entropy regularization and an information sharing mechanism based on the adjacency list, promote collaboration among agents, enabling them to jointly optimize the overall traffic performance of the road network, rather than just the local optimum of a single intersection;
[0162] Enhance scalability: Adopt a distributed learning framework and a local communication mechanism to reduce the computational complexity and communication overhead of the algorithm, enabling this system to be applied to large-scale and high-density urban road networks.
Claims
1. A distributed multi-agent traffic signal collaborative control method, characterized in that: The following steps are involved: S1, initialization stage: Initialize the urban road network environment, and build an adjacency table according to the topological structure of the urban road network. Each traffic light-controlled intersection in the urban road network is regarded as an independent intelligent agent, and each intelligent agent is initialized. Among them, the intelligent agents share local information through the adjacency table; S2, environment interaction stage: each agent independently obtains the local observation state of the intersection where it is located; S3, decision-making stage: based on the local observation state obtained by each agent, the action probability distribution is obtained, and the agent selects the action through sampling according to the action probability distribution, where the action probability distribution represents the probability of the agent selecting each available traffic signal phase; S4, execution phase: Each agent sends the selected action to the traffic simulator, updates the state of the traffic lights and simulates the movement of vehicles, and uses the traffic simulator to calculate and return the agent information of the next time step, where the agent information includes the reward of each agent and the next local observation state of each agent; S5, information sharing stage: each agent sends the action probability distribution to all its neighboring agents, where each agent only shares the action probability distribution, and the neighboring agents are determined by the adjacency list; S6, learning phase: according to the action probability distribution obtained in S3, the reward and next local observation state of each agent obtained in S4, and the action probability distribution received from the neighboring agents obtained in S5, the parameters are updated, and the Actor network, Critic network and adaptive temperature coefficients alpha and beta are updated by calculating individual entropy and joint entropy. The individual entropy is used to measure the randomness of the action selection of a single agent, and the joint entropy is used to measure the joint distribution of actions between adjacent agents. The updated expression of the adaptive temperature coefficient alpha is as follows: L_alpha=-log(alpha)*(H(π(·|s))+target_entropy_individual) Wherein, L_alpha represents the loss function value of the adaptive temperature coefficient alpha, alpha represents the adaptive temperature coefficient used to adjust the individual entropy regularization strength, H(π(·|s)) represents the individual entropy, target_entropy_individual represents the preset target individual entropy, and the preset target individual entropy is set to the size of the negative action space, and log(alpha) represents the logarithm of the adaptive temperature coefficient alpha; The updated expression of the adaptive temperature coefficient beta is as follows: L_beta=-log(beta)*(H(π_joint)+target_entropy_joint) Wherein, L_beta represents the loss function value of the adaptive temperature coefficient beta, beta represents the adaptive temperature coefficient used to adjust the strength of the joint entropy regularization, log(beta) represents the logarithm of the adaptive temperature coefficient beta, H(π_joint) represents the joint entropy, target_entropy_joint represents the preset target joint entropy, and the preset target joint entropy is set to the product of the size of the negative action space and the number of joint agents; The expression of the loss function of the Critic network is as follows: L_critic=(R+γ*V(s')-V(s))^2 Where L_critic represents the loss function of the Critic network, R represents the immediate reward, γ represents the discount factor, V(s) and V(s') represent the value estimates of the current state and the next state respectively; The expression of the loss function of the Actor network is as follows: L_actor=-(log(π(a|s))*A(s, a)+ alpha*H(π(·|s))+ beta*H(π_joint)) Where L_actor represents the loss function of the Actor network, π(a|s) represents the probability of taking action a in state s, and A(s, a) represents the advantage function; S7, loop iteration: Repeat S2 to S6 until the preset training rounds are reached. The agent learns to obtain the optimal signal control decision under different traffic conditions, and works with all its neighboring agents to complete the coordinated control of traffic signals.
2. The distributed multi-agent traffic signal cooperative control method according to claim 1 is characterized in that: The intelligent agent in S1 includes: Actor network, which is used to generate action probability distribution based on the local observation state observed by the agent, where the local observation state includes the lane queue length, average vehicle speed and number of vehicles at each entrance of the current intersection, as well as the phase state of the current traffic light and the duration of the phase; The Critic network is used to evaluate the value of the current state based on the local observation state observed by the agent, that is, to predict the long-term cumulative return of the current state; Optimizer, used to update the parameters of the Actor network, Critic network, and adaptive temperature coefficients alpha and beta; The entropy regularization module is used to calculate the individual entropy according to the action probability distribution output by the Actor network, and calculate the joint entropy based on each agent sending the action probability distribution to all neighboring agents determined by the adjacency table at each time step, wherein the adaptive temperature coefficients alpha and beta are used to adjust the strength of the individual entropy regularization and the joint entropy regularization respectively. The agent identification and adjacency relationship module is used to set a unique identifier for each agent and set the identifiers of all neighboring agents for each agent according to the adjacency table.
3. The distributed multi-agent traffic signal collaborative control method according to claim 2 is characterized in that: The expression of the individual entropy is as follows: H(π(·|s))=-Σπ(a|s)*log(π(a|s)) Among them, H(π(·|s)) represents individual entropy, H() represents the operation of finding entropy, π(a|s) represents the probability of taking action a in state s, and Σ represents the sum of all actions a.
4. The distributed multi-agent traffic signal cooperative control method according to claim 2 is characterized in that: The calculation process of the joint entropy is as follows: A1. Obtain the action probability distribution of the current agent output by the Actor network, and obtain the action probability distribution of all neighboring agents of the current agent according to the adjacency list; A2, perform outer product operation on the action probability distribution obtained in A1; A3, reshape the tensor obtained in A2, and calculate the reshaped tensor to complete the calculation of joint entropy. The reshaped tensor becomes a two-dimensional tensor with a shape of (batch_size, A). , where each row contains the probability of all joint actions in a batch, batch_size represents the batch size, A represents the number of all joint action combinations, and action_dim_N represents the size of the action space of the Nth agent.
5. The distributed multi-agent traffic signal cooperative control method according to claim 4 is characterized in that: The expression of the joint entropy is as follows: H(π_joint)=-ΣP*log(P) Among them, H(π_joint) represents the joint entropy and P represents the joint probability distribution.
6. The distributed multi-agent traffic signal cooperative control method according to claim 4 is characterized in that: The outer product operation is specifically as follows: Expand the probability distribution of the agent's actions in sequence; Multiply the expanded tensors to get a shape of The outer product operation is completed on the tensor of , where each element in the obtained tensor represents the probability of a set of joint actions, and action_dim_N represents the size of the action space of the Nth agent.
Citation Information
Patent Citations
Pathological image segmentation method based on multi-agent deep reinforcement learning
CN117115182A
Method and system for determining action of device for given state using model trained based on risk-measure parameter
US20220198225A1