Multi-agent traffic signal control method based on space-time convolution attention mechanism
By introducing a spatiotemporal convolutional attention mechanism and deep reinforcement learning algorithm in multi-agent signal control, the problems of single environmental interaction and low state space utilization in multi-intersection signal control are solved, and more efficient coordinated traffic signal control is achieved.
Patent Information
- Application Number
- CN202510259602.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing multi-agent signal control methods have problems such as single environmental interaction and low state space utilization in independent and centralized control, which makes it difficult to efficiently coordinate multi-intersection signal control.
The multi-agent traffic signal control method based on the spatiotemporal convolutional attention mechanism is adopted to preprocess the state through the convolutional attention mechanism network, highlight important features and suppress redundant information, and design a deep reinforcement learning algorithm to consider the influence of neighboring agent state and reward when interacting with the environment.
The performance and effect of regional traffic signal coordinated control is improved, the signal control strategy is more flexible and proactive, the amount of information transmitted is reduced, and effective communication is ensured.
Smart Images

Figure CN120048134A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent traffic control, and specifically to a multi-agent traffic signal control method based on a spatio-temporal convolutional attention mechanism. Background Art
[0002] At present, multi-agent signal control can be divided into three categories: independent, centralized, and collaborative. Independent multi-agent algorithms directly train each agent without any cooperation between agents. The disadvantage of this method is that the influence of the behavior of other agents is completely regarded as part of the environment. Agents only focus on the environmental information of their respective intersections and do not consider the dynamic randomness of the strategies of other agents when learning local agents. Centralized control can be understood as a single agent jointly controlling the intersections in the area for all traffic lights in the area. Collaborative control can be achieved by sharing states or introducing heuristic operation decisions. In the observation of the environment by local agents, the states of some neighboring agents are introduced, and the value of each action is evaluated according to the cooperation state.
[0003] In order to improve the sensitivity of the model to traffic states and thus enable agents to make better action decisions, we often need to improve its network structure, which can help the model pay more attention to the distribution and dynamics of vehicles near intersections. For multi-agent signal control models, there are mainly the following problems: 1. Independent signal control only considers the local traffic environment during training, and the environment in the intersection lanes is complex, and it is impossible to correctly identify the impact of vehicle states in different spatio-temporal on traffic flow; 2. The state-action space of centralized signal control will increase sharply after the road network scale increases, consuming a large amount of resources; 3. As the number of intersections increases exponentially, the amount of communication information to be transmitted increases significantly. Therefore, while reducing the amount of transmitted information, effective communication needs to be ensured.
[0004] Therefore, in traffic signal control, there is a problem that it is difficult to efficiently coordinate the signals of adjacent intersections. In view of the above situation, there is an urgent need to develop a multi-agent traffic signal control method based on a spatio-temporal convolutional attention mechanism to overcome the deficiencies in current practical applications. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-agent traffic signal control method based on a spatio-temporal convolutional attention mechanism to solve the problems raised in the above background art.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A multi-agent traffic signal control method based on a spatio-temporal convolutional attention mechanism specifically includes the following steps:
[0008] Step 1: Model the multi-intersection traffic network in the area as a multi-agent system, and define the state, action, and reward;
[0009] Step 2: Use the convolutional attention mechanism network to preprocess the state. By introducing two sub-modules, channel attention and spatial attention, the channel dimension and spatial dimension of the state feature map are weighted respectively to highlight important features and suppress redundant information;
[0010] Step 3: Design an intersection signal control neural network model based on the deep reinforcement learning algorithm with spatio-temporal attention mechanism. The model includes two convolutional layers and a CBAM module, and the CBAM module is used to enhance the important features in the feature map;
[0011] Step 4: Train the multi-intersection signal coordination control model through the multi-agent collaborative control algorithm. When the agent interacts with the environment, the influence of the states and rewards of neighboring agents is considered to generate the best action response to the current global state.
[0012] As a further solution of the present invention: In step 1, the definition of the traffic state at the regional intersection includes:
[0013] Regard the traffic environment as a traffic grid composed of multiple intersections. Each intersection has four directions, and the edge of each direction consists of four lanes, namely the straight lane, the left-turn lane, and the right-turn lane;
[0014] The state space of each intersection is defined as:
[0015]
[0016] Among them, represents the position matrix of the vehicles in the incoming lane l of intersection i in the road network, represents the speed matrix of the vehicles in the incoming lane l of intersection i. The current joint state space of intersection i is composed of the current agent's local state space and the state spaces of neighboring agents, and can be expressed as:
[0017]
[0018] Among them, represents the local state information of the intersection at time t; represents the set of adjacent intersection state information; N i is the set of neighbor nodes.
[0019] As a further solution of the present invention: In step 1, the action space contains four phase spaces, namely straight and right-turn for north-south direction, left-turn for north-south direction, straight and right-turn for east-west direction, and left-turn for east-west direction.
[0020] As a further solution of the present invention: in step 1, the reward function is defined as a multi-objective optimization reward, expressed as:
[0021]
[0022] where q i represents the sum of the vehicle queue lengths of all lanes at the intersection at the previous moment; d i represents the sum of the waiting times of all lanes at the intersection at the previous moment; L jt is the red light duration of lane j at time t, and C is the maximum red light time that a driver can tolerate.
[0023] As a further solution of the present invention: in step 1, the agent reward and the rewards of adjacent domains are combined using the weighted sum method, and a spatial discount factor is introduced to make the reward value of the agent in the neighborhood positively correlated with the distance between the two agents. The closer the distance, the greater the reward value of the adjacent agent obtained. The spatial discount factor is expressed as:
[0024]
[0025] where represents the distance between node i and node j, and j ∈ N i ; is the discount coefficient;
[0026] The agent reward function is defined as the sum of the rewards obtained from the current local environment and the rewards of all agents in the neighborhood, expressed as:
[0027]
[0028] where is the local reward obtained by the agent; is the reward information of the adjacent agent j obtained by the agent, is the spatial discount factor.
[0029] As a further solution of the present invention: the calculation formula of the channel attention mechanism is:
[0030]
[0031] where and respectively represent the average pooling and max pooling operation numbers in the channel dimension. W 0 ∈ R C / r×C and W 1 ∈ R C×C / r represent the weight parameters of the shared multi-layer perceptron. σ represents the Sigmoid activation function;
[0032] The calculation formula of the spatial attention mechanism is as follows:
[0033]
[0034] Among them, and respectively represent the average pooling and max pooling operations in the spatial dimension. f 7×7 represents a kernel size of 7×7. σ represents the Sigmoid activation function.
[0035] As a further solution of the present invention: In step 3, the network structure of the deep reinforcement learning algorithm includes:
[0036] The first convolutional layer is used to extract the features of the normalized vehicle speed information and vehicle position information;
[0037] The CBAM module is used to perform channel attention and spatial attention processing on the feature map;
[0038] The second convolutional layer is used to further process the feature map;
[0039] The fully connected layer is used to generate the state value and the action advantage value.
[0040] As a further solution of the present invention: In step 4, the process of the multi-agent traffic signal cooperative control method includes:
[0041] Initialize the state space, action space, main network, and target network;
[0042] The agent makes decisions using the ε-greedy strategy to balance exploration and exploitation;
[0043] After the agent executes the selected action, the environment feedbacks the corresponding information and stores the experience sequence in the Replay buffer;
[0044] According to the prioritized experience replay strategy, extract experience data from the experience replay pool, calculate the loss function, and update the main network parameters.
[0045] As a further solution of the present invention: In step 4, the calculation formula of the loss function is:
[0046]
[0047] Among them, r t is the reward value, γ is the discount factor, Q(s t ,a; θ) is the action value function of the main network, θ - is the parameter of the target network, represents the action corresponding to the highest action value output by the online network in the state s t+1 below, Represents the action value output by the target network according to the action selected by the online network in state s t+1 when the state is s.
[0048] A multi-agent traffic signal control system based on a spatio-temporal convolutional attention mechanism, including a traffic signal control module for executing the multi-agent traffic signal dynamic collaborative control method based on the spatio-temporal convolutional attention mechanism described above, further including:
[0049] A traffic simulation environment module for simulating a multi-intersection traffic network and providing traffic flow data and status information;
[0050] An agent module for interacting with the traffic simulation environment module, executing signal control strategies and collecting feedback information;
[0051] An experience replay module for storing the experience data of the interaction between the agent and the traffic simulation environment and for training the traffic signal control module.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] 1. The present invention proposes a multi-agent traffic signal dynamic collaborative control method based on a spatio-temporal convolutional attention mechanism for the regional signal control problem, combining the attention mechanism with the regional signal control, which improves the performance and effect of the regional traffic signal collaborative control.
[0054] 2. The previous traffic signal control based on the deep reinforcement learning algorithm only mechanically defines the vehicle states in the intersection lanes and cannot better focus on the states of vehicles with strong correlations. However, the present invention incorporates a convolutional attention mechanism after each convolutional layer, which can better focus on the states of vehicles near the intersection, making the signal control strategy more flexible and proactive.
[0055] 3. In the multi-agent signal control, aiming at the problems of relatively single environmental interaction and high utilization rate of the state space existing in the independent control and centralized control in the regional signal control, the multi-agent dynamic collaborative control method proposed by the present invention can consider the influence of the states and rewards of neighboring agents when the agent interacts with the environment, generate the best action response to the current state, and improve the coordination control effect of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is a schematic diagram of multi-agent modeling in the regional road network in the embodiment of the present invention.
[0057] Figure 2 It is a schematic diagram of the action space of the traffic lights in the embodiment of the present invention.
[0058] Figure 3Schematic diagram of the D3QN_CBAM algorithm network structure in the embodiments of the present invention.
[0059] Figure 4 Schematic diagram of the multi-agent traffic signal collaborative control framework in the embodiments of the present invention. Detailed implementation manners
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0061] The following describes the specific implementation of the present invention in detail in conjunction with specific embodiments.
[0062] Please refer to Figures 1 - 4 , the multi-agent traffic signal control method based on the spatio-temporal convolutional attention mechanism provided by the embodiments of the present invention specifically includes the following steps:
[0063] Step 1: Model the multi-intersection traffic network in the area as a multi-agent system, and define the state, action, and reward;
[0064] Step 2: Use the convolutional attention mechanism network to preprocess the state. By introducing two sub-modules of channel attention and spatial attention, the channel dimension and spatial dimension of the state feature map are weighted respectively to highlight important features and suppress redundant information;
[0065] Step 3: Design an intersection signal control neural network model based on the deep reinforcement learning algorithm (D3QN_CBAM) of the spatio-temporal attention mechanism. The model includes two convolutional layers and a CBAM module, and the CBAM module is used to enhance important features in the feature map;
[0066] Step 4: Train the multi-intersection signal coordination control model through the multi-agent collaborative control algorithm. When the agent interacts with the environment, the influence of the states and rewards of neighboring agents is considered to generate the best action response to the current global state.
[0067] The multi-agent signal light coordination control method based on spatio-temporal convolutional attention mechanism and deep reinforcement learning designed by the present invention first models the multi-intersection traffic network of a region as a multi-agent system. Each multi-agent simultaneously considers the influence of the actions of adjacent multi-agents at adjacent moments during the learning strategy process, enabling multiple multi-agents to collaboratively control the signal lights at multiple intersections and establishing a tensor that can reflect the current state of the regional traffic network. Secondly, the convolutional attention mechanism is incorporated into the deep reinforcement learning neural network so that the agent can better focus on the states of vehicles at adjacent intersections. Finally, a multi-agent intersection signal control model based on the deep reinforcement learning algorithm D3QN_CBAM is established, and the multi-agent collaborative control algorithm is used to train the multi-intersection signal coordination control model. The actual traffic state information of the current intersection and adjacent intersections is input into the neural network model to obtain a signal control scheme that is conducive to improving the current traffic passing indicators.
[0068] In an embodiment of the present invention, please refer to Figures 1 - 4 , in step 1, the definition of the traffic state of the regional intersection includes:
[0069] The traffic environment is regarded as a traffic grid composed of multiple intersections. Each intersection has four directions, and the edge of each direction consists of four lanes, namely the straight lane, the left-turn lane, and the right-turn lane.
[0070] The state space of each intersection is defined as:
[0071]
[0072] Among them, represents the position matrix of vehicles in the incoming lane l of intersection i in the road network, represents the speed matrix of vehicles in the incoming lane l of intersection i. The combined state space of the current intersection i consists of the current agent's local state space and the state spaces of neighboring agents, and can be expressed as:
[0073]
[0074] Among them, represents the local state information of the intersection at time t; represents the set of adjacent intersection state information; N i is the set of neighbor nodes.
[0075] In step 1, the action space contains four phase spaces, namely north-south straight and right-turn (NSG), north-south left-turn (NSLG), east-west straight and right-turn (EWG), and east-west left-turn (EWLG).
[0076] In step 1, the reward function is defined as a multi-objective optimization reward, expressed as:
[0077]
[0078] where q i represents the sum of the vehicle queue lengths of all lanes at the intersection at the previous moment; d i represents the sum of the waiting times of all lanes at the intersection at the previous moment; L jt is the red light duration of lane j at time t, and C is the maximum tolerable red light time for the driver.
[0079] In step 1, the agent reward and the rewards of adjacent domains are combined using the weighted sum method, and a spatial discount factor is introduced to make the reward value of the agent in the neighborhood positively correlated with the distance between the two agents. The closer the distance, the greater the reward value of the adjacent agent obtained. The spatial discount factor is expressed as:
[0080]
[0081] where represents the distance between node i and node j, j ∈ N i ; is the discount coefficient;
[0082] The agent reward function is defined as the sum of the reward obtained from the current local environment and the rewards of all agents in the neighborhood, expressed as:
[0083]
[0084] where is the local reward obtained by the agent; is the reward information of the adjacent agent j obtained by the agent, is the spatial discount factor.
[0085] In an embodiment of the present invention, please refer to Figures 1 - 4 , in step 2, the calculation formula of the channel attention mechanism is:
[0086]
[0087] where and respectively represent the average pooling and maximum pooling operation numbers in the channel dimension. W 0 ∈ R C / r×C and W 1 ∈ R C×C / r represent the weight parameters of the shared multi-layer perceptron. σ represents the Sigmoid activation function;
[0088] The calculation formula of the spatial attention mechanism is as follows:
[0089]
[0090] Among them, and respectively represent the average pooling and max pooling operations in the spatial dimension. f 7×7 represents a kernel size of 7×7. σ represents the Sigmoid activation function.
[0091] Please refer to Figures 1 - 4 , in step 3, the network structure of the D3QN_CBAM algorithm includes:
[0092] The first convolutional layer is used to extract the features of the normalized vehicle speed information and vehicle position information;
[0093] The CBAM module is used to perform channel attention and spatial attention processing on the feature map;
[0094] The second convolutional layer is used to further process the feature map;
[0095] The fully connected layer is used to generate the state value and the action advantage value.
[0096] Please refer to Figures 1 - 4 , in step 4, the process of the multi-agent traffic signal cooperative control method includes:
[0097] Initialize the state space, action space, main network, and target network;
[0098] The agent makes decisions using the ε-greedy strategy to balance exploration and exploitation;
[0099] After the agent executes the selected action, the environment feedbacks the corresponding information and stores the experience sequence in the Replay buffer;
[0100] According to the priority experience replay strategy, extract experience data from the experience replay pool, calculate the loss function, and update the main network parameters.
[0101] In step 4, the calculation formula of the loss function is:
[0102]
[0103] Among them, r t is the reward value, γ is the discount factor, Q(s t ,a; θ) is the action value function of the main network, θ - is the parameter of the target network, represents the action corresponding to the highest action value output by the online network in state s t+1 ; Represents the action value output by the target network according to the action selected by the online network in state s t+1
[0104] The multi-agent traffic signal control method based on the spatio-temporal convolutional attention mechanism of the present invention optimizes the network structure, incorporates the convolutional attention mechanism, adds the attention mechanism behind each convolutional layer of the initial network, can better focus on the states of vehicles at adjacent intersections, and makes the signal control strategy more flexible and proactive. At the same time, the improved network is applied to the multi-agent road network to achieve cooperative control. It includes the following steps:
[0105] Step 1: Model the multi-intersection traffic network of a region as a multi-agent system, and define the state, action, and reward. The steps for defining the traffic state of the regional intersection are as follows:
[0106] (1) As Figure 1 shown, regard the traffic environment as a traffic grid composed of multiple intersections. Each intersection has four directions. The edge of each direction consists of four lanes, namely the straight lane, the left-turn lane, and the right-turn lane. In the given traffic grid, each intersection has an agent to control the traffic signal, and n1, n2, n3, n4 are its adjacent intersections. Each agent has at least 2 neighbors and at most 4 neighbors. For any adjacent intersection j of agent i, (i, j) ∈ I (i,j) .
[0107] In the environment of multiple intersections, the state definition space of intersection i is defined as:
[0108]
[0109] Among them, represents the position matrix of vehicles in the incoming lane l of intersection i in the road network, represents the speed matrix of vehicles in the incoming lane l of intersection i. The current joint state space of intersection i is composed of the current agent's local state space and the state spaces of neighboring agents, and can be expressed as:
[0110]
[0111] Among them, represents the local state information of intersection i at time t; represents the set of adjacent intersection state information; N i is the set of neighbor nodes as shown in formula (1).
[0112] (2) The action space is as Figure 2 As shown, it includes four phase spaces, which are, from left to right, straight and right turn in the north-south direction (NSG), left turn in the north-south direction (NSLG), straight and right turn in the east-west direction (EWG), and left turn in the east-west direction (EWLG).
[0113] (3) Use multi-objective optimization reward as the reward representation of the reinforcement learning algorithm. The reward function is defined as:
[0114]
[0115] Among them, q i represents the sum of the vehicle queue lengths of all lanes at the intersection at the previous moment; d i represents the sum of the waiting times of all lanes at the intersection at the previous moment; L jt is the red light duration of lane j at time t, C is the maximum red light time that the driver can tolerate. When the red light durations of all lanes are less than the maximum tolerance time, the result of the following penalty term is zero. If there are some lanes whose red light durations exceed the maximum tolerance time, a certain degree of penalty will be given at this time.
[0116] Use the weighted sum method to combine the agent reward with the rewards of adjacent domains, and introduce a spatial discount factor to make the reward value of the agent in the neighborhood positively correlated with the distance between the two agents. The closer the distance, the greater the reward value of the adjacent agent obtained. The spatial discount factor is expressed as:
[0117]
[0118] Among them, represents the distance between node i and node j, j ∈ N i ; is the discount coefficient.
[0119] The agent reward function is defined as the sum of the reward obtained from the current local environment and the rewards of all agents in the neighborhood, which is expressed as:
[0120]
[0121] Among them, is the local reward obtained by the agent; is the reward information of the adjacent agent j obtained by the agent, is the spatial discount factor.
[0122] Step 2: Use the convolutional attention mechanism network to preprocess the state. The model effectively highlights important features and suppresses redundant information by introducing two sub-modules, channel attention and spatial attention, to weight the channel dimension and spatial dimension of the state feature map respectively.
[0123] The present invention integrates channel information into the feature map by applying average pooling and max pooling operations, and uses a shared multi-layer perceptron to separately generate channel attention maps. These generated attention maps will be added element-wise, and the Sigmoid function is used to obtain the weight information for each channel. The calculation formula (6) is as follows:
[0124]
[0125] where, and respectively represent the average pooling and max pooling operation numbers in the channel dimension. W 0 ∈R C / r×C and W 1 ∈R C×C / r represent the weight parameters of the shared multi-layer perceptron. σ represents the Sigmoid activation function.
[0126] To calculate the spatial attention, the present invention first applies average pooling and max pooling operations along the spatial dimension to obtain the spatial average and maximum value at each spatial position. Then, these average values and maximum values are combined by concatenation. Next, a convolution operation is performed to generate a three-dimensional spatial attention map. Finally, the Sigmoid function is used to obtain the weight information. The calculation formula (7) is as follows:
[0127]
[0128] where, and respectively represent the average pooling and max pooling operations in the spatial dimension. f 7×7 represents a kernel size of 7×7. σ represents the Sigmoid activation function.
[0129] Step 3: Design an intersection signal control neural network model based on the spatio-temporal attention mechanism (Dueling Double Deep Q Network with CBAM, D3QN_CBAM), as Figure 3 shown.
[0130] D3QN_CBAM allows the agent to automatically focus on the key aspects of the state (such as vehicles approaching intersections) without being disturbed by noise or irrelevant features. The state in this system consists of two parts: normalized vehicle speed information and vehicle position information. First, the normalized vehicle speed information and vehicle position information are fed into the first convolutional layer to extract relevant features, thereby generating a feature map; this feature map is input into the CBAM module. In the CBAM module, the feature map is processed by channel attention and spatial attention, enabling the model to pay more attention to important features. The channel attention mechanism enhances the features in each channel, while the spatial attention mechanism captures important information at different positions.
[0131] The feature map processed by the CBAM module is passed to the second convolutional layer for further processing. The features obtained through convolutional operations are again input into the CBAM module to enhance the attention feature weights. Subsequently, the output data is flattened and concatenated with the phase state information. Finally, these concatenated data are processed through two fully connected layers and a ReLU operation to generate an output for use by two subsequent branches. The first branch is used to calculate the state value, while the second branch is used to calculate the action advantage value, and for each operation, its dimension matches that of the output layer.
[0132] Step 4: The process of the multi-agent traffic signal collaborative control method is as Figure 4 shown.
[0133] Initialize the state space, action space, main network, and target network, initializing the network structure and parameters. In addition, it is also necessary to initialize the capacities of the experience replay pool and the training pool. During the interaction between the agent and the environment, first, the agent will obtain the current state of the environment and the state information of neighboring intersections to form a joint state and input it into the neural network.
[0134] The agent uses the ε-greedy strategy to make decisions, balancing exploration and exploitation. After taking the corresponding action, the traffic signal returns a reward value, and each obtained reward value is saved and calculated, and then the next iteration begins. After the agent executes the selected action, the environment will feedback the corresponding information according to the result of this action. Specifically, the agent receives the state s t , selects the operation a t . During the execution of a t action, the effective phase in a cycle is preferentially activated, and then the reward r t is obtained, and it enters the new state s t+1 , and the experience sequence (s t , a t , r t , st+1 It is stored in the Replay buffer.
[0135] According to the priority experience replay strategy, B groups of experience data are sampled from the experience replay pool, and the loss function L(θ) is calculated using formula (8).
[0136]
[0137] Among them, r t is the reward value, γ is the discount factor, Q(s t , a; θ) is the action value function of the main network, and θ - are the parameters of the target network;
[0138] Subsequently, the parameters θ of the main network are updated through TD-error, and the priority of the transfer sequence is adjusted at the same time. After the main network is trained for a fixed number of iterations, the parameters in the target network are synchronously updated, that is, θ - = θ.
[0139] In an embodiment of the present invention, a multi-agent traffic signal control system based on a spatio-temporal convolutional attention mechanism includes a traffic signal control module for executing the multi-agent traffic signal dynamic cooperative control method based on the spatio-temporal convolutional attention mechanism described above, and further includes:
[0140] A traffic simulation environment module for simulating a multi-intersection traffic network and providing traffic flow data and status information;
[0141] An agent module for interacting with the traffic simulation environment module, executing a signal control strategy, and collecting feedback information;
[0142] An experience replay module for storing the experience data of the interaction between the agent and the traffic simulation environment and for training the traffic signal control module.
[0143] In summary, the present invention designs a multi-agent signal lamp coordination control method based on a spatio-temporal convolutional attention mechanism and deep reinforcement learning, enabling local agents to consider the dynamic changes and responses of neighboring agents. Agents can exchange state information, and in order to better focus on the states of vehicles at neighboring intersections, a spatio-temporal convolutional attention mechanism is introduced. When agents interact with the environment, they consider the influence of the states and rewards of neighboring agents. After reaching the Nash equilibrium through Markov games, the actions generated can be expressed as the best action responses to the current global state. Agents improve the coordination control effect of the algorithm by sharing state information and reward information within the neighborhood at intersections during the decision-making process.
[0144] The method was verified in the SUMO traffic simulation system, including two scenarios (peak traffic flow and off-peak traffic flow), as follows:
[0145] Table 1 Evaluation Metrics of Each Algorithm under Low Peak Flow
[0146]
[0147] In the off-peak traffic scenario, compared with the D3QN algorithm, the average queue length is shortened by 19.6%, the average waiting time of vehicles is reduced by 22.6%, and the average driving speed is increased by 8.8%.
[0148] Table 2 Evaluation Metrics of Each Algorithm under High Peak Flow
[0149]
[0150] In the high-flow traffic scenario, compared with the D3QN algorithm, the average queue length is shortened by 17.8%, the average waiting time of vehicles is reduced by 39.6%, and the average driving speed is increased by 17.2%.
[0151] The experimental results show that under the same traffic flow conditions, compared with the D3QN algorithm and the Webster method, the improved algorithm has better performance in terms of both control effect and convergence speed.
[0152] It should be noted that in the present invention, it should be understood that although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism, characterized in that: The specific steps include: Step 1: Model the multi-intersection traffic network in the region as a multi-agent system and define states, actions, and rewards; Step 2: Use the convolutional attention mechanism network to preprocess the state. By introducing two sub-modules, channel attention and spatial attention, the channel dimension and spatial dimension of the state feature map are weighted respectively to highlight important features and suppress redundant information. Step 3: Design an intersection signal light control neural network model based on a deep reinforcement learning algorithm with a spatiotemporal attention mechanism, wherein the model includes two convolutional layers and a CBAM module, and the CBAM module is used to enhance important features in the feature map; Step 4: Train the multi-intersection signal coordination control model through the multi-agent collaborative control algorithm. When the agent interacts with the environment, the influence of the state and reward of the neighboring agents is considered to generate the best action response to the current global state.
2. The multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism according to claim 1 is characterized in that: In step 1, the definition of regional intersection traffic status includes: The traffic environment is considered as a traffic grid consisting of multiple intersections. Each intersection has four directions, and the edge of each direction consists of four lanes, namely the straight lane, the left turn lane, and the right turn lane. The state space of each intersection is defined as: in, represents the position matrix of the vehicle in the incoming lane l of intersection i in the road network, represents the velocity matrix of vehicles entering lane l at intersection i. The current joint state space of intersection i consists of the local state space of the current agent and the state space of the neighboring agents, which can be expressed as: in, Represents the local state information of the intersection at time t; Represents the set of adjacent intersection status information; N i is the set of neighbor nodes.
3. The multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism according to claim 2 is characterized in that: In step 1, the action space contains four phase spaces, namely, turning right when going straight in the north-south direction, turning left when going straight in the north-south direction, turning right when going straight in the east-west direction, and turning left when going straight in the east-west direction.
4. The multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism according to claim 3 is characterized in that: In step 1, the reward function is defined as the multi-objective optimization reward, expressed as: Among them, q i It represents the sum of the lengths of vehicle queues in all lanes of the intersection at the previous moment; d i represents the sum of vehicles waiting in all lanes of the intersection at the previous moment; L jt is the duration of the red light in lane j at time t, and C is the maximum duration of the red light that the driver can tolerate.
5. The multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism according to claim 4 is characterized in that: In step 1, the agent reward is combined with the reward of the neighboring domain in a weighted sum manner. The spatial discount factor is introduced to make the reward value of the agent in the neighborhood positively correlated with the distance between the two agents. The closer the distance, the greater the reward value of the neighboring agent. The spatial discount factor is expressed as: in, Represents the distance between node i and node j, j∈N i ; is the discount factor; The agent reward function is defined as the sum of the reward obtained by the local environment at the current moment and the rewards of all agents in the neighborhood, expressed as: in, The agent obtains local rewards; is the reward information obtained by the agent from the neighboring agent j, is the space discount factor.
6. The multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism according to claim 1 is characterized in that: The calculation formula of the channel attention mechanism is: in, and Represents the number of average pooling and maximum pooling operations in the channel dimension respectively. W0∈R C / r×C and W1∈R C×C / r represents the weight parameter of the shared multilayer perceptron. σ represents the Sigmoid activation function; The calculation formula of the spatial attention mechanism is: in, and Represent the average pooling and maximum pooling operations in the spatial dimension respectively. 7×7 Indicates that the kernel size is 7 × 7. σ represents the Sigmoid activation function.
7. The multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism according to claim 1 is characterized in that: In step 3, the network structure of the deep reinforcement learning algorithm includes: The first convolutional layer is used to extract the features of normalized vehicle speed information and vehicle position information; CBAM module, used to perform channel attention and spatial attention processing on feature maps; The second convolutional layer is used to further process the feature map; Fully connected layer, used to generate state value and action advantage value.
8. The multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism according to claim 1 is characterized in that: In step 4, the process of the multi-agent traffic signal collaborative control method includes: Initialize the state space, action space, main network and target network; The agent adopts the ε-greedy strategy to make decisions and balance exploration and exploitation; After the agent performs the selected action, the environment feeds back the corresponding information and stores the experience sequence in the Replay buffer; According to the priority experience replay strategy, experience data is extracted from the experience replay pool, the loss function is calculated and the main network parameters are updated.
9. The multi-agent traffic signal control method based on spatiotemporal convolutional attention mechanism according to claim 8 is characterized in that: In step 4, the loss function is calculated as: Among them, r t is the reward value, γ is the discount factor, Q(s t ,a;θ) is the action value function of the main network, θ - are the parameters of the target network, Indicates that the online network is in state s t+1 The action corresponding to the highest action value outputted below, Indicates that the target network is in state s t+1 Next, the action value is output based on the action selected by the online network.
10. A multi-agent traffic signal control system based on a spatiotemporal convolutional attention mechanism, comprising a traffic signal control module, for executing a multi-agent traffic signal dynamic collaborative control method based on a spatiotemporal convolutional attention mechanism according to any one of claims 1 to 9, characterized in that: Also includes: Traffic simulation environment module, used to simulate multi-intersection traffic network and provide traffic flow data and status information; The intelligent agent module is used to interact with the traffic simulation environment module, execute signal control strategies and collect feedback information; The experience replay module is used to store the experience data of the interaction between the intelligent agent and the traffic simulation environment, and is used to train the traffic signal control module.
Citation Information
Patent Citations
Traffic light control method and system based on multi-agent reinforcement learning in control area
CN115631638A
Deep reinforcement learning traffic signal control method based on self-attention mechanism
CN115762128A
Deep reinforcement learning traffic signal control method based on attention mechanism
CN117746651A
Apparatus and method for controlling traffic signals of traffic lights in sub-area by using reinforcement learning model
US20240013654A1
Cited By
Traffic signal cooperative control method based on multi-agent reinforcement learning
CN120340272A