A deep reinforcement learning traffic signal control method based on attention mechanism

CN117746651BActive Publication Date: 2026-09-29SICHUAN TIANAO KONGTIAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311794316.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2026-09-29
Estimated Expiration
2043-12-25

AI Technical Summary

Technical Problem

[0005]针对背景技术所提出的问题,本发明目的在于一种基于注意力机制的深度强化学习交通信号控制方法,解决了多交叉口的城市交通拥堵的问题

Benefits of technology

[0040]构建交通道路路网模型,获得各个道路每个车道的车辆交通信息;分析当前道路路网模型,建立以各个交叉口为代理的多智能体深度学习框架,设定抽象定义及集合;采用去中心化思想,基于D3QN增强学习基础网络结构,构建Q学习离线策略以及γ注意力奖励策略;基于γ注意力奖励策略,设计基于注意力机制的修正回放数据缓冲层算法;选取模拟数据和实际检测交通数据,采用Colight算法中超参数邻居作用域确定各个交叉口代理的邻居数目初始化,根据交通流实际数据带入进行仿真迭代,快速得到最优决策仿真结果,解决了多交叉口的城市交通拥堵的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117746651B_ABST
    Figure CN117746651B_ABST
Patent Text Reader

Abstract

The application discloses a kind of deep reinforcement learning traffic signal control methods based on attention mechanism, constructs traffic road network model, obtains the vehicle traffic information of each lane of each road;Analysis of the current road network model, establish the multi-agent deep learning framework with each intersection as agent, set abstract definition and set;Adopt the decentralized thought, based on D3QN reinforcement learning basic network structure, construct Q learning offline strategy and gamma attention reward strategy;Based on gamma attention reward strategy, design the correction playback data buffer layer algorithm based on attention mechanism;Select simulation data and actual detection traffic data, determine the number of each intersection agent neighbors initialization using the neighbor scope of Colight algorithm hyperparameter, according to the actual data of traffic flow is brought into simulation iteration, quickly obtain the optimal decision simulation result, solve the problem of urban traffic congestion of multiple intersections.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of traffic signal control technology, and more specifically to a deep reinforcement learning-based traffic signal control method based on an attention mechanism. Background Technology

[0002] Currently, urban traffic congestion is becoming increasingly serious, resulting in huge economic costs and wasted time. Traffic congestion is caused by a variety of factors, such as vehicle overloading and poor lane structure design. Some factors require complex policies or long-term planning. Effective traffic signal control is the most direct and cost-effective way to improve lane efficiency and alleviate traffic congestion.

[0003] Traffic signal control is crucial for building smart cities. With the development of deep reinforcement learning technology, many studies have applied it to traffic signal control at single intersections. However, urban traffic typically consists of multiple intersections, and the interactions between these intersections should not be ignored during modeling. Analyzing only single intersections will not alleviate urban traffic congestion; the city will remain congested.

[0004] Therefore, there is an urgent need for a deep reinforcement learning-based traffic signal control method that comprehensively analyzes multiple intersections to alleviate urban traffic congestion. Summary of the Invention

[0005] In response to the problems raised in the background technology, the present invention aims to provide a deep reinforcement learning-based traffic signal control method based on attention mechanism, which solves the problem of urban traffic congestion at multiple intersections.

[0006] This invention is achieved through the following technical solution:

[0007] This invention provides a deep reinforcement learning-based traffic signal control method based on an attention mechanism, comprising the following steps:

[0008] Step S1: Construct a traffic road network model and obtain vehicle traffic information; construct a multi-agent deep learning model using the traffic road network model and the vehicle traffic information.

[0009] Step S2: Construct a Q-learning strategy and a γ-attention reward strategy based on the D3QN reinforcement learning network structure, and modify the multi-agent deep learning model through the Q-learning strategy and the γ-attention reward strategy;

[0010] Step S3: Obtain simulated data and / or historical data, determine training parameters using the Colight algorithm, and input the training parameters and the simulated data and / or historical data into the corrected multi-agent deep learning model for training;

[0011] Step S4: Obtain actual traffic flow data and input the actual traffic flow data into the trained multi-agent deep learning model to obtain a traffic signal control strategy.

[0012] In the above technical solution, a traffic road network model is constructed based on open-source OSM data or actual data. Vehicle traffic information for each lane of each road is obtained through video analysis combined with fusion data interfaces such as test radar. The current road network model, especially road intersections, is analyzed, and a multi-agent deep learning framework (MADRL) is established with each intersection as an agent. Then, based on the internal agent relationships, abstract definitions and sets are defined for environmental state sets, action sets, environmental state transition probabilities, reward value functions, discount value functions, and learning rules. A decentralized approach is adopted, based on the D3QN reinforcement learning network structure, to construct an offline Q-learning strategy that combines exploration and consistency, as well as a γ-attention reward strategy. Based on the γ-attention reward strategy, a correction playback data buffer layer algorithm based on an attention mechanism is developed and designed for the actual system. Simulated data and actual detected traffic data that conform to the intersection road information and road network model are selected. The number of neighbors for each intersection agent is initialized using the hyperparameter neighbor scope in the Colight algorithm. Simulation iterations are performed based on actual traffic flow data to quickly obtain the optimal decision simulation results, thus solving the problem of urban traffic congestion at multiple intersections.

[0013] In one optional embodiment, constructing the traffic road network model includes:

[0014] Road network files are obtained through OSM (OpenStreetMap), and road network data is obtained by processing the road network files using the JSOM open-source software netconvert.

[0015] Modify the road network data using netedit;

[0016] Configure information on the modified road network data; the information configuration includes configuring intersection object information for all intersections and road segment attributes for all roads within the scope of the road network data.

[0017] In one alternative embodiment, the modification includes: deleting irrelevant road and river information from the road network data and improving the road network data.

[0018] In one alternative embodiment, the intersection object information includes an encoding, name, type, and coordinate values.

[0019] In one optional embodiment, the road segment attributes include road segment number, road segment name, lane driving direction, intersection codes before and after, restriction information, and normal driving speed.

[0020] In one optional embodiment, obtaining vehicle traffic information includes:

[0021] Obtain the average traffic flow, number of vehicles waiting, and average vehicle speed for each road at each intersection;

[0022] The traffic flow information body for each intersection is set, which includes the intersection number, time node, average traffic flow at the intersection, number of vehicles waiting at the intersection, and average speed of vehicles traveling at the intersection.

[0023] In one optional embodiment, constructing a multi-agent deep learning model using the traffic road network model and the vehicle traffic information includes:

[0024] Analyze the traffic road network model, establish an independent agent at each intersection, and set the agent's environment state, actions, environment state transition probability, reward function, discount value function, and learning rules;

[0025] The agents at each intersection are integrated to generate a multi-agent deep learning model, wherein the multi-agent deep learning model is <O,A,P,R,π,γ>.

[0026] In one alternative embodiment, constructing the γ attention reward policy includes:

[0027] Construct an attention component and construct an attention reward based on the attention component;

[0028] The attention construction part includes:

[0029] Each agent node is used to obtain observation values ​​through a multi-layer sensing mechanism;

[0030] The hidden parameters are obtained by calculating the weights of the corresponding proxy nodes based on the neighboring nodes of the proxy node, and then the hidden parameters are normalized.

[0031] The normalized hidden parameters are calculated using the ReLU activation function to obtain the performance parameters of the proxy node.

[0032] In an alternative embodiment, the attention portion is constructed... Attention rewards include:

[0033] Construct spatial difference formulas and attention score update formulas based on the attention component;

[0034] Design an attention-based algorithm for a modified replay data buffer layer.

[0035] In one optional embodiment, the actual traffic flow data is input into the trained multi-agent deep learning model to obtain a traffic signal control strategy, including:

[0036] The actual traffic flow data is input into the trained multi-agent deep learning model to obtain the optimal action sequence.

[0037] The optimal action sequence is combined with the agent node of the intersection corresponding to the environmental state sequence and mapped to obtain the traffic light control phase;

[0038] The traffic signal control strategy is determined based on the traffic light command phase.

[0039] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0040] A traffic network model was constructed to obtain vehicle traffic information for each lane on each road. The current road network model was analyzed, and a multi-agent deep learning framework with each intersection as an agent was established, defining abstract concepts and sets. A decentralized approach was adopted, based on the D3QN reinforcement learning network structure, to construct a Q-learning offline strategy and a γ-attention reward strategy. Based on the γ-attention reward strategy, an attention-based correction and replay data buffer layer algorithm was designed. Simulated data and actual traffic data were selected, and the number of neighbors for each intersection agent was initialized using the hyperparameter neighbor scope in the Colight algorithm. Simulation iterations were performed based on actual traffic flow data to quickly obtain the optimal decision simulation results, thus solving the problem of urban traffic congestion at multiple intersections. Attached Figure Description

[0041] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:

[0042] Figure 1 This is a schematic diagram of traffic intersection phase types provided in an embodiment of the present invention;

[0043] Figure 2 This is a diagram showing the relationship between the action set and the environment state set of each proxy node provided in this embodiment of the invention.

[0044] Figure 3 This is a schematic diagram of the network structure provided in an embodiment of the present invention;

[0045] Figure 4 A flowchart illustrating the Q-value learning and updating process is provided for embodiments of the present invention.

[0046] Figure 5 A flowchart illustrating the algorithm for correcting the playback data buffer layer is provided for embodiments of the present invention;

[0047] Figure 6 This is a schematic diagram of the position matrix and velocity matrix of the input layer provided in an embodiment of the present invention. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0049] Example

[0050] This invention provides a deep reinforcement learning-based traffic signal control method based on an attention mechanism, wherein the deep reinforcement learning-based traffic signal control method includes the following steps:

[0051] Step S1: Construct a traffic road network model and obtain vehicle traffic information; construct a multi-agent deep learning model using the traffic road network model and the vehicle traffic information.

[0052] In one optional embodiment, constructing the traffic road network model includes:

[0053] Road network files are obtained through OSM (OpenStreetMap), and road network data is obtained by processing the road network files using the JSOM open-source software netconvert.

[0054] Modify the road network data using netedit;

[0055] Configure information on the modified road network data; the information configuration includes configuring intersection object information for all intersections and road segment attributes for all roads within the scope of the road network data.

[0056] The modifications include: deleting irrelevant road and river information from the road network data, and improving the road network data.

[0057] The intersection object information includes code, name, type, and coordinate values;

[0058] The road segment attributes include the road segment number, road segment name, lane direction of travel, intersection codes before and after, restriction information, and normal traffic flow speed.

[0059] In one optional embodiment, obtaining vehicle traffic information includes:

[0060] Obtain the average traffic flow, number of vehicles waiting, and average vehicle speed for each road at each intersection;

[0061] The traffic flow information body for each intersection is set, which includes the intersection number, time node, average traffic flow at the intersection, number of vehicles waiting at the intersection, and average speed of vehicles traveling at the intersection.

[0062] In this embodiment, vehicle traffic information for each lane on each road can be obtained through video analysis combined with fusion data interfaces such as test radar, including basic traffic flow information such as the number of vehicles waiting, average vehicle speed, and the lane to which the vehicle belongs.

[0063] In one optional embodiment, constructing a multi-agent deep learning model using the traffic road network model and the vehicle traffic information includes:

[0064] Analyze the traffic road network model, establish an independent agent at each intersection, and set the agent's environment state, actions, environment state transition probability, reward function, discount value function, and learning rules;

[0065] The agents at each intersection are integrated to generate a multi-agent deep learning model, wherein the multi-agent deep learning model is <O,A,P,R,π,γ>.

[0066] Furthermore, for each independent agent built at each intersection, the phase settings for each lane are set according to the actual traffic conditions at the intersection. Each lane includes straight-ahead, left-turn, right-turn, and non-U-turn lanes. The traffic intersection phase types are as follows: Figure 1 As shown, its phase type is abstractly represented.

[0067] In this embodiment, setting the environment state of the agent includes: the agent observing the number of vehicles in the approach lane of its corresponding intersection and constructing the environment state. In this context, each element in the environmental state is a multi-dimensional vector, with the direction of the vector being the lane direction and the number of vectors being the normalized result of the number of vehicle queues.

[0068] The normalized result for the number of vehicles in the queue includes the normalized average vehicle speed and the number of vehicles. The formula for the normalized average vehicle speed is as follows: V max This is the maximum speed limit for this section of road.

[0069] Setting the agent's actions includes: combining the traffic light scheduling application, setting the actions as a traffic phase setting sequence, wherein the actions are constructed as follows:

[0070] Among them, the environment status of each proxy node itself. With action Relationship diagram as follows Figure 2 As shown.

[0071] The transition probability of the agent's environment state is set as follows: It indicates from the state The transition probability to the next possible state.

[0072] Wherein, the agent's reward value Set to complete the action The system reward obtained subsequently can be determined in this embodiment by judging the environmental state. Generally, the waiting vehicle entering the lane is used as the initial reward judgment value.

[0073] Here, the learning rule π represents the agent learning method in deep reinforcement learning. The purpose of setting the learning rule π in traffic signal control is to reduce learning time and increase learning speed. For a single agent, it can be set as follows:

[0074] The discount value γ' is used in time difference calculation and is a numerical parameter.

[0075] The reward function is set as follows:

[0076]

[0077]

[0078] In the above formula, α represents the learning rate, which is a numerical parameter.

[0079] Step S2: Construct a Q-learning strategy and a γ-attention reward strategy based on the D3QN reinforcement learning network structure, and modify the multi-agent deep learning model through the Q-learning strategy and the γ-attention reward strategy;

[0080] The process of constructing the Q-learning strategy is as follows: Figure 4 As shown, firstly, various types of data are initialized, including: constructing N proxy nodes and initializing them; initializing the replay storage and its capacity; initializing the original replay storage and its capacity; initializing the action value function Q and its random weights; initializing the target action value function and its weights; and initializing the attention score.

[0081] Then, the algorithm iterates by judging T and episode. The algorithm iterates as follows: T is initially assigned a value of 0. The episode is judged. If the episode is less than M, it is incremented and the environment is updated to obtain the initial state of each agent node. t is judged. If t is less than T, t is incremented and an action set is selected for each environment and each agent node. The actions are sent to all parallel computing environments. The algorithm receives the environment state and reward value of all agent nodes and stores them in the original playback storage.

[0082] Finally, T is compared with min{steps / (time update + n)}. If it is greater than min{steps / (time update + n)}, the processing for each agent node in each environment is repeated; otherwise, the data buffer is replayed and revised. After the data buffer is replayed and revised, the judgment on t is repeated.

[0083] In one alternative embodiment, constructing the γ attention reward policy includes:

[0084] Constructing the attention component and building upon the attention component Attention reward;

[0085] The attention construction part includes:

[0086] Each agent node is used to obtain observation values ​​through a multi-layer sensing mechanism;

[0087] The hidden parameters are obtained by calculating the weights of the corresponding proxy nodes based on the neighboring nodes of the proxy node, and then the hidden parameters are normalized.

[0088] The normalized hidden parameters are calculated using the ReLU activation function to obtain the performance parameters of the proxy node.

[0089] In this embodiment, the attention construction process includes first obtaining the environmental observation value z by applying a multilayer perceptron to each agent node. i =W l o i +b, and then obtain the hidden parameters based on the weights of neighboring nodes to that node. Then it is normalized. Impact value v ij =α ij z j After unification, using the ReLU activation function, it can be expressed as: Obtain the final performance parameters of the proxy node. Among them, W l 'b' and 'w' are parameters obtained by training the agent nodes at each intersection using historical data through a multilayer perceptron, and 'w' is the parameter obtained by training them. lLet b be the weight matrix and b be the offset vector.

[0090] In an alternative embodiment, the attention portion is constructed... Attention rewards include:

[0091] Construct spatial difference formulas and attention score update formulas based on the attention component;

[0092] Design an attention-based algorithm for a modified replay data buffer layer.

[0093] In this embodiment, construct Attention rewards include:

[0094] By using an attention mechanism, adding a fraction before the sum can yield a more accurate and interpretable formula for spatial difference. The formula for updating attention scores is: The reward value is R. i γ′ is the attention parameter, set to 0.9, and the target attention score. The Q-learning method, combined with a deviation strategy, is used to calculate the value from the target attention layer. and To learn and update the parameters of the process cycle through Q-value learning;

[0095] Design an attention-based algorithm to modify the replay data buffer layer. Define the buffer size as N, and iteratively update the replay data buffer layer data. The specific process is as follows: Figure 5 As shown.

[0096] It should be noted that the basic network structure is as follows: Figure 3 The diagram shows the overall structure. The reward value R update formula includes the Q-learning offline policy and the attention construction part. Reinforcement learning obtains the reward value R feedback value by updating the reward value R to make a decision. The attention construction part utilizes the attention mechanism to establish feedback based on the different effects caused by the spatial and temporal relationships of traffic data. This part runs in the Q-learning offline policy, storing historical data of past policy interactions with the environment in an experience pool. When updating the current policy, it samples from the experience pool and uses the obtained samples to update the policy. The resulting policy will continue to interact with the environment, and the data obtained from the interaction will also be stored in the experience pool. The purpose is to improve the training using static historical data.

[0097] Because this invention uses proxy mapping modeling for each traffic node, the input layer uses the lanes corresponding to each intersection, along with their included vehicle positions and normalized vehicle speed data, as a matrix. The vehicle speed normalization formula is spd = v / V_max, and the matrix is ​​similar to... Figure 6As shown, the traffic conditions in several directions of the road are combinations of these matrices in a certain order, and the parameters of convolutional layer 1 and convolutional layer 2 need to be fitted to approximate values ​​of the Q function.

[0098] Step S3: Obtain simulated data and / or historical data, determine training parameters using the Colight algorithm, and input the training parameters and the simulated data and / or historical data into the corrected multi-agent deep learning model for training;

[0099] In this embodiment, simulated traffic flow data or actual historical data with similar characteristics to or characteristic of this road network model are selected. Referring to the hyperparameter design in the Colight algorithm, training parameters are set, including an empirical replay unit size of 2000, a learning rate (α) of 0.001, a batch size of 20, a discount factor γ of 0.8, an epoch of 200, an iteration cycle of 200, and a simulation training period of 3600, thus initializing the algorithm training. The trained model is then used to access traffic flow data for simulation calculations. The optimal action sequence corresponding to the simulation results is mapped to the relevant traffic intersection proxy nodes based on the environmental state sequence to obtain the corresponding traffic light control phase information, guiding traffic control.

[0100] In this embodiment, simulated data or actual historical data of traffic flow information that is similar to or has the characteristics of this road network model are used;

[0101] Step S4: Obtain actual traffic flow data and input the actual traffic flow data into the trained multi-agent deep learning model to obtain a traffic signal control strategy.

[0102] In one optional embodiment, the actual traffic flow data is input into the trained multi-agent deep learning model to obtain a traffic signal control strategy, including:

[0103] The actual traffic flow data is input into the trained multi-agent deep learning model to obtain the optimal action sequence.

[0104] The optimal action sequence is combined with the agent node of the intersection corresponding to the environmental state sequence and mapped to obtain the traffic light control phase;

[0105] The traffic signal control strategy is determined based on the traffic light command phase.

[0106] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A deep reinforcement learning-based traffic signal control method based on attention mechanism, characterized in that, Includes the following steps: Step S1: Construct a traffic road network model and obtain vehicle traffic information; construct a multi-agent deep learning model using the traffic road network model and the vehicle traffic information; constructing the multi-agent deep learning model using the traffic road network model and the vehicle traffic information includes: Analyze the traffic road network model, establish an independent agent at each intersection, and set the agent's environment state, actions, environment state transition probability, reward function, discount value function, and learning rules; The agents at each intersection are integrated to generate a multi-agent deep learning model, wherein the multi-agent deep learning model is... ; O represents the environment state, and the environment state is constructed as A represents the action, and the action is constructed as follows: P represents the transition probability of the environmental state, set as follows: , indicating from state The probability of transitioning to the next state; R is the reward value, which is set to the value after the action is completed. The system reward obtained afterward; π is the learning rule, set as follows for a single agent. Constructing a gamma attention reward strategy includes: Constructing an attention component and constructing a γ attention reward based on the attention component; The attention construction part includes: Each agent node is used to obtain observation values ​​through a multi-layer sensing mechanism; The hidden parameters are obtained by calculating the weights of the corresponding proxy nodes based on the neighboring nodes of the proxy node, and then the hidden parameters are normalized. The normalized hidden parameters are calculated using the ReLU activation function to obtain the performance parameters of the proxy node; Step S2: Construct a Q-learning strategy and a γ-attention reward strategy based on the D3QN reinforcement learning network structure, and modify the multi-agent deep learning model through the Q-learning strategy and the γ-attention reward strategy; Step S3: Obtain simulated data and / or historical data, determine training parameters using the Colight algorithm, and input the training parameters and the simulated data and / or historical data into the corrected multi-agent deep learning model for training; Step S4: Obtain actual traffic flow data and input the actual traffic flow data into the trained multi-agent deep learning model to obtain a traffic signal control strategy.

2. The deep reinforcement learning-based traffic signal control method based on attention mechanism according to claim 1, characterized in that, Constructing a traffic road network model includes: Road network files are obtained through OSM, and road network data is obtained by processing the road network files using the JSOM open-source software netconvert. Modify the road network data using netedit; Configure information on the modified road network data; the information configuration includes configuring intersection object information for all intersections and road segment attributes for all roads within the scope of the road network data.

3. The deep reinforcement learning-based traffic signal control method based on an attention mechanism according to claim 2, characterized in that, The modifications include: deleting irrelevant road and river information from the road network data, and improving the road network data.

4. The deep reinforcement learning-based traffic signal control method based on an attention mechanism according to claim 2, characterized in that, The intersection object information includes code, name, type, and coordinate values.

5. The deep reinforcement learning-based traffic signal control method based on an attention mechanism according to claim 2, characterized in that, The road segment attributes include road segment number, road segment name, lane driving direction, intersection codes before and after, restriction information, and normal driving speed.

6. The deep reinforcement learning-based traffic signal control method based on an attention mechanism according to claim 2, characterized in that, Obtaining vehicle traffic information includes: Obtain the average traffic flow, number of vehicles waiting, and average vehicle speed for each road at each intersection; The traffic flow information body for each intersection is set, which includes the intersection number, time node, average traffic flow at the intersection, number of vehicles waiting at the intersection, and average speed of vehicles traveling at the intersection.

7. The deep reinforcement learning-based traffic signal control method based on attention mechanism according to claim 1, characterized in that, Constructing a γ-attention reward based on the aforementioned attention component includes: Construct spatial difference formulas and attention score update formulas based on the attention component; Design an attention-based algorithm for a modified replay data buffer layer.

8. The deep reinforcement learning-based traffic signal control method based on an attention mechanism according to claim 1, characterized in that, The actual traffic flow data is input into the trained multi-agent deep learning model to obtain traffic signal control strategies, including: The actual traffic flow data is input into the trained multi-agent deep learning model to obtain the optimal action sequence. The optimal action sequence is combined with the agent node of the intersection corresponding to the environmental state sequence and mapped to obtain the traffic light control phase; The traffic signal control strategy is determined based on the traffic light command phase.

Citation Information

Patent Citations

  • Deep reinforcement learning traffic signal control method combined with state prediction

    CN113963555A

  • Deep reinforcement learning traffic signal control method based on self-attention mechanism

    CN115762128A