Multi-intersection traffic signal control method fusing traffic flow prediction
By adopting heterogeneous graph representation and deep reinforcement learning algorithms in multi-intersection traffic signal control, combined with long-term and short-term memory neural networks and multi-scale spatiotemporal feature aggregation technology, the shortcomings of existing methods in utilizing spatiotemporal information are solved, and more robust and forward-looking traffic signal control is achieved, and the traffic efficiency of the traffic network is improved.
Patent Information
- Application Number
- CN202510291323.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multi-intersection traffic signal control method is difficult to effectively utilize the space-time information of the traffic network, resulting in insufficient utilization of timing characteristics and spatial structure characteristics when facing a complex dynamic traffic environment, lack of a global perspective, and it is difficult to deal with future traffic flow changes in advance.
The road network structure is modeled using heterogeneous graph representation, combined with Markov's decision-making process and deep reinforcement learning algorithm, a traffic flow prediction model based on long and short-term memory neural network is constructed, and a multi-scale spatiotemporal heterogeneous graph feature aggregation technology and a reward function for perceived spatiotemporal information is used to optimize traffic signal control.
By making full use of heterogeneous information and historical vehicle traffic data of the transportation network, the robustness and forward-looking nature of the signal control model are improved, and the traffic efficiency of the transportation network and the overall efficiency of the transportation system are improved.
Smart Images

Figure CN120071647A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of urban intelligent traffic signal control, and in particular relates to a multi-intersection traffic signal control method integrating traffic flow prediction. Background Art
[0002] With the continuous growth of the urban vehicle ownership, the problem of traffic congestion has become increasingly severe. In the densely trafficked urban central areas, the traffic flows between adjacent intersections affect each other, making it difficult for independent signal control to effectively solve the overall traffic problem. In recent years, traffic signal control methods based on multi-agent deep reinforcement learning algorithms have been proven to be able to effectively alleviate regional traffic congestion and promote the coordinated control of traffic signals at multiple intersections. In this context, we need to coordinate the traffic signals at multiple adjacent or interconnected intersection roadways in the urban road network through intelligent information interaction and signal decision-making mechanisms. This will enable us to maximize traffic throughput, alleviate congestion, and improve the overall efficiency of the traffic system.
[0003] Although the existing research on the problem of multi-intersection traffic signal coordinated control has made progress, there are still certain limitations: First, existing methods often have difficulty in comprehensively capturing the diversity and dynamics of traffic network state information. In the face of a complex and dynamic traffic environment, the insufficient utilization of temporal features and spatial structure features restricts their performance. Second, the action decision-making methods of existing agents mainly rely on the traffic state information of their own intersections for optimization, and fail to fully consider the global spatio-temporal information in the traffic network. This approach results in a lack of a global perspective and makes it difficult to anticipate future traffic flow changes in advance. Finally, existing methods mainly use global graphs to represent the traffic network structure, increasing the computational complexity of learning the heterogeneous graph structure across multiple intersections.
[0004] Therefore, how to improve a multi-intersection traffic signal control method integrating traffic flow prediction based on spatio-temporal deep reinforcement learning algorithm that can effectively utilize spatio-temporal information such as historical, current, and future information in the road network is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0005] The object of the present invention is to solve the above problems and provide a multi-intersection traffic signal control method integrating traffic flow prediction.
[0006] To achieve the above object, the present invention provides the following technical solution: A multi-intersection traffic signal control method integrating traffic flow prediction, comprising the following steps:
[0007] Step 1, model the road network structure using the heterogeneous graph representation method, determine the signal light nodes, intersection nodes in the heterogeneous graph of the road network, and the spatial association relationships between the nodes;
[0008] Step 2: Model the multi-intersection traffic signal control problem based on the Markov decision process, and formulate the state space S, action space A, and reward function R for the research object;
[0009] Step 3: Establish a historical data experience pool based on the dynamic traffic flow of the input traffic network, construct a traffic flow prediction model using a long short-term memory neural network, and train the traffic flow data in the experience pool;
[0010] Step 4: Build a joint simulation platform for multi-intersection traffic network simulation and optimization training based on software, train and optimize the deep reinforcement learning model, and realize the testing and verification of the proposed method.
[0011] In the above multi-intersection traffic signal control method integrating traffic flow prediction, the state space S is defined as the feature information of each signal light node obtained by the intelligent agent n at time step t ( : signal light state), intersection node feature information ( : lane occupancy rate), and lane node feature information set,
[0012] The action space is determined by the phase sequence and the corresponding green light duration. The proposed method adopts a fixed phase sequence that changes sequentially from phase 1 to 4. The intelligent agent realizes the dynamic optimization control of traffic by selecting the green light duration of the phase. The definition of the action space is shown in Equation (1),
[0013] (1)
[0014] where: action represents the set of actions that the intelligent agent n can select at time step t, which consists of the green light durations of the turning lanes from phase 1 to 4 at time step t (Phase 1: north-south - straight and right turn; Phase 2: north-south - left turn; Phase 3: east-west - straight and right turn; Phase 4: east-west - left turn). represents the signal light state, for example represents the green light duration of phase 1 at time step t.
[0015] The reward function R is used to evaluate the return value of taking a certain action in a specific state. The intelligent agent selects appropriate actions to maximize the long-term cumulative reward. Its definition is shown in Equation (2),
[0016] (2)
[0017] where, is an adjustable weight coefficient used to control the traffic flow parameter The degree of influence, N is the total number of intersections, i represents the number of lanes, and I represents the maximum number of lanes. represents the queuing length of the lane, represents the predicted traffic flow. For example, represents the queuing length of lane i at intersection n within time step t, represents the traffic flow on lane i at intersection j at time step t + 1. represents the betweenness centrality coefficient of intersection j, and its calculation formula is as follows,
[0018] (3)
[0019] In the formula, b and c represent any two intersections in the traffic network, represents the number of shortest paths from b to c, represents the number of shortest paths from b to c passing through intersection j.
[0020] In the above multi-intersection traffic signal control method for fusing traffic flow prediction, a traffic flow prediction model based on a long short-term memory neural network (LSTM) is constructed. In this model, the agent stores historical traffic flow data in the experience pool. After the experience pool is full, the LSTM neural network is used to analyze the time-series data of the above traffic flow and perform short-term prediction of future traffic flow. The core calculation process of LSTM is shown in Equation (4),
[0021] (4)
[0022] In the formula, is the output value of the forget gate in the LSTM neural network at time step t, determining how much information from the historical time step t - 1 is absolutely retained and passed to the current state The information, with a value between 0 and 1, is the weight matrix of the forget gate is the bias term of the forget gate, represents the activation function. is the output value of the input gate, determining how much new information is to be added to the current state in, represents the weight matrix of the input gate, represents the bias term of the input gate. is the candidate cell state at time step t, represents the non-linear activation function, represents the weight matrix of the candidate cell state gate, is the bias term of the candidate cell state gate. and represent the cell states at the current time step t and the historical time step t - 1 respectively. The output value of the output gate at time step t, which is used to determine how much information of the cell state is output as the output value of the current time step. Represents the weight matrix of the output gate. Represents the bias term of the output gate. Represents the hidden state at time step t, which is the output of the LSTM and is used for the next calculation.
[0023] In the above multi-intersection traffic signal control method for fusing traffic flow prediction, step four specifically includes the following content:
[0024] S1. Construct a multi-intersection traffic network joint simulation platform based on Python and SUMO, set traffic simulation parameters in the Python environment, and design a multi-intersection traffic signal control method for fusing traffic flow prediction using a deep reinforcement learning algorithm.
[0025] S2. Use the constructed joint simulation platform to train the model and conduct simulation tests in the simulated traffic road network environment to verify the effectiveness of the proposed method.
[0026] Compared with the prior art, the present invention provides a multi-intersection traffic signal control method for fusing traffic flow prediction, having the following beneficial effects:
[0027] 1. Different from other multi-intersection traffic signal control methods for fusing traffic flow prediction, the multi-intersection traffic signal control method for fusing traffic flow prediction of the present invention makes full use of diverse heterogeneous information of the traffic network, and the constructed heterogeneous graph representation model of the road network can make full use of the temporal information and spatial structure information of multiple time steps in the road network, which can improve the robustness of the signal control model.
[0028] 2. Different from other traffic signal control models that only optimize the signal timing plan, the multi-intersection traffic signal control method for fusing traffic flow prediction of the present invention makes full use of historical traffic flow data in the traffic network to construct a traffic flow prediction model, which can enable the signal decision-making to utilize the predicted future traffic flow information, making the action decision of the signal model more forward-looking.
[0029] 3. The multi-intersection traffic signal control method for fusing traffic flow prediction of the present invention designs a multi-scale spatio-temporal heterogeneous graph feature aggregation technology. This technology uses a long short-term memory neural network to predict the traffic flow of the traffic network. By applying a graph neural network, it captures the temporal and spatial structure information of multiple time steps from historical, current, and predicted future data, enriching the feature representation of the traffic network.
[0030] 4. The multi-intersection traffic signal control method for integrated traffic flow prediction proposes a reward function that perceives spatio-temporal information. This function uses betweenness centrality to evaluate the spatial importance of each intersection and introduces the predicted traffic flow information as a parameter of the function, which can effectively improve the ability of the agent to perceive spatio-temporal information. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram of the control method of the present invention;
[0032] Figure 2 Schematic diagram of the multi-intersection traffic network of the present invention;
[0033] Figure 3 Modeling the traffic network structure using the heterogeneous graph representation method in the present invention;
[0034] Figure 4 Schematic diagram of the multi-scale spatio-temporal heterogeneous graph feature aggregation technology in the present invention;
[0035] Figure 5 Schematic diagram of the deep reinforcement learning algorithm principle combined with traffic flow prediction in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0036] The following embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention.
[0037] Please refer to Figures 1-5 , a multi-intersection traffic signal control method for integrated traffic flow prediction, including the following steps:
[0038] Step 1: Model the road network structure using the heterogeneous graph representation method, determine the signal light nodes, intersection nodes, and lane nodes in the heterogeneous graph of the road network, and the spatial association relationships between the nodes;
[0039] Step 2: Model the multi-intersection traffic signal control problem according to the Markov decision process, and formulate a state space S, an action space A, and a reward function R for the research object;
[0040] Step 3: Establish a historical data experience pool based on the dynamic traffic flow of the input traffic network, use a long short-term memory neural network to construct a traffic flow prediction model, and train the traffic flow data in the experience pool;
[0041] Step 4: Build a joint simulation platform for multi-intersection traffic network simulation and optimization training based on software, train and optimize the deep reinforcement learning model, and realize the test and verification of the proposed method.
[0042] The state space S is defined as the feature information of each signal light node obtained by the agent n at time step t ( : Signal light status), intersection node feature information ( : Lane occupancy rate) and lane node feature information set.
[0043] The action space is determined by the phase sequence and the corresponding green light duration. The proposed method adopts a fixed phase sequence that changes sequentially from phase 1 to 4. The agent realizes the dynamic optimization control of traffic by selecting the green light duration of the phase. The definition of the action space is shown in Equation (1).
[0044] (1)
[0045] In the formula: action represents the set of actions that agent n can select at time step t, which consists of the green light durations on the turning lanes of phases 1 to 4 at time step t (Phase 1: North-South - Straight and right turn; Phase 2: North-South - Left turn; Phase 3: East-West - Straight and right turn; Phase 4: East-West - Left turn). represents the signal light status, for example represents the green light duration of phase 4 at time step t.
[0046] The reward function R is used to evaluate the return value of taking a certain action in a specific state. The agent selects appropriate actions to maximize the long-term cumulative reward. Its definition is shown in Equation (2).
[0047] (2)
[0048] Among them, is an adjustable weight coefficient used to control the influence degree of traffic flow parameters , N is the total number of intersections, i represents the number of lanes, and I represents the maximum number of lanes. represents the queue length of the lane, represents the predicted traffic flow. For example, represents the queue length of lane i in intersection n at time step t, represents the traffic flow of lane i in intersection j at time step t + 1. represents the betweenness centrality coefficient of intersection j, and its calculation formula is as follows,
[0049] (3)
[0050] In the formula, b and c represent any two intersections in the traffic network, represents the number of shortest paths from b to c, represents the number of shortest paths from b to c passing through intersection j.
[0051] Build a traffic flow prediction model based on the Long Short-Term Memory Neural Network (LSTM). In this model, the agent stores historical traffic flow data in the experience pool. After the experience pool is full, the LSTM neural network is used to analyze the time series data of the above traffic flow and make short-term predictions for future traffic flow. The core calculation process of LSTM is shown in Equation (4).
[0052] (4)
[0053] In the formula, is the output value of the forget gate in the LSTM neural network at time step t, which determines how much information from the historical time step t - 1 is retained and passed to the current state The information passed to the current state ranges from 0 to 1, is the weight matrix of the forget gate, is the bias term of the forget gate, represents the activation function. is the output value of the input gate, which determines how much new information is to be added to the current state is the weight matrix of the input gate, is the bias term of the input gate. is the candidate cell state at time step t, represents the non-linear activation function, is the weight matrix of the candidate cell state gate, is the bias term of the candidate cell state gate. and represent the cell states at the current time step t and the historical time step t - 1 respectively. represents the output value of the output gate at time step t, which is used to determine how much information of the cell state is output as the output value of the current time step, is the weight matrix of the output gate, is the bias term of the output gate. represents the hidden state at time step t, which is the output of the LSTM and is used for the next calculation.
[0054] In Step 4, build a joint simulation platform for multi-intersection traffic network simulation and optimization training to train and optimize the deep reinforcement learning model, which specifically includes the following contents:
[0055] Define the heterogeneous graph structure of the road network and build a Multi-Intersection Traffic Signal Control (MITSC) model;
[0056] Initialize the number of training rounds , the number of agents , the simulation time ;
[0057] Initialize the Markov elements of the MITSC problem , the relevant parameters of the MITSC model, and the traffic flow dataset;
[0058] Conduct experiments using traffic simulation software (SUMO), store historical traffic flow data, and perform predictions through a traffic flow prediction model based on long short-term neural networks;
[0059] Obtain traffic state data using sensors on the lanes in traffic simulation software (SUMO), such as signal node feature information ( : signal state), intersection node feature information ( : lane occupancy rate), and the set of lane node feature information .
[0060] Based on the constructed heterogeneous graph neural network, capture the temporal and spatial information of multiple time steps in the traffic network through multi-scale spatio-temporal heterogeneous graph feature aggregation technology.
[0061] Calculate the loss according to formula (5) and update the neural network parameters,
[0062] (5);
[0063] In the formula, represents the loss value of intersection n at time step t, E represents the expectation, represents the reward of intersection n at time step t, represents the weight coefficient. Among them, represents when intersection n is in state , takes action and neural network parameters the Q value at this time, represents the action with the maximum execution return value, then represents the Q value when the neural network parameters are at this time.
[0064] Calculate the action reward according to formula (6) and update the action value,
[0065] (6).
[0066] Among them, is an adjustable weight coefficient used to control the influence degree of traffic flow parameter , N is the total number of intersections, i represents the number of lanes, and I represents the maximum number of lanes. represents the queue length of the lane, represents the predicted traffic flow. For example, represents the queue length of lane i in intersection n at time step t It represents the traffic flow on lane i within intersection j at time step t+1. It represents the betweenness centrality coefficient of intersection j.
[0067] Through multiple rounds of training of the model, a multi-intersection traffic signal control method integrating traffic flow prediction is designed using the deep reinforcement learning algorithm. After experimental testing, the proposed method can output the optimal signal timing plan, effectively improving the traffic network's traffic efficiency.
[0068] As can be seen from the above technical solutions, the present invention discloses a multi-intersection traffic signal control method integrating traffic flow prediction designed using the deep reinforcement learning algorithm. Compared with the prior art, this method combines traffic flow prediction with the multi-intersection traffic signal control method, including the following beneficial effects:
[0069] Different from other multi-intersection traffic signal control methods integrating traffic flow prediction, the present invention makes full use of the diverse heterogeneous information of the traffic network and establishes a heterogeneous graph representation model of the road network, which can make full use of the temporal information and spatial structure information of multiple time steps in the road network, and can improve the robustness of the signal control model;
[0070] Different from other traffic signal control models, the present invention makes full use of the historical traffic flow data in the traffic network to construct a traffic flow prediction model, enabling the signal decision-making to utilize the predicted future traffic flow information and making the action decision of the signal model more forward-looking;
[0071] The present invention designs a multi-scale spatio-temporal heterogeneous graph feature aggregation technology, which uses a long short-term memory neural network to predict the traffic flow of the traffic network. By applying a graph neural network, it captures the temporal and spatial structure information of multiple time steps from historical, current, and predicted future data, enriching the feature representation of the traffic network;
[0072] The present invention proposes a new reward function that perceives spatio-temporal information. This function uses betweenness centrality to evaluate the spatial importance of each intersection and introduces the predicted traffic flow information as a parameter of this function, which can effectively improve the agent's ability to perceive spatio-temporal information.
[0073] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-intersection traffic signal control method integrating traffic flow prediction, characterized in that: The steps include: Step 1: Use heterogeneous graph representation to model the road network structure, determine the signal light nodes, intersection nodes and lane nodes of the heterogeneous graph in the road network, and the spatial association relationship between the nodes; Step 2: Model the multi-intersection traffic signal control problem based on the Markov decision process, and formulate the state space S, action space A and reward function R for the research object; Step 3: Establish a historical data experience pool based on the dynamic traffic flow of the input traffic network, use the long short-term memory neural network to build a traffic flow prediction model, and train the traffic flow data in the experience pool; Step 4: Build a joint simulation platform for multi-intersection traffic network simulation and optimization training based on the software, train and optimize the deep reinforcement learning model, and test and verify the proposed method.
2. A multi-intersection traffic signal control method integrating traffic flow prediction according to claim 1, characterized in that: The state space S is defined as the feature information of each traffic light node obtained by agent n at time step t ( : Traffic light status), intersection node feature information ( : Lane occupancy rate) and lane node feature information A collection of The action space is determined by the phase sequence and the corresponding green light duration. The proposed method adopts a fixed phase sequence that changes from phase 1 to 4. The agent realizes dynamic optimization control of traffic by selecting the green light duration of the phase. The definition of the action space is shown in formula (1): (1) Where: action represents the set of actions that agent n can choose at time step t, which consists of the duration of the green light on the turning lanes from phase 1 to 4 at time step t (Phase 1: North-South - go straight and turn right; Phase 2: North-South - turn left; Phase 3: East-West - go straight and turn right; Phase 4: East-West - turn left). Indicates the status of the signal light, for example represents the green light duration of phase 1 at time step t.
3. The reward function R is used to evaluate the reward value of taking an action in a specific state. The agent maximizes the long-term cumulative reward by selecting appropriate actions. Its definition is shown in formula (2): (2) in, is an adjustable weight coefficient used to control traffic flow parameters The degree of influence is given by N, N is the total number of intersections, i is the number of lanes, and I is the maximum number of lanes. represents the queue length of the lane, represents the predicted traffic flow, for example, represents the queue length on lane i in intersection n at time step t, .
4. Represents the traffic flow on lane i in intersection j at time step t+1. represents the betweenness centrality coefficient of intersection j, and its calculation formula is as follows: (3) Where b and c represent any two intersections in the traffic network. represents the number of shortest paths from b to c, Represents the number of shortest paths from b to c through intersection j.
5. The method for controlling traffic signals at multiple intersections integrating traffic flow prediction according to claim 2 is characterized in that: Construct a traffic flow prediction model based on long short-term memory neural network (LSTM). In this model, the agent uses historical traffic flow data The data is stored in the experience pool. When the experience pool is full, the LSTM neural network is used to analyze the time series data of the above traffic flow and make short-term predictions for future traffic flow. The core calculation process of LSTM is shown in formula (4). (4) In the formula, is the output value of the forget gate in the LSTM neural network at time step t, and how much information from the historical time step t-1 is absolutely retained Pass to current state The information has a value between 0 and 1. is the weight matrix of the forget gate, is the bias term of the forget gate, Represents the activation function. The output value of the input gate determines how much new information is added to the current state middle, represents the weight matrix of the input gate, Represents the bias term of the input gate. is the candidate cell state at time step t, represents a nonlinear activation function, represents the weight matrix of the candidate cell state gate, is the bias term for the candidate cell state gate. and Represent the cell states at the current time step t and the historical time step t-1 respectively. Represents the output value of the output gate at time step t, which is used to determine how much cell state information is output as the output value of the current time step. represents the weight matrix of the output gate, Represents the bias term of the output gate. Represents the hidden state at time step t, which is the output of LSTM and is used for the next step of calculation.
6. The method for controlling traffic signals at multiple intersections integrating traffic flow prediction according to claim 1 is characterized in that: The step 4 specifically includes the following contents: S1. Build a multi-intersection traffic network joint simulation platform based on Python and SUMO, set traffic simulation parameters in the Python environment, and use deep reinforcement learning algorithm to design a multi-intersection traffic signal control method integrating traffic flow prediction; S2. The model is trained using the constructed joint simulation platform and simulation tests are performed in a simulated traffic network environment to verify the effectiveness of the proposed method.
Citation Information
Cited By
Adaptive signal control method for malformed intersection based on LSTM-GNN
CN120636155A
Traffic jam prediction method and system, electronic equipment and computer program product
CN121191327A
Traffic congestion prediction method, system, electronic device and computer program product
CN121191327B