Vehicle-signal cooperative signal control method based on double-layer AMOC

Through the double-layer AMOC vehicle-signal cooperative signal control method, combined with the NSGA-II and DDPG algorithms, efficient traffic flow optimization and coordination are achieved, solving the problems of high computational complexity and high real-time data requirements in existing technologies, and improving the efficiency and safety of the transportation system.

CN118692250BActive Publication Date: 2025-09-26SOUTHEAST UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410685909.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-30
Publication Date
2025-09-26
Estimated Expiration
2044-05-30

AI Technical Summary

Technical Problem

Existing multi-agent algorithms in regional signal control have high computational complexity, high real-time data requirements and lack of interpretability, making it difficult to achieve efficient traffic flow optimization and coordination.

Method used

A vehicle-signal cooperative signal control method based on a double-layer AMOC is adopted. Through information sharing and collaboration between the central coordinator and the lower-level intelligent agents, the signal control strategy is optimized in combination with the NSGA-II and DDPG algorithms to achieve coordinated control of signals and guidance of vehicles in the area.

Benefits of technology

It improves traffic smoothness and safety, reduces traffic congestion and vehicle delays, reduces transportation costs and energy consumption, enhances the adaptability and flexibility of the transportation network, and improves the explainability of the decision-making process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118692250B_ABST
    Figure CN118692250B_ABST
Patent Text Reader

Abstract

This invention discloses a vehicle-signal cooperative signal control method based on a two-layer AMOC. At the lower level, a single vehicle is used as an intelligent agent to perform cooperative control at a single intersection. At the higher level, a single traffic signal is used as an intelligent agent to coordinate traffic flow and optimize traffic efficiency across the entire area. Vehicles establish communication with traffic signals through vehicle-to-everything (V2X) communication, exchanging information such as position, speed, and steering intent to establish a single-point intersection control strategy. At the higher level, a collaborative multi-agent reinforcement learning algorithm is used to achieve cooperative control between traffic signals. This method achieves intelligent collaboration between vehicles and traffic signals, optimizing traffic flow, reducing congestion, and improving the efficiency and safety of the entire transportation system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of regional signal control and multi-agent reinforcement learning, and specifically relates to a vehicle-signal machine collaborative signal control method based on a double-layer AMOC. Background Art

[0002] Regional signal control has been widely researched and applied in the transportation sector. Regional signal control is a crucial component of intelligent transportation systems (ITS). With the development and widespread adoption of ITS technologies, regional signal control will be integrated with other ITS components (such as intelligent traffic management and intelligent vehicles) to achieve higher levels of traffic optimization and coordination. Regional signal control is increasingly based on real-time traffic data and vehicle demand. Traffic conditions are monitored and assessed in real time using traffic flow data, GPS data, and other sensor data. Signal control strategies are adjusted based on this real-time data to adapt to changes in traffic flow.

[0003] Multi-agent algorithms can be used to coordinate multiple signals in traffic signal control systems to optimize traffic flow, reduce delays, and improve traffic conditions. Each signal acts as an agent, communicating and coordinating with each other to jointly optimize the signal control strategy for the entire traffic network. Current multi-agent algorithms typically need to process a large number of state and action spaces and require extensive computation and optimization. This results in high computational complexity, requiring more computing resources and time. Furthermore, multi-agent algorithms require coordination and communication between multiple agents to jointly make decisions and adjust strategies. This involves the design of information sharing, communication protocols, and collaborative mechanisms, and requires addressing the issues of information transfer and coordination between agents.

[0004] Multi-agent reinforcement learning is an effective approach for regional traffic signal control. Patent CN113436443A discloses a method for accelerating reinforcement learning (RL) algorithms using an improved generative adversarial network (WGAN-GP) for regional traffic signal control. This method leverages the advantages of GANs in generating data and reinforcement learning algorithms in learning control strategies. However, this reinforcement learning algorithm requires a large amount of real-time data, and the decision-making process lacks interpretability, making it difficult to understand why the agents make specific decisions.

[0005] The Adaptive Multi-Objective Coordination (AMOC) algorithm coordinates the control strategies of multiple traffic signals to optimize traffic flow throughout an intersection or region. The AMOC algorithm can be combined with reinforcement learning algorithms to leverage the learning capabilities of intelligent agents to continuously optimize signal control strategies. Through interaction with intelligent agents, AMOC can gradually learn and improve control strategies to adapt to different traffic environments and demands. Summary of the Invention

[0006] Purpose of the invention: To address the above problems, the present invention proposes a vehicle-signal collaborative signal control method based on a double-layer AMOC, which enables the central control coordinator to coordinate the timing plans of various signal machines in the control area to obtain a regional timing plan based on the system optimal principle. At the same time, the signal machine guides the autonomous driving vehicles within the intersection range through the vehicle-signal machine network, thereby improving the efficiency and safety of the entire transportation system from point to surface.

[0007] Technical solution: To achieve the purpose of the present invention, the technical solution adopted by the present invention is: a vehicle-signal coordinated signal control method based on a double-layer AMOC, comprising the following steps:

[0008] Step 1: Using a single vehicle as a low-level agent and a single traffic signal as a high-level agent, initialize the strategies and value functions of the two-level agents, as well as the strategies and value functions of the central coordinator and each low-level coordinator.

[0009] Step 2: The low-level agent receives intersection environment information sent by the traffic light, including the current vehicle status, the status of surrounding vehicles, traffic signal status, road structure and speed limit information;

[0010] Step 3: The signal unit collects information from lower-level agents and uses it to update its own strategy and value function. Based on the intersection geometry data and the relationships between the agents, the AMOC algorithm is used to generate a single-point intersection signal control strategy.

[0011] Step 4: Create a network topology for the regional road network and connect all traffic lights at intersections within the region. Each traffic light transmits its own strategy and value function to the central coordinator, which uses the AMOC algorithm to update the strategy and value function to obtain the regional signal control strategy.

[0012] Step 5: The central coordinator uses the Option-Critic algorithm to provide strategy instructions to each agent, allowing the agent to make decisions under the road environment and traffic flow scenarios; it sends the regional signal control strategy to each signal machine, applies the parameters of the optimal strategy to the strategy network of each high-level agent, and updates the strategy of the high-level agent.

[0013] Furthermore, the signal unit collects information from lower-level agents and uses the received information to update its own strategy and value function. The steps include:

[0014] 1.1) Establishing a vehicle-signal connection: The low-level agent connects to the wireless network provided by the signal and communicates with the signal through a wireless communication protocol;

[0015] 1.2) Establish an application layer communication protocol on the wireless network. The lower-layer agent uses this protocol to encapsulate the value function and policy information into data packets and transmit them via the wireless communication protocol. The signal machine, as the receiver, receives the data packets and parses the value function and policy information.

[0016] 1.3) The vehicle’s state and the states of surrounding vehicles are converted into vector form and used as the input of the DDPG algorithm. Based on the initialized low-level agent strategy and value function, the traffic light’s strategy and value function are updated, and its own state, strategy action, and reward information are sent to the low-level coordinator.

[0017] Furthermore, based on the intersection geometry data and the relationships between the agents, the AMOC algorithm is used to generate a single-point intersection signal control strategy, which specifically includes:

[0018] 2.1) defining a state vector based on intersection geometry data and intersection signal status, wherein the state vector includes the number of lanes and vehicle positions;

[0019] The number of lanes and vehicle positions are discretized according to the lane channelization of each entrance. Each entrance corresponds to a processor, and finally the state vectors of all entrances are spliced ​​to obtain the intersection state vector;

[0020] 2.2) Define the action space, which is the timing plan for the signal;

[0021] The action space is defined as a three-dimensional matrix. The first dimension of the matrix represents the number of phases N, the second dimension represents the number of entrance lanes M, and the third dimension represents the passage order K. Each element represents the green light time G of the phase on the corresponding entrance lane. i , the action space A is represented as A=[N,M,K];

[0022] Set the yellow light duration Y and the full red light duration R; the entire signal cycle C is:

[0023]

[0024] When selecting an action, the agent determines an action by specifying the travel time and pass order of each phase on each entrance lane; the signal machine updates the signal timing plan based on the action selected by the agent;

[0025] 2.3) With the goal of minimizing the total delay of vehicles at the intersection and maximizing the efficiency of green light utilization, define a reward function r for the intersection;

[0026] r=(1-x1)*D+x2E

[0027] Where x1 and x2 are linear weight coefficients, D is the total vehicle delay, and E is the green light utilization efficiency. The total vehicle delay D is the time it takes for all vehicles in an intersection to pass through the intersection minus the free-flow time in a complete cycle period. The green light utilization efficiency E is the number of vehicles passing through the intersection divided by the saturation flow rate. The green light utilization rate and total vehicle delay are standardized and mapped to the range between 0 and 1.

[0028] 2.4) Define the state space;

[0029] NSGA-II is applied to the DDPG algorithm. Based on the reward function, the objective function of DDPG is split into green light utilization rate and total vehicle delay, which are used as the optimization objectives of NSGA-II respectively.

[0030] According to the NSGA-II algorithm process, NSGA-II selection and evolution operations are used to generate a set of Pareto front solutions. The Pareto front solutions have different green light timing schemes corresponding to different strategies. By evaluating all strategies and following the update rules of the DDPG algorithm, the optimal strategy is selected for update.

[0031] Furthermore, the state vector is defined in step 2.1 as follows:

[0032]

[0033] P i =[S2,S3];

[0034] Among them, S1 represents a single lane state vector, S2 represents a single entrance state vector, S3 represents the intersection state vector; F is the number of vehicles, V F is the speed of vehicle F, P1 is the lane position code of vehicle 1, P F is the lane position code of vehicle F, P i is the lane position code of vehicle i, Indicates the phase state; L1 is the first lane state, is the N1 lane state, N1 represents the total number of lanes at the intersection; E1 is the first entrance state, is the N2th entrance state, where N2 represents the total number of entrance lanes at the intersection; the lane position encoding represents the one-hot encoding vector of the position information of each vehicle, and the phase state is the green light duration of the lane phase.

[0035] Furthermore, the definition of the state space in step 2.4 specifically includes:

[0036] Initialize NSGA-II parameters, including population size, maximum number of iterations, crossover rate, and mutation rate;

[0037] Initialize DDPG parameters, including experience replay buffer size, batch size, discount factor, exploration noise, actor network learning rate, and critic network learning rate;

[0038] Initialize the DDPG network and experience replay buffer: Initialize the Actor network, initialize the Critic network, initialize the NSGA-II population, use DDPG to evaluate individual strategies and calculate fitness;

[0039] NSGA-II is used for selection and evolution operations, and the best individual is selected from the individuals generated by NSGA-II. The best individual is optimized by the DDPG algorithm, and the optimized and updated individual is used as the basis of the next generation population.

[0040] Furthermore, the regional signal control strategy is obtained, and the process is as follows:

[0041] 3.1) Use the adjacency matrix to create a topology for the road network in the area, and assign each node to the traffic lights at each intersection;

[0042] 3.2) Define the objective function:

[0043]

[0044] Among them, ω is the total reward, r i is the reward obtained at intersection i, where n represents the number of traffic lights in the area, i = 0, 1, …, n;

[0045] Constraints:

[0046] ω j+1 ≥ω j

[0047] ω j+1 -ω j ≤M

[0048] Among them, j is the current iteration number, ω j and ω j+1 are the Pareto front solutions obtained after the j-th and j+1-th iterations of the NSGA-II algorithm, i.e., the total reward; the optimal solution ω obtained at the end of the iteration is the maximum total reward, and M is the convergence threshold;

[0049] 3.3) The central coordinator uses the road network topology data to obtain the regional signal control strategy and runs the NSGA-II algorithm through the central coordinator to obtain the Pareto optimal solution set;

[0050] The NSGA-II algorithm uses simulated binary SBX operators and NDX operators to encode real numbers, obtains strategies through local search, and then performs global search optimization strategies to propose a hybrid operator:

[0051]

[0052] Among them, P n represents the usage probability of the NDX operator, P b represents the probability of using the SBX operator, and j represents the current number of iterations;

[0053] Finally, the tools.selSPEA2 function is used to select the Pareto optimal solution set from the population.

[0054] Furthermore, in step 1.2, use the Tornado library in Python to establish a WebSocket application layer connection and create a WebSocket handler class named SignalHandler;

[0055] When the vehicle establishes a WebSocket connection and sends a message, the on_message method is called to process the received message, and a response message is sent to the vehicle through the self.write_message method.

[0056] Furthermore, step 1.3 specifically includes:

[0057] 1.31) Convert the strategy and value functions into vectors, encapsulate them into data packets, and extend the SignalHandler class;

[0058] First, the low-level agent uses the on_message method to parse the received message and extract the policy and value function. Then, it uses the convert_to_vector function to convert the policy and value function into vector form. The converted vector data is encapsulated into a data packet. Finally, it uses json.dumps to convert the data packet into JSON format and sends it to the signaler using self.write_message.

[0059] 1.32) The signal machine receives JSON data, including policy and value function weight parameters. The signal machine parses the JSON formatted policy and value function. It then extracts the policy weights and value function weights, constructs a policy and value function network, and uses NumPy arrays to represent the weights. Finally, it uses a multilayer perceptron to construct a policy network using the policy weights and a value function network using the value function weights.

[0060] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:

[0061] This invention, through a vehicle-signal cooperative signal control method based on a dual-layer AMOC, can improve traffic flow and safety, achieve more precise and efficient signal control, thereby reducing traffic congestion and vehicle delays, and improving traffic flow and road safety. It also reduces transportation costs and energy consumption. The optimized signal control strategy can reduce vehicle parking time and queue time at intersections, lowering transportation costs and reducing vehicle fuel consumption, thereby conserving energy resources.

[0062] This invention combines the NSGA-II and DDPG algorithms to enhance the adaptability and flexibility of traffic networks and dynamically optimize and adjust traffic signal control strategies to better adapt to varying traffic flows and road conditions. By coordinating the search processes of multiple optimization algorithms within the timing solution implementation path and designing information sharing, communication protocols, and collaborative mechanisms, this method effectively improves solution quality and convergence speed while reducing reliance on real-time data. Furthermore, by designing the state representation and decision rules of intelligent agents, the interpretability of the decision-making process is enhanced, making the algorithm more understandable and acceptable.

[0063] Through information sharing and collaboration between intelligent entities and the guidance and management of a central coordinator, the present invention realizes intelligent coordination between vehicles and traffic signals, improves the intelligence level of urban traffic management, and actually achieves a balanced management and control result of the system. It provides more decision-making support and optimization solutions for guiding each vehicle and traffic signal to make the optimal system strategy in the future under fully L5 autonomous driving vehicles. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0065] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0066] This embodiment takes two low-level intelligent agents: car 1 and car 2, two high-level intelligent agents: machine 1 and machine 2, and a central coordinator as an example, wherein car 1 and car 2 are both at the intersection where machine 1 is located.

[0067] The vehicle-signal cooperative signal control method based on double-layer AMOC of the present invention is as follows: Figure 1 , including the following steps:

[0068] Step 1: Initialize the strategies a1, a2, A1, A2 and value functions of the two-layer agents (Car 1, Car 2, Machine 1, Machine 2), as well as the strategies and value functions of the central coordinator and each layer of lower-level coordinators. The lower layer uses a single vehicle as an agent, and the higher layer uses a single traffic signal as an agent. The strategy and value function initialization operations of the agents are as follows:

[0069] 1.1) Initialize the layers of the neural network using linear (fully connected) layers. The network has two hidden layers with 64 neurons. Initialize as follows:

[0070] The policy network consists of an input layer, two hidden layers, and an output layer.

[0071] The process expression from input layer to hidden layer is: h1=ReLU(W1x+b1);

[0072] The expression of the hidden layer to hidden layer process is:

[0073] The process from hidden layer to output layer is expressed as:

[0074] Where x is the input vector, output_dim is the output dimension, ReLU() is the rectified linear unit activation function, tanh() is the hyperbolic tangent activation function, W1 and W2 are weight matrices of size 64, and b1 and b2 are bias vectors of size 64.

[0075] The value network consists of an input layer, two hidden layers, and an output layer.

[0076] The process expression from input layer to hidden layer is: h1=ReLU(W1x+b1);

[0077] The expression of the hidden layer to hidden layer process is:

[0078] The process from hidden layer to output layer is expressed as:

[0079] 1.2) Define the forward propagation of the neural network. It applies the ReLU activation function to the outputs of the first two layers and the hyperbolic tangent (tanh) activation function to the output of the third layer. Because there are two hidden layers, there are two hidden layer activation functions. The forward propagation method is as follows:

[0080] The policy network consists of an input layer, two hidden layers, and an output layer.

[0081] Input vector: x;

[0082] The hidden layer activation process expression is: h1 = ReLU (W1x + b1);

[0083] The hidden layer activation process expression is: h2 = ReLU (W2h1 + b2);

[0084] The output layer activation process expression is: y = tanh (W ) h2+b ) ).

[0085] The value network consists of an input layer, two hidden layers, and an output layer.

[0086] The hidden layer activation process expression is: h1 = ReLU (W1x + b1);

[0087] The hidden layer activation process expression is: h2 = ReLU (W2h1 + b2);

[0088] The output layer activation process expression is: y = W ) h2+b ) .

[0089] The specific forward propagation process is:

[0090] 1.21) The input vector x is transformed linearly and activated by ReLU in the first hidden layer to obtain h1.

[0091] 1.22) h1 is transformed linearly and activated by ReLU in the second hidden layer to obtain h2.

[0092] 1.23) Finally, h2 undergoes a linear transformation and tanh activation in the output layer to obtain the final output y.

[0093] 1.3) Define the value function. It has three layers: an input layer and two hidden layers with 64 neurons. The output layer has only one neuron and does not use an activation function because it predicts a single scalar value.

[0094] 1.4) Initialize the dimensions of the state space and action space.

[0095] Step 2: Low-level intelligent vehicles 1 and 2 receive intersection traffic environment information 1, including the current vehicle status, the status of surrounding vehicles, traffic signal status, road structure and speed limit information.

[0096] The low-level intelligent agents, Car 1 and Car 2, are connected to the Wi-Fi network provided by Signal 1 and communicate with Signal 1 via the Wi-Fi 6 protocol. Signal 1 provides intersection road environment information services on the Wi-Fi network, and the intelligent agents request intersection road environment information from the signal through the Wi-Fi connection.

[0097] Step 3: Signal unit 1 collects information about the lower-level intelligent agents Car 1 and Car 2, and uses the received information about the lower-level intelligent agents Car 1 and Car 2 to update its own strategy and value function.

[0098] 1.1) First, establish the vehicle-signal connection;

[0099] The lower-level intelligent vehicles 1 and 2 are connected to the Wi-Fi network provided by signal 1 and communicate with the signal through wireless communication protocols.

[0100] 1.2) Establishing application layer communication protocols on Wi-Fi networks;

[0101] Low-level intelligent agents, Cars 1 and 2, encapsulate value functions and policy information into data packets using this protocol and transmit them via the Wi-Fi 6 protocol. Traffic light 1, as the receiver, receives these data packets and parses the value functions and policy information a1 and a2.

[0102] 1.21) Use the Tornado library in Python to establish a WebSocket application layer connection;

[0103] Create a WebSocket handler class named SignalHandler;

[0104] When vehicle 1 and vehicle 2 establish a WebSocket connection and send a message, the on_message method is called to process the received message and send a response message to the vehicle through the self.write_message method.

[0105] 1.3) The states c1 and c2 of the two vehicles are converted into vector form as the input of the DDPG algorithm. Based on the initialized low-level intelligent agents’ strategies a1 and a2 and value functions, each agent updates its own strategy and value function and sends its own state, strategy action and reward information to the low-level coordinator.

[0106] 1.31) Convert the strategy and value functions into vectors, encapsulate them into data packets, and extend the SignalHandler class;

[0107] First, the logic in the on_message method is expanded to parse the received message and extract the policy and value functions. Then, the policy and value functions are converted into vectors using the convert_to_vector function. Next, the converted vector data is encapsulated into a data packet and sent to the signaler. Finally, the data packet is converted into JSON format using json.dumps and sent to the signaler using self.write_message.

[0108] 1.32) First, receive JSON data, including the weight parameters of the policy and value function. The signal machine parses the JSON formatted policy and value function. Then, extract the policy weights and value function weights, construct the policy and value function network, and use NumPy arrays to represent the weights. Finally, use the multilayer perceptron to construct the policy network using the policy weights and the value function weights to construct the value function network.

[0109] Based on the existing geometric data of intersection 1 and considering the relationship between vehicles 1 and 2, the AMOC algorithm is used to generate the signal control strategy for intersection 1, and the effect and performance of the current signal control strategy are evaluated.

[0110] Step 4: Conduct network topology analysis on the regional road network and connect signal machines 1 and 2 at the intersections within the region.

[0111] Signal machine 1 and signal machine 2 transmit their respective strategies and value functions to the central coordinator, which updates the strategies and value functions using the same AMOC method to obtain the regional signal control strategy.

[0112] 2.1) A state vector is defined based on intersection geometry data and intersection signal status. The state vector includes the number of lanes and vehicle positions.

[0113] The number of lanes and vehicle positions are discretized according to the lane channeling of each entrance. Each entrance corresponds to a processor, and finally the state vectors of all entrances are spliced ​​to obtain the intersection state vector.

[0114]

[0115] P1=[S2,S ) ];

[0116] Among them, S1 represents a single lane state vector, S2 represents a single entrance state vector, and S3 represents the intersection state vector; F is the number of vehicles, V1 and V2 are the speeds of vehicle 1 and vehicle 2 respectively, and P1 and P2 are the lane position codes of vehicle 1 and vehicle 2 respectively. Indicates the phase state; the leftmost lane facing the intersection is the reference lane; L1 is lane state 1. For unified representation, the entrance in the due north direction is set as the reference entrance. If there is no due north entrance, the northwest entrance is used as the reference entrance. Then the order of each entrance lane is calibrated in a clockwise direction. E1 is entrance state 1; N1 represents the total number of lanes in the intersection; N2 represents the total number of entrance lanes in the intersection; the lane position code represents the one-hot encoding vector of the position information of each vehicle, and the phase state is the green light duration of the phase of this lane.

[0117] Phase state In binary encoding, phase states are represented using 0s and 1s. For example, if there are four phase states, a 4-bit binary number can be used, with each bit representing the state of a phase, 0 for closed and 1 for open. For example, a phase state of 1010 indicates that phases 1 and 3 are open, while phases 2 and 4 are closed. Assuming that phases 1 and 3 are open, both vehicles are in the leftmost lane of the northbound entrance, and intersection 1 is a crossroads; the state vector for intersection 1 is [2, V1, V2, 10001000, 10001000, 1010].

[0118] 2.2) Define the action space A1 of the signal at intersection 1. This action space is the timing plan of the signal.

[0119] The action space is defined as a three-dimensional matrix. The first dimension of the matrix represents the number of phases N, the second dimension represents the number of entrance lanes M, and the third dimension represents the passage order K. Each element represents the green light time G of the phase on the corresponding entrance lane. i , the action space A is expressed as A=[N,M,K].

[0120] Set the yellow light duration Y and the full red light duration R; the entire signal cycle is:

[0121]

[0122] When selecting an action, the agent determines an action by specifying the travel time and travel order of each phase on each entrance lane; the traffic light updates the signal timing plan according to the action selected by the agent to achieve optimal control of traffic flow.

[0123] 2.3) Define a reward function for the intersection with the goal of minimizing the total delay of vehicles at the intersection and maximizing the efficiency of green light utilization;

[0124] r=(1-x1)*D+x2E

[0125] Among them, x1 and x2 are linear weight coefficients, D is the total vehicle delay, and E is the green light utilization efficiency; the total vehicle delay D is the time it takes for all vehicles in the intersection to pass through the intersection in a complete cycle period minus the sum of the free-flowing time; the green light utilization efficiency E is the number of passing vehicles divided by the saturated flow rate.

[0126] Normalize the green light utilization rate and total vehicle delay separately, mapping both to the range between 0 and 1:

[0127]

[0128] Among them, S D is the standardized total vehicle delay, S is the total vehicle delay, S max is the maximum total vehicle delay, S min is the minimum total vehicle delay; S U is the standardized green light utilization rate, U is the green light utilization rate, U min is the minimum green light utilization rate, U max It is the maximum green light utilization rate.

[0129] 2.4) Define the state spaces C1 and C2 of intersection 1 and intersection 2.

[0130] After defining the state vector, action space, and reward function, NSGA-II is applied to the DDPG algorithm. Based on the reward function, the DDPG objective function is split into green light utilization rate and total vehicle delay, which serve as the optimization objectives of NSGA-II respectively.

[0131] According to the NSGA-II algorithm flow, NSGA-II selection and evolution operations are used to generate a set of Pareto front solutions representing different strategies. The Pareto front solutions have different green light timing schemes corresponding to different strategies. By evaluating all strategies and following the update rules of the DDPG algorithm, the optimal strategy is selected for update and improvement.

[0132] 2.41) Initialize NSGA-II parameters, including population size, maximum number of iterations, crossover rate, and mutation rate.

[0133] 2.42) Initialize DDPG parameters, including experience replay buffer size, batch size, discount factor, exploration noise, actor network learning rate, and critic network learning rate.

[0134] 2.43) Initialize the DDPG network and experience replay buffer: Initialize the Actor network, initialize the Critic network, initialize the NSGA-II population, use DDPG to evaluate individual strategies and calculate fitness.

[0135] 2.44) Using NSGA-II for selection and evolution operations, the best individual is selected from the individuals generated by NSGA-II. The best individual is optimized using the DDPG algorithm, and the optimized and updated individual serves as the basis for the next generation population.

[0136] The regional signal control strategy is sent to signal machine 1 and signal machine 2, the parameters of the optimal individual are directly applied to the strategy network of each signal machine 1 and signal machine 2, and the strategies A1 and A2 of signal machine 1 and signal machine 2 are updated.

[0137] 3.1) Use the adjacency matrix to perform topology analysis on the road network in the area, and assign each node to the traffic lights at each intersection.

[0138] 3.2) Define the objective function:

[0139]

[0140] Among them, ω is the total reward, r i is the reward obtained at intersection i, n is the number of traffic lights in the area, i = 0, 1, ..., n.

[0141] Constraints:

[0142] ω 891 ≥ω8

[0143] ω 891 -ω8≤M

[0144] Among them, j is the current iteration number, ω8 and ω 891 are the Pareto frontier solutions obtained after the jth and j+1th iterations of the NSGA-II algorithm, i.e., the total reward. The optimal solution ω obtained by the iteration is the maximum value of the total reward, and M is the convergence threshold. However, in order to shorten the training time, the Pareto optimal solution can be obtained after ω becomes stable. M is a smaller number, which can be ω×10 a2 .

[0145] 3.3) The central coordinator uses the road network topology data to obtain the regional signal control strategy and runs the NSGA-II algorithm through the central coordinator to obtain the Pareto optimal solution set.

[0146] The NSGA-II algorithm uses simulated binary SBX operators and NDX operators to encode real numbers, obtains strategies through local search, and then performs global search optimization strategies to propose a hybrid operator:

[0147]

[0148] Among them, P2 represents the usage probability of NDX operator, P Brepresents the probability of using the SBX operator, and j represents the current iteration number. Finally, the tools.selSPEA2 function is used to select the Pareto optimal solution set from the population.

[0149] Step 5: Based on the guidance of the central coordinator, the Option-Critic provides strategy instructions A1 and A2 to Signal 1 and Signal 2, enabling the agents to make decisions under specific road conditions and traffic flow scenarios.

Claims

1. The vehicle-signal cooperative signal control method based on double-layer AMOC is characterized by: The following steps are involved: Step 1: Using a single vehicle as a low-level agent and a single traffic signal as a high-level agent, initialize the strategies and value functions of the two-level agents, as well as the strategies and value functions of the central coordinator and each low-level coordinator. Step 2: The low-level agent receives intersection environment information sent by the traffic light, including the current vehicle status, the status of surrounding vehicles, traffic signal status, road structure and speed limit information; Step 3: The signal unit collects information from lower-level agents and uses it to update its own strategy and value function. Based on the intersection geometry data and the relationships between the agents, the AMOC algorithm is used to generate a single-point intersection signal control strategy. 1) defining a state vector based on intersection geometry data and intersection signal status, wherein the state vector includes the number of lanes and vehicle positions; The number of lanes and vehicle positions are discretized according to the lane channelization of each entrance. Each entrance corresponds to a processor, and finally the state vectors of all entrances are spliced ​​to obtain the intersection state vector; 2) Define the action space, which is the timing plan for the signal; The action space is defined as a three-dimensional matrix. The first dimension of the matrix represents the number of phases N, the second dimension represents the number of entrance lanes M, and the third dimension represents the passage order K. Each element represents the green light time G of the phase on the corresponding entrance lane. i , the action space A is represented as A=[N,M,K]; Set the yellow light duration Y and the full red light duration R; the entire signal cycle C is: When selecting an action, the agent determines an action by specifying the travel time and pass order of each phase on each entrance lane; the signal machine updates the signal timing plan based on the action selected by the agent; 3) With the goal of minimizing the total delay of vehicles at the intersection and maximizing the efficiency of green light utilization, a reward function r is defined for the intersection; r=(1-x1)*D+x2E Where x1 and x2 are linear weight coefficients, D is the total vehicle delay, and E is the green light utilization efficiency. The total vehicle delay D is the time it takes for all vehicles in an intersection to pass through the intersection minus the free-flow time in a complete cycle period. The green light utilization efficiency E is the number of vehicles passing through the intersection divided by the saturation flow rate. The green light utilization rate and total vehicle delay are standardized and mapped to the range between 0 and 1. 4) Define the state space; NSGA-II is applied to the DDPG algorithm. Based on the reward function, the objective function of DDPG is split into green light utilization rate and total vehicle delay, which are used as the optimization objectives of NSGA-II respectively. According to the NSGA-II algorithm process, NSGA-II selection and evolution operations are used to generate a set of Pareto front solutions with different green light timing schemes corresponding to different strategies. By evaluating all strategies and using the DDPG algorithm update rules, the optimal strategy is selected for update. Step 4: Create a network topology for the regional road network and connect all traffic lights at intersections within the region. Each traffic light transmits its own strategy and value function to the central coordinator, which uses the AMOC algorithm to update the strategy and value function to obtain the regional signal control strategy. Step 5: The central coordinator uses the Option-Critic algorithm to provide strategy instructions to each agent, allowing the agent to make decisions under the road environment and traffic flow scenarios; it sends the regional signal control strategy to each signal machine, applies the parameters of the optimal strategy to the strategy network of each high-level agent, and updates the strategy of the high-level agent.

2. The method according to claim 1, characterized in that The signal unit collects information from lower-level agents and uses the received information to update its own strategy and value function. The steps include: 1.1) Establishing a vehicle-signal connection: The low-level agent connects to the wireless network provided by the signal and communicates with the signal through a wireless communication protocol; 1.2) Establish an application layer communication protocol on the wireless network. The lower-layer agent uses this protocol to encapsulate the value function and policy information into data packets and transmit them via the wireless communication protocol. The signal machine, as the receiver, receives the data packets and parses the value function and policy information. 1.3) The vehicle’s state and the states of surrounding vehicles are converted into vector form and used as the input of the DDPG algorithm. Based on the initialized low-level agent strategy and value function, the traffic light’s strategy and value function are updated, and its own state, strategy action, and reward information are sent to the low-level coordinator.

3. The method according to claim 1, characterized in that In step 1), the state vector is defined as follows: P i =[S2,S ) ]; Among them, S1 represents a single lane state vector, S2 represents a single entrance state vector, S3 represents the intersection state vector; F is the number of vehicles, V F is the speed of vehicle F, P1 is the lane position code of vehicle 1, P F is the lane position code of vehicle F, P i is the lane position code of vehicle i, Indicates the phase state; L1 is the first lane state, is the N1 lane state, N1 represents the total number of lanes at the intersection; E1 is the first entrance state, is the N2th entrance state, where N2 represents the total number of entrance lanes at the intersection; the lane position encoding represents the one-hot encoding vector of the position information of each vehicle, and the phase state is the green light duration of the lane phase.

4. The method according to claim 1, wherein The definition of the state space in step 4) specifically includes: Initialize NSGA-II parameters, including population size, maximum number of iterations, crossover rate, and mutation rate; Initialize DDPG parameters, including experience replay buffer size, batch size, discount factor, exploration noise, actor network learning rate, and critic network learning rate; Initialize the DDPG network and experience replay buffer: Initialize the Actor network, initialize the Critic network, initialize the NSGA-II population, use DDPG to evaluate individual strategies and calculate fitness; NSGA-II is used for selection and evolution operations, and the best individual is selected from the individuals generated by NSGA-II. The best individual is optimized by the DDPG algorithm, and the optimized and updated individual is used as the basis of the next generation population.

5. The method according to claim 1, characterized in that Get the regional signal control strategy, the process is as follows: 3.1) Use the adjacency matrix to create a topology for the road network in the area, and assign each node to the traffic lights at each intersection; 3.2) Define the objective function: Among them, ω is the total reward, r i is the reward obtained at intersection i, where n represents the number of traffic lights in the area, i = 0, 1, …, n; Constraints: oh j+1 ≥ω j oh j+1 -oh j ≤M Among them, j is the current iteration number, ω j and ω j+1 are the Pareto front solutions obtained after the j-th and j+1-th iterations of the NSGA-II algorithm, i.e., the total reward; the optimal solution ω obtained at the end of the iteration is the maximum total reward, and M is the convergence threshold; 3.3) The central coordinator uses the road network topology data to obtain the regional signal control strategy and runs the NSGA-II algorithm through the central coordinator to obtain the Pareto optimal solution set; The NSGA-II algorithm uses simulated binary SBX operators and NDX operators to encode real numbers, obtains strategies through local search, and then performs global search optimization strategies to propose a hybrid operator: Among them, P n represents the probability of using the NDX operator, P represents the probability of using the SBX operator, and j represents the current number of iterations; Finally, the tools.selSPEA2 function is used to select the Pareto optimal solution set from the population.

6. The method according to claim 2, characterized in that In step 1.2, use the Tornado library in Python to establish a WebSocket application layer connection and create a WebSocket handler class named SignalHandler. When the vehicle establishes a WebSocket connection and sends a message, the on_message method is called to process the received message, and a response message is sent to the vehicle through the self.write_message method.

7. The method according to claim 2, characterized in that Step 1.3 specifically includes: 1.31) Convert the strategy and value functions into vectors, encapsulate them into data packets, and extend the SignalHandler class; First, the low-level agent uses the on_message method to parse the received message and extract the policy and value function. Then, it uses the convert_to_vector function to convert the policy and value function into vector form. The converted vector data is encapsulated into a data packet. Finally, it uses json.dumps to convert the data packet into JSON format and sends it to the signaler using self.write_message. 1.32) The signal machine receives JSON data, including policy and value function weight parameters. The signal machine parses the JSON formatted policy and value function. It then extracts the policy weights and value function weights, constructs a policy and value function network, and uses NumPy arrays to represent the weights. Finally, it uses a multilayer perceptron to construct a policy network using the policy weights and a value function network using the value function weights.

Citation Information

Patent Citations

  • Distributed traffic signal control method based on generative adversarial network and reinforcement learning

    CN113436443A

  • Multi-ramp cooperative control method based on model predication control

    CN108898854A

  • Regional traffic signal coordination optimization control system and method

    CN109785619A