Urban traffic signal coordinated control method based on multi-agent deep reinforcement learning

Through the multi-agent deep reinforcement learning method, a coordinated urban traffic signal control system was constructed, which solved the state space explosion and agent coupling problems of multi-intersection traffic signal control, realized efficient coordinated control of the urban road network, reduced congestion delays, and improved traffic efficiency.

CN114995119BActive Publication Date: 2025-09-09NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210151210.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-16
Publication Date
2025-09-09
Estimated Expiration
2042-02-16

AI Technical Summary

Technical Problem

Existing urban traffic control systems suffer from the problem of state space explosion when processing high-dimensional traffic data. In addition, the degree of coupling between agents in multi-agent reinforcement learning is high, making training difficult to converge and making it impossible to effectively coordinate the control of traffic signals at multiple intersections.

Method used

A multi-agent deep reinforcement learning method is used to construct a traffic control unit module, a road network collaborative control module, a traffic information collection module and a playback memory pool. The collaborative control module guides the real-time response signal control of multiple intersections, reducing the congestion delay time of intersections and balancing the utilization of the road network.

Benefits of technology

It achieves overall coordinated control of multiple intersections, reduces the risk of state space explosion, improves computing efficiency and control system flexibility, reduces traffic congestion delays, and improves the traffic efficiency of urban road networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114995119B_ABST
    Figure CN114995119B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for coordinated control of urban traffic signals based on multi-agent deep reinforcement learning, comprising: collecting traffic status information vectors of the urban road network; coordinating the control strategies of the intersections of each sub-region, and generating control strategies for the intersections of the sub-regions. Traffic signal timing is optimized through a deep reinforcement learning algorithm, and traffic flow at intersections is dynamically adjusted in real time to reduce congestion and delays. The signal timing of all intersections is optimized with the goal of reducing the total travel time, preventing the optimization of a single intersection from having an adverse effect on the road network, and continuously updating the optimization strategy through reinforcement learning. The present invention can meet the complexity, real-time, and adaptability requirements of urban traffic signal control problems, improve the overall traffic efficiency of the urban road network, and alleviate traffic congestion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for coordinated control of urban traffic signals based on multi-agent deep reinforcement learning, and belongs to the technical field of intelligent traffic control. Background Art

[0002] With the continuous growth of my country's urban population and vehicle fleet, urban traffic often experiences periodic, prolonged, and widespread congestion during peak hours. my country's urban transportation system is characterized by high vehicle volumes, uneven temporal and spatial distribution of vehicles, and significant influence from intersection signal control. Given the limited availability of urban land resources, simply increasing transportation infrastructure cannot solve this problem, necessitating the development of advanced urban traffic control systems. Existing adaptive traffic control systems, such as SCOOT and SCAT, require complex mathematical models, and their effectiveness depends on the accuracy of these models. Furthermore, the higher the model's accuracy, the more complex and time-consuming the structure and parameter adjustments become. This creates a conflict between real-time performance and reliability, a conflict that becomes particularly pronounced when improving control efficiency. Furthermore, existing control systems rely on manual parameter adjustment experience, resulting in signal timing schemes that often lag behind, sometimes exacerbating traffic congestion.

[0003] Reinforcement learning, a key branch of machine learning, eliminates the need for precise modeling of the traffic environment. Instead, it continuously interacts with the environment to obtain feedback on the effects of different signal control strategies. It then learns control strategies for different random traffic environments, ultimately achieving the optimal signal control strategy for dynamic traffic conditions. Furthermore, offline reinforcement learning technology can separate training and control. Pre-training neural networks using empirical and simulation data enables the system to possess preliminary adaptive control capabilities before use. During actual use, the latest traffic data is continuously collected to optimize the control neural network, achieving dynamic control based on the ever-changing random traffic flow.

[0004] Existing reinforcement learning control systems are primarily implemented for single intersections. A mature solution for collaborative control across multiple intersections is lacking. This is because the time complexity of reinforcement learning algorithms increases exponentially with the size of the state space, making single-intersection control solutions inapplicable to the coordinated control of multiple intersections. Furthermore, existing multi-agent reinforcement learning collaborative control theory is complex, and its practicality needs further improvement. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the shortcomings of the existing technology and provide a method for coordinated control of urban traffic signals based on multi-agent deep reinforcement learning. By constructing traffic control unit modules (agents), road network coordinated control modules (agents), traffic information collection modules, and playback memory pools at multiple intersections, real-time response signal control of multiple intersections is achieved, and the agents can collaborate under the guidance of the coordinated control module. This system can effectively reduce congestion delays at intersections, balance the utilization of each intersection in the road network, and reduce traffic delays. This system solves the state space explosion problem of existing reinforcement learning technology when processing high-dimensional traffic data, as well as the problem of high coupling between agents in multi-agent reinforcement learning, which makes training difficult to converge.

[0006] To achieve the above objectives, the present invention provides a method for coordinated urban traffic signal control based on multi-agent deep reinforcement learning, which is characterized by comprising:

[0007] Collect traffic status information vectors of urban road networks;

[0008] Coordinate the control strategies of each sub-area intersection and generate the control strategies of the sub-area intersections.

[0009] Prioritize the collection of traffic status information vectors of the urban road network, including:

[0010] The urban road network is divided into N sub-areas containing traffic lights, and the traffic status information vector Traffic status information for all sub-areas The set of , i∈[1,N], is the total number of sub-areas in the urban road network;

[0011] Traffic status information of sub-area i Including time Sub-area Congestion delay vector and time The state vector of the traffic light in sub-region i , For the moment Sub-area i section Congestion delay value, k∈[1,K], K is the number of road sections in the sub-area;

[0012] If the moment Sub-area i section If there is no car inside, ,otherwise = , For the moment road section The total number of vehicles, For road sections The actual travel time of the vehicle, For road sections travel time at free-flow speed;

[0013] time The state vector of the traffic light in sub-region i , For the moment Phase timing of traffic lights in sub-area i, For the moment Sub-area i phase timing The execution time of .

[0014] Prioritize and coordinate the control strategies of each sub-region intersection and generate the control strategies of the sub-region intersection, including:

[0015] Get traffic status information vector , dynamically generate the control strategy for each sub-area intersection, and convert the control strategy for each sub-area intersection into the phase timing of the traffic lights in the corresponding sub-area.

[0016] Prioritize, obtain traffic status information vector and , dynamically generate the control strategy for each sub-region intersection, and convert the control strategy for each sub-region intersection into the phase timing of the traffic lights in the corresponding sub-region, including:

[0017] Traffic status information of sub-area i , sent to the execution network obtained by training ;

[0018] Execution Network Dynamically generate known optimal control strategies ;

[0019] The optimal control strategy Converted into the phase timing of the traffic light in the corresponding sub-area i :

[0020] Get the current traffic light phase timing Phase Timing Execution time , phase timing Latest execution time ;like , the traffic light will jump to the next phase timing ;

[0021] from Extract the congestion delay status of the sub-area , calculate the reward function ;

[0022] Get the traffic information status at the next moment and ,Will Save as experience data.

[0023] Prioritize the local evaluation network based on training , train the execution network ,include:

[0024] Train the local evaluation network ,include:

[0025] collection , obtain historical experience data , is the historical moment of sub-region i Traffic status information, for The corresponding historical control strategy, For control strategy The reward value, For the moment +1 Traffic status information of the sub-area, For the moment +1 Traffic status information of the entire road network, is the capacity of historical experience data;

[0026] Random selection The historical experience data constitutes the training data set ;

[0027] Use the training dataset to evaluate the local network Conduct training, including:

[0028] Extract reward vector from training dataset , traffic status information vector and ;

[0029] The trained global execution network Calculate the global optimization strategy Control strategy component ;

[0030] Update the target evaluation network using soft update method Weight :

[0031] ,

[0032] Where τ is the set coefficient, is the weight before update, is the updated target evaluation network Weight , for The weight of

[0033] According to the traffic status information vector , control strategy component , reward vector and target evaluation network Weight , solve the control target vector that maximizes the cumulative reward :

[0034] ,

[0035] Where, is the set discount factor;

[0036] Computing local evaluation network Value and the control target vector The loss value between:

[0037] ,

[0038] Where, yes and the control target vector The loss value between It is a local review network The weight vector is updated iteratively using the Adam optimizer with the goal of minimizing the loss value ; is the expected value of loss calculated from the training data set, , ;

[0039] judge Can it converge to ,like Converges to Then the final local evaluation network is output .

[0040] Prioritize the local evaluation network based on training , train the execution network ,include:

[0041] Step 1: From the training dataset , from which the traffic information state vector is extracted ;

[0042] Step 2: Call the execution network Calculated Corresponding strategies , ;

[0043] Step 3, and Substitute into the local evaluation network , calculation strategy Score ;

[0044] Step 4: Score Substitute the following equation and use the Adam optimizer and deterministic policy gradient method to update the execution network Weight , so that The highest score is achieved by:

[0045] ,

[0046] Where, is the learning rate, Is the execution network The network weight The policy gradient of Is the execution network Strategy Vector The policy gradient of The status is When the strategy is The conditional probability of yes In the The update amount of the step;

[0047] Step 5, when Stop updating when is the similarity threshold, and the final execution network is obtained by output . Prioritize the global execution network obtained by training Calculate the global optimization strategy Control strategy component ,include:

[0048] Global Execution Network Get the traffic status information vector of the urban road network ; Based on traffic status information vector , global execution network Calculate the global optimization strategy , , is the total number of sub-areas in the urban road network, Decompose into , Global optimization strategy control strategy component.

[0049] Prioritize the global evaluation network obtained through training Train to obtain the global execution network ,include:

[0050] Train to obtain a global evaluation network ,include:

[0051] Step 1: Get the current time of the city road network The control strategies of all N sub-regions are combined into a global control strategy ; Get the global traffic status information vector at the current moment and the global traffic state information vector at the next moment ;

[0052] Step 2: Calculate the global reward value based on the total congestion delay time of the urban road network , , obtain global experience data ;

[0053] Step 3, by continuously collecting Obtaining experience data for the entire road network , D is the capacity, For urban road network Historical traffic status information at all times, for The control strategy of all sub-areas at the moment, For control strategy The reward value, For urban road network Historical traffic status information at the next moment;

[0054] Step 4, random selection Group data constitutes the training set , extract the reward value from the training set Constructing the reward vector , extract traffic status information from the training set Constructing traffic status information vector , and according to Generate global control strategy ,Right now = ;

[0055] Step 5: Use soft update method to update the global target evaluation network Weight :

[0056] ;

[0057] Where τ is the set coefficient, is the global target evaluation network before updating The weight of is the updated global target evaluation network Weight , yes The weight of

[0058] Update the global control objective function to maximize the global cumulative reward;

[0059] The global control objective function is:

[0060] ,

[0061] Where, is the global control target, Evaluate the network for the global goal, It is a global target evaluation network The weight of

[0062] Step 6: Update the global evaluation network through iteration The weight 𝒘 is minimized ; The calculation formula is:

[0063] ,

[0064] Where, Based on the training set The loss value obtained is The expected loss value calculated for the training set; call the Adam optimizer to iteratively update the global evaluation network with the goal of minimizing the loss value The weight vector 𝒘 of

[0065] If Converges to , then the operation ends and the final global evaluation network is obtained .

[0066] Prioritize the global evaluation network obtained through training Train to obtain the global execution network ,include:

[0067] Step 1: Get the training set , from which the traffic information state vector is extracted ;

[0068] Step 2: Call the global execution network Calculated Corresponding strategies , ;

[0069] Step 3, and Substitute into the global evaluation network , calculation strategy Score ;

[0070] Step 4: Score Substitute the following equation and use the Adam optimizer to perform the network globally Weight Update to improve The score reached the highest;

[0071] Global Execution Network Weight The update formula is:

[0072]

[0073] Where, is the learning rate, is the global execution network Weight The policy gradient of is the control strategy vector The policy gradient of The status is When the strategy is The conditional probability of yes In the The update increment of the step;

[0074] like Then the operation ends. is the similarity threshold, outputting the global execution network Otherwise, go to step 1.

[0075] The beneficial effects achieved by the present invention are:

[0076] (1) The system and method described in this invention provide for the coordinated control of an urban road network as a whole, no longer limited to point-by-point control at a single intersection. By integrating the phase control of traffic lights at multiple intersections through a multi-agent deep reinforcement learning approach, the system can improve the efficiency of entire bottleneck sections and reduce urban traffic congestion. Furthermore, through collaboration, the input state size of each traffic control unit module can be reduced, avoiding state space explosion.

[0077] (2) This invention proposes a collaborative control system based on a "global-local" hierarchical control framework, which decomposes the collaborative control problem of traffic lights at multiple intersections in a city into a series of local optimization problems of limited scale, which not only reduces the computational burden but also ensures the convergence of the algorithm. At the global level, a loosely coupled design is adopted, allowing the collaborative control module to provide strategy guidance and optimization during the training of the traffic control unit module, without directly participating in the control of the intersection through methods such as shared weights. This reduces the coupling between the collaborative control module and the traffic control unit module, thereby reducing the difficulty of training. At the local level, autonomous control is adopted, allowing the intelligent agent to balance local and global optimization strategies according to actual conditions, thereby increasing the flexibility of control and targeting specific problems.

[0078] (3) This invention reduces the number of global evaluation networks in the control system and simplifies the algorithm structure. In contrast, the newer foreign Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm requires a global evaluation network for each agent to monitor changes in global state information. However, this invention only requires one global evaluation network in the collaborative control module. By enabling other traffic control unit modules to share the results of the global evaluation network through the communication module, collaborative effects can be achieved, improving computational efficiency.

[0079] (4) The present invention helps to improve the level of intelligent management and control of urban road traffic in my country. It can solve the problem that traditional single-point traffic control is difficult to handle congestion and delays at the road network level, and has good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 It is a schematic diagram of the overall structure of the present invention;

[0081] Figure 2 A schematic diagram of a typical urban road network division and corresponding traffic status information of the present invention;

[0082] Figure 3 A schematic diagram of typical intersection signal phase timing according to the present invention;

[0083] Figure 4 A training flow chart of the traffic control unit module of the present invention;

[0084] Figure 5 is an execution flow chart of the collaborative control module of the present invention;

[0085] Figure 6 This is a training flow chart of the collaborative control module of the present invention. DETAILED DESCRIPTION

[0086] The following examples are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0087] A multi-agent deep reinforcement learning urban traffic signal collaborative control system includes a traffic information collection module, a traffic control unit module, a collaborative control module and a communication module.

[0088] The traffic information collection module is used to collect urban road traffic information;

[0089] The traffic control unit module is used to dynamically adjust the control strategy of the intersection traffic flow to reduce congestion and delays;

[0090] The collaborative control module is used to coordinate the control strategies of various intersections in the urban road network to prevent the traffic flow control at a certain intersection from affecting the traffic flow at other intersections, thereby reducing the total travel time of the road network traffic;

[0091] The communication module is used for communication and interaction between the traffic control unit module and the collaborative control module.

[0092] Furthermore, the traffic information collection module in this embodiment uses vehicle-road cooperative technology, or obtains traffic status information available for deep reinforcement learning through devices such as coils or high-definition cameras. , and then sent to the traffic control unit module and the collaborative control module. Traffic status information The sampling interval is 5 seconds; after each control strategy update, the latest control strategy of each traffic control unit module is recorded and sent to the local replay memory pool and the global replay memory pool.

[0093] Traffic status information Congestion delay vectors for each road segment in the urban road network , and the state vector of the traffic light ,Right now ; The latest control strategy includes the updated phase and phase duration .

[0094] Get traffic status information vector The method is:

[0095] by The urban road network is gridded with a spacing of meters, and a total of grid, forming A vector of dimension, is the total number of grids, and each intersection establishes a separate grid; at the same time, the road network is divided into sub-areas, each of which contains a traffic light, an intersection, and the Road grid. Traffic status information vector Traffic status information for all sub-areas i A collection of The value is 200 meters;

[0096] Traffic status information of sub-area i Including time Sub-area Congestion delay vector and time The state vector of the traffic light in sub-region i , For the moment Sub-area i section Congestion delay value, i∈[1,N], k∈[1,K]; if time Sub-area i section If there is no car inside, ,otherwise = , For the moment road section The total number of vehicles, For road sections The actual travel time of the vehicle, For road sections travel time at free-flow speed;

[0097] time The state vector of the traffic light in sub-region i , For the moment Phase timing of traffic lights in sub-area i, For the moment Sub-area i phase timing The execution time of .

[0098] like Figure 3 As shown, phase timing is divided into P1, P2, P3 and P4. In each phase timing, the green light duration, yellow light duration and red light duration of the four traffic lights in the same intersection are executed in sequence. For example, in phase P1, vehicles in the north-south direction can go straight and turn right, in phase P2, vehicles in the east-west direction can turn left, in phase P3, vehicles in the east-west direction can go straight and turn right, and in phase P4, vehicles in the east-west direction can turn left.

[0099] Traffic information collection module transmits traffic status information vectors through the communication module Sent to the traffic control unit module and the collaborative control module respectively.

[0100] The traffic signal control module obtains the traffic status information vector , dynamically generate the control strategy for each sub-region intersection, and convert the control strategy for each sub-region intersection into the phase timing of the traffic lights in the corresponding sub-region to guide vehicle driving behavior and implement traffic control at the intersection. The main steps are:

[0101] Step 1: Obtain traffic status information from the traffic information collection module , and extract the traffic status information of sub-area i , sent to the local execution network ;

[0102] Step 2: Execute the network Dynamically generate known optimal control strategies ;

[0103] Step 3: The optimal control strategy Converted into the phase timing of the traffic lights in the corresponding sub-area ;

[0104] Get the current traffic light phase timing Phase Timing Execution time , latest duration .like , the traffic light will jump to the next phase timing , and repeat over and over again;

[0105] Step 4: Update the latest control strategy A copy of the data is sent to the collaborative control module to generate global decision information.

[0106] The local playback memory pool is used to store the intersection Historical experience data , the capacity is set to .in is the historical moment of sub-region i Traffic status information, for The corresponding historical control strategy, For control strategy The reward value, For the moment +1 Traffic status information of the sub-area, For the moment +1 traffic status information. In each sampling period, the traffic information collection module puts the latest traffic status information into the local playback memory pool. When the memory pool is full, the earliest stored traffic status information is deleted. Historical experience data.

[0107] The strategy optimization module is used to optimize the intersection Control strategy The policy optimization module includes a local evaluation (neural) network and target evaluation (neural) networks Local review network Used to evaluate the short-term behavior of the execution network and the target evaluation network Used to evaluate the long-term behavior of the execution network.

[0108] The network structures of the execution network, local evaluation network, target evaluation network, global execution network, global evaluation network and global target evaluation network in the present invention are all prior arts and will not be elaborated in detail in this embodiment.

[0109] The steps to train the local evaluation network are:

[0110] Step 1: Randomly select from the local playback memory pool The historical experience data constitutes the training data set , extract the reward vector from the training dataset , traffic status information vector ,and ;

[0111] Step 2: The strategy optimization module requests the collaborative control module to provide a global optimization strategy, including:

[0112] Step 2-1: Transform the traffic status information vector Send to collaborative control module;

[0113] Step 2-2, obtain the global optimization strategy component generated by the collaborative control module ;

[0114] Step 3, the strategy optimization module is based on , , and target evaluation network Weight Find the control target vector that maximizes the cumulative reward :

[0115]

[0116] Where, =0.9 is the discount factor, which reflects that the influence of historical experience data decreases over time.

[0117] Step 4: Make the local evaluation network of Value and control target vector The loss value between them is the smallest:

[0118]

[0119] Where, is the loss function, representing the local evaluation network Control target vector The loss value between It is a local review network The weight vector of is the expected loss value of the batch of samples, is the historical traffic status information vector, yes Corresponding to the historical control strategy vector.

[0120] Iteratively update the local evaluation network through the Adam optimizer The weight vector , to minimize the loss value, so that Gradually converge to , then the optimal local evaluation network is obtained and target evaluation network, ending the run.

[0121] Update the target evaluation network using soft update method Weight :

[0122]

[0123] Where τ is the set coefficient, which is set to 0.9. is the weight of the target evaluation network before updating, is the weight of the updated target evaluation network, for The weight of .

[0124] Step 5: Based on the local evaluation network, train the local execution network:

[0125] Step 5-1, obtain the training data set , from which the traffic information state vector is extracted ;

[0126] Step 5-2, call the execution network Calculated Corresponding strategies , ;

[0127] Step 5-3, and Substitute into the evaluation network , calculation strategy Score ;

[0128] Step 5-4, score Substitute the following equation and use the Adam optimizer and deterministic policy gradient method to update the execution network Weight ,make Score as high as possible:

[0129] ,

[0130] Where, is the learning rate, Is the execution network The network weight The policy gradient of Is the execution network Strategy Vector The policy gradient of The status is When the strategy is The conditional probability of yes In the The update amount of the step;

[0131] Step 5-5, when Stop updating when is the similarity threshold, and the final execution network is obtained by output .

[0132] The collaborative control module is used to generate a global optimization strategy, and includes a global strategy execution module, a global memory pool and a global strategy optimization module.

[0133] The global memory pool is used to store the experience data of the entire road network , the capacity is .in, For the road network Historical traffic status information at all times, For all traffic control units corresponding to The historical control strategy of the moment, For control strategy The reward value, For the road network The historical traffic status information at each moment. In each sampling period, the traffic information collection module puts the latest experience data into the global playback memory pool, and the reward value Calculated according to the control target. When the memory pool is full, delete the earliest pieces of data.

[0134] The global strategy execution module receives the request from the traffic control unit module to obtain the global optimization strategy, and obtains the traffic status information vector of the urban road network from the traffic information collection module. , and then through the global execution network Dynamically generate a global optimization strategy, and then send the global optimization strategy to the local evaluation network of each traffic control unit module In this way, the traffic control unit module can take into account the optimization goals of both local intersections and the global road network.

[0135] Get the global optimization strategy Control strategy component The specific steps are:

[0136] Step 1: Global execution network Get traffic status information vector ;

[0137] Step 2: Global execution network Compute global optimization strategy ,in , The number of sub-areas in which the urban road network is divided (each sub-area has only one traffic control unit). Decompose into , and then Return to the traffic control unit module of sub-area i.

[0138] The global strategy optimization module includes a global evaluation network and global target evaluation network . Global strategy optimization module training obtained and update Optimize and then pass Execute the network globally The global strategy optimization module is trained to obtain and update The main steps are:

[0139] Step 1: Get the current time of the city road network The control strategies of all N sub-regions are combined into a global control strategy ; Get the global traffic status information vector at the current moment and the global traffic state information vector at the next moment ;

[0140] Step 2: Calculate the global reward value based on the total congestion delay time of the urban road network , , obtain global experience data ;

[0141] Step 3, by continuously collecting Obtaining experience data for the entire road network , D is the capacity, For urban road network Historical traffic status information at all times, for The control strategy of all sub-areas at the moment, For control strategy The reward value, For urban road network Historical traffic status information at all times;

[0142] Step 4, random selection Group data constitutes the training set , extract the reward value from the training set Constructing the reward vector , extract traffic status information from the training set Constructing traffic status information vector , and according to Generate global control strategy ,Right now = ;

[0143] Step 5: Update the global control objective function to maximize the global cumulative reward. The global control objective function is:

[0144]

[0145] Where, is the global control target, Evaluate the network for the global goal, It is a global target evaluation network The weight of

[0146] Step 6: Update the global evaluation network through iteration The weight 𝒘 is minimized , The calculation formula is:

[0147]

[0148] Where, Based on the training set The loss value obtained is The expected value of loss calculated for this batch of training samples;

[0149] If Converges to , then the operation ends and the final global evaluation network is obtained , otherwise go to step 7;

[0150] Step 7: Call the Adam optimizer to iteratively update the global evaluation network with the goal of minimizing the loss value The weight vector ;

[0151] Step 8: Use soft update method to update the global target evaluation network Weight :

[0152]

[0153] Where τ is the set coefficient, which is set to 0.9. is the weight before update, is the updated weight, yes The global strategy optimization module is Get the global execution network for training , so the main steps of optimization are:

[0154] Step 1: Get the training dataset , from which the traffic information state vector is extracted ;

[0155] Step 2: Call the global execution network Calculated Corresponding strategies , ;

[0156] Step 3, and Substitute into the evaluation network , calculation strategy Score ;

[0157] Step 4: Score Substitute the following equation and use the Adam optimizer to perform the network globally Weight Update to make The score of is as high as possible. Global execution network Weight The update formula is:

[0158]

[0159] Where, is the learning rate, is the global execution network Weight The policy gradient of is the control strategy vector The policy gradient of The status is When the strategy is The conditional probability of yes In the The update increment of the step, when Stop updating when is the similarity threshold, and the final global execution network is obtained .

[0160] The communication module uses wired (such as fiber optic communication) or wireless (such as 5G communication or Dedicated Short Range Communication (DSRC)) technology to enable the various modules of the system to exchange traffic information through the module, including: a communication submodule between the traffic information collection module and the traffic control unit module, which is used to transmit traffic status information and control strategy information; and a communication submodule between the collaborative control module and the traffic control unit module, which is used to transmit global optimization decision information.

[0161] The structure of traffic control system is as follows Figure 1 As shown, the traffic control unit module includes a (local) execution network, a local playback memory pool, and a local evaluation network. The current moment is obtained through vehicle-road cooperative technology, video, coils, etc. Road network observable traffic information data , obtain urban road traffic information through reinforcement learning.

[0162] The local evaluation network randomly selects a small batch (512) of experience data from the local replay memory pool each time to train the parameter vector of the local evaluation network. , by updating Parameters minimize the loss function , and then update the local execution network parameters through the local evaluation network The local evaluation network and execution network include input layer, fully connected layer and output layer in sequence. The fully connected layer has 128 neurons and uses ReLU activation function to evaluate the score of the control strategy output by the network ( value), execute network output phase encoding (corresponding to Figure 3 in , , , The duration of each phase ranges from 0 to 1 minute. The output layer uses the Sigmoid activation function to ensure that the output value is bounded. The value is 0.99.

[0163] Collaborative control modules such as Figure 1 As shown in , it includes a global execution network, a global memory pool, and a global evaluation network, which are used to generate a global optimization strategy and then optimize the strategy of the traffic control unit module.

[0164] The global execution network obtains the global traffic status of the urban road network ,according to Generate global optimization strategy The collaborative control module uses a request-response method to interact with the traffic control unit module. The method is to monitor the request of the traffic control unit module. When the traffic control unit module When a collaborative request is sent, the collaborative control module responds to the request and optimizes the global strategy The amount Sent to the requester.

[0165] The global evaluation network randomly selects a small batch (512 pieces) of experience data from the global replay memory pool each time to train the parameter vector of the global evaluation network. , by updating Parameters minimize the loss function , and then update the global execution network parameters through the evaluation network The evaluation network and execution network include input layer, fully connected layer and output layer respectively. The fully connected layer has 256 neurons and uses ReLU activation function. The output of the evaluation network is the score of the strategy ( The output of the execution network is the phase and duration of each signal light in the road network, and the duration of each phase ranges from 0 to 1 minute. The output layer uses the Sigmoid activation function. Learning rate The value is 0.99.

[0166] The communication module uses optical fiber, 5G communication or DSRC technology to enable each module in the system to exchange traffic status information and control decision information.

[0167] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0168] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0170] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A collaborative urban traffic signal control method based on multi-agent deep reinforcement learning, characterized by: include: Collect traffic status information vectors of urban road networks; Coordinate the control strategies of each sub-area intersection and generate the control strategies of the sub-area intersections, including: Get traffic status information vector , dynamically generate the control strategy for each sub-region intersection, and convert the control strategy for each sub-region intersection into the phase timing of the traffic lights in the corresponding sub-region, specifically including: Traffic status information of sub-area i , sent to the execution network obtained by training ; Execution Network Dynamically generate known optimal control strategies ; The optimal control strategy Converted into the phase timing of the traffic light in the corresponding sub-area i : Get the current traffic light phase timing Phase Timing Execution time , phase timing Latest execution time ;like , the traffic light will jump to the next phase timing ; from Extract the congestion delay status of the sub-area , calculate the reward function ; Get the traffic information status at the next moment and ,Will Save as experience data; Based on the local evaluation network obtained by training , train the execution network ,include: Train the local evaluation network ,include: collection , obtain historical experience data , is the historical moment of sub-region i Traffic status information, for The corresponding historical control strategy, For control strategy The reward value, For the moment +1 Traffic status information of the sub-area, For the moment +1 Traffic status information of the entire road network, is the capacity of historical experience data; Random selection The historical experience data constitutes the training data set ; Use the training dataset to evaluate the local network Conduct training, including: Extract reward vector from training dataset , traffic status information vector and ; The trained global execution network Calculate the global optimization strategy The control strategy component ; Update the target evaluation network using soft update method Weight : , Where τ is the set coefficient, is the weight before update, is the updated target evaluation network Weight , for The weight of According to the traffic status information vector , control strategy component , reward vector and target evaluation network Weight , solve the control target vector that maximizes the cumulative reward : , Where, is the set discount factor; Computing local evaluation networks Value and the control target vector The loss value between: , Where, yes and the control target vector The loss value between It is a local review network The weight vector is updated iteratively using the Adam optimizer with the goal of minimizing the loss value ; is the expected value of loss calculated from the training data set, , ; judge Can it converge to ,like Converges to Then the final local evaluation network is output .

2. The urban traffic signal coordinated control method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: Collect traffic status information vectors of urban road networks, including: The urban road network is divided into N sub-areas containing traffic lights, and the traffic status information vector Traffic status information for all sub-areas The set of , i∈[1,N], is the total number of sub-areas in the urban road network; Traffic status information of sub-area i Including time Sub-area Congestion delay vector and time The state vector of the traffic light in sub-region i , For the moment Sub-area i section Congestion delay value, k∈[1,K], K is the number of road sections in the sub-area; If the moment Sub-area i section If there is no car inside, ,otherwise = , For the moment road section The total number of vehicles, For road sections The actual travel time of the vehicle, For road sections travel time at free-flow speed; time The state vector of the traffic light in sub-region i , For the moment Phase timing of traffic lights in sub-area i, For the moment Sub-area i phase timing The execution time of .

3. The urban traffic signal coordinated control method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: Based on the local evaluation network obtained by training , train the execution network ,include: Step 1: From the training dataset , from which the traffic information state vector is extracted ; Step 2: Call the execution network Calculated Corresponding strategies , ; Step 3, and Substitute into the local evaluation network , calculation strategy Score ; Step 4: Score Substitute the following equation and use the Adam optimizer and deterministic policy gradient method to update the execution network Weight , so that The highest score is achieved by: , Where, is the learning rate, Is the execution network The network weight The policy gradient of Is the execution network Strategy Vector The policy gradient of The status is When the strategy is The conditional probability of yes In the The update amount of the step; Step 5, when Stop updating when is the similarity threshold, and the final execution network is obtained by output .

4. The urban traffic signal coordinated control method based on multi-agent deep reinforcement learning according to claim 3 is characterized in that: The trained global execution network Calculate the global optimization strategy The control strategy component ,include: Execute the network globally Get the traffic status information vector of the urban road network ; Based on traffic status information vector , global execution network Calculate the global optimization strategy , , is the total number of sub-areas in the urban road network, Decompose into , Global optimization strategy control strategy component.

5. The urban traffic signal coordinated control method based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: Based on the global evaluation network obtained by training Train to obtain the global execution network ,include: Train to obtain a global evaluation network ,include: Step 1: Get the current time of the city road network The control strategies of all N sub-regions are combined into a global control strategy ; Get the global traffic status information vector at the current moment and the global traffic state information vector at the next moment ; Step 2: Calculate the global reward value based on the total congestion delay time of the urban road network , , obtain global experience data ; Step 3, by continuously collecting Obtaining experience data for the entire road network , D is the capacity, For urban road network Historical traffic status information at all times, for The control strategy of all sub-areas at the moment, For control strategy The reward value, For urban road network Historical traffic status information at the next moment; Step 4, random selection Group data constitutes the training set , extract the reward value from the training set Constructing the reward vector , extract traffic status information from the training set Constructing traffic status information vector , and according to Generate global control strategy ,Right now = ; Step 5: Use soft update method to update the global target evaluation network Weight : ; Where τ is the set coefficient, is the global target evaluation network before updating The weight of is the updated global target evaluation network Weight , yes The weight of Update the global control objective function to maximize the global cumulative reward; The global control objective function is: , Where, is the global control target, Evaluate the network for the global goal, It is a global target evaluation network The weight of Step 6: Update the global evaluation network through iteration The weight 𝒘 is minimized ; The calculation formula is: , Where, Based on the training set The loss value obtained is The expected loss value calculated for the training set; call the Adam optimizer to iteratively update the global evaluation network with the goal of minimizing the loss value The weight vector 𝒘 of If Converges to , then the operation ends and the final global evaluation network is obtained .

6. The urban traffic signal coordinated control method based on multi-agent deep reinforcement learning according to claim 5 is characterized in that: Based on the global evaluation network obtained by training Train to obtain the global execution network ,include: Step 1: Get the training set , from which the traffic information state vector is extracted ; Step 2: Call the global execution network Calculated Corresponding strategies , ; Step 3, and Substitute into the global evaluation network , calculation strategy Score ; Step 4: Score Substitute the following equation and use the Adam optimizer to perform the network globally Weight Update by adjusting make The score reaches the highest; the global execution network Weight The update formula is: , Where, is the learning rate, is the global execution network Weight The policy gradient of is the control strategy vector The policy gradient of The status is When the strategy is The conditional probability of yes In the The update increment of the step; like Then the operation ends. is the similarity threshold, outputting the global execution network Otherwise, go to step 1.

Citation Information

Patent Citations

  • Multi-intersection traffic light control method and system based on reinforcement learning, and storage medium

    CN113223305A