Multi-intersection asynchronous traffic signal control method and system in communication-free environment

By using the partially observable Markov decision-making process and Mac-I2Q algorithm based on macro action in a communication-free environment, the traffic signal phase and duration are optimized, and the inefficiency problem of traditional traffic light control methods in dynamic traffic environments is solved, and more efficient traffic management is achieved.

CN120544408APending Publication Date: 2025-08-26SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410193516.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-21
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Traditional traffic light control methods are difficult to adapt to highly dynamically changing urban traffic environments, resulting in inefficient traffic management during periods of fluctuations in traffic flows, and existing methods pose security risks in communication-free environments.

Method used

Using partially observable Markov decision-making process and Mac-I2Q algorithm based on macro actions, a multi-intersection asynchronous traffic signal control system is designed in a communication-free environment, and through distributed reinforcement learning agents to train in a simulation environment, optimize the phase and duration of traffic signals to improve traffic efficiency.

Benefits of technology

It shortens the average driving time of the vehicle, improves vehicle throughput, provides more stable traffic control signals, and improves the efficiency of transportation management and the stability of traffic order.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544408A_ABST
    Figure CN120544408A_ABST
Patent Text Reader

Abstract

The invention provides a multi-intersection asynchronous traffic signal control method and system in a communication-free environment, and the method comprises the steps: S1, designing a traffic signal control system, and controlling the traffic signal and duration of an intersection; s2, modeling a multi-intersection traffic signal control problem into a partially observable Markov decision process based on a macro action, and designing a state, an action, an award and a target of an intelligent agent; s3, proposing an independent ideal Q algorithm based on macro action, and realizing phase and duration control of traffic signals; and S4, deploying the trained Mac-I2Q algorithm into an intersection agent, training in a simulated traffic environment, interacting the agent with the environment, storing and transferring data, updating network parameters, and repeating the process until the whole simulation process is finished. According to the invention, more stable traffic control signals can be provided for a traffic signal lamp system, and the efficiency of traffic transportation management can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of reinforcement learning and traffic signal control, and in particular to a method and system for controlling asynchronous traffic signals at multiple intersections in a non-communication environment. Background Art

[0002] Traffic congestion has become a serious challenge in urban management. It not only restricts urban traffic efficiency and affects residents' travel experience, but also has a negative impact on economic development and the ecological environment. Traffic congestion can also cause a series of environmental problems such as air pollution, noise pollution and urban heat island effect.

[0003] Traffic lights, the "metronome" of urban traffic, are the primary means of managing urban traffic. Their proper control is crucial to alleviating traffic congestion. However, traditional traffic light control methods often rely on fixed-duration rules, making them difficult to adapt to the highly dynamic urban traffic environment. Most existing traffic signal control systems rely on pre-designed signal rules, which often fail to effectively direct traffic in practice, especially during periods of high traffic flow fluctuations.

[0004] In recent years, with the rapid development of machine learning, especially deep reinforcement learning, many researchers have begun applying this technology to traffic signal control, aiming to achieve more efficient dynamic signal regulation. Deep reinforcement learning is particularly well-suited to solving dynamic control problems in complex environments, as reinforcement learning agents continuously learn and optimize strategies through trial and error in their interaction with the environment, gradually adapting to changes in traffic flow. For example, deep reinforcement learning technology has achieved excellent results in robot control and drone collaborative control tasks.

[0005] Patent document CN113643528A discloses a signal light control method for controlling the operating state of a signal light at a target intersection. The method comprises: obtaining traffic state information at the target intersection, the traffic state information including driving state information of a first vehicle at the target intersection; predicting a signal light state strategy based on the traffic state information to obtain a signal light control strategy for the target intersection; and controlling the operating state of the signal light at the target intersection based on the signal light control strategy. The first vehicle driving state information comprises driving state statistical characteristics of vehicles traveling in different directions on each fork road at the target intersection. The first vehicle driving state information is obtained by: obtaining driving state characteristics of the vehicles traveling on each fork road at the target intersection; and grouping and statistically analyzing the driving state characteristics of the vehicles according to their driving directions on the fork roads to obtain the driving state statistical characteristics of vehicles traveling in different directions on each fork road at the target intersection. Although existing methods can extend signal duration by continuously selecting the same signal action, this may pose safety risks in practical applications. Summary of the Invention

[0006] In view of the defects in the prior art, the object of the present invention is to provide a method and system for controlling asynchronous traffic signals at multiple intersections in a non-communication environment.

[0007] The method for controlling asynchronous traffic signals at multiple intersections in a non-communication environment provided by the present invention includes:

[0008] Step S1: Design a traffic signal control system to control the traffic signals and duration at the intersection, and design system optimization objectives, including average vehicle travel time and throughput;

[0009] Step S2: Model the multi-intersection traffic signal control problem as a partially observable Markov decision process based on macro-actions, and design the state, action, reward, and goal of the agent;

[0010] Step S3: Implementing phase and duration control of traffic signals based on an independent ideal Q algorithm based on macro actions;

[0011] In step S4, the Mac-I2Q algorithm is deployed to the intersection agent and trained in a simulated traffic environment. The agent interacts with the environment, stores and transfers data, and updates network parameters. This process is repeated until the entire simulation process is completed.

[0012] Preferably, the step S1 includes:

[0013] Step S1.1: The vehicle in the system enters the road network. The road network is the environment in which motor vehicles move, which includes N intersections and roads connecting the intersections. The road network is described as a directed graph. in represents an intersection, and ij represents the road leading to intersection j; ji represents the road leading to intersection i from intersection j;

[0014] Step S1.2: A vehicle enters an intersection in the system. For an intersection i, the exit road is the road leading from intersection i to a certain intersection, i.e., ij∈ε. Conversely, ij is the entrance road to intersection j. An intersection is connected to 8 roads, 4 of which are entrance roads and 4 are exit roads.

[0015] Step S1.3: When a vehicle enters an intersection, it selects a lane based on its direction of travel. A road has three lanes, corresponding to the three possible directions of travel for vehicles at the intersection: left turn, straight ahead, and right turn. A lane specifies the direction of travel for vehicles within it. Assume that vehicles in each lane do not change lanes within the lane. The lane entering a road is called the entry lane, and the lane exiting a road is called the exit lane. The eight roads at an intersection contain 24 lanes.

[0016] Step S1.4: A vehicle enters an intersection from an entry lane and then leaves from an exit lane. This behavior is called a traffic movement. Traffic movements are classified into left turns, straight ahead, and right turns based on the direction of the vehicle passing through the intersection. Each lane of a road corresponds to a traffic movement.

[0017] Step S1.5: The intersection decides the next signal phase when switching signals. A signal phase is a combination of traffic signals that lasts for a period of time. There are four signal phases available at an intersection: east-west straight ahead, north-south straight ahead, east-west left turn, and north-south left turn. When the intersection makes a traffic signal decision, it makes the decision among the available signal phases.

[0018] Step S1.6: The system records the driving trajectory information of each vehicle entering the road network and calculates the average vehicle driving time and throughput. The average vehicle driving time is the average length of the driving time of all vehicles. The throughput refers to the number of vehicles that complete the driving process in the road network per unit time. The driving time of each vehicle is calculated from the starting position of the vehicle to the time it takes to reach the destination. When a vehicle enters an intersection, its movement is affected by traffic signals. When it encounters a red light, it waits in the intersection until the corresponding green light comes on and then continues to drive.

[0019] Preferably, step S2 includes:

[0020] Step S2.1: Model the multi-agent system as a distributed partially observable Markov decision process based on macro-actions, written as a tuple in It's a picture Every intelligent agent, is the communication link, and also the road between the agents at adjacent intersections. The number of agents is is the global state space, is the global action space, each agent can only observe the local information o of the global state s i , S i is the local state space of agent i, and The joint observation space of the agent is Accordingly, is the local action space of agent i, and is a collection of joint macro actions, is a finite set of macro-actions for each agent i; is the joint macro-measurement set, is the macroscopic measurement of agent i, setting

[0021] Step S2.2: At each time step t, the environment state starts to transfer from s, and the system takes joint action a and transfers to the new state s ′ , the transfer process is based on the state transfer function P(s ′ |s,a), O(o,a,s ′ )=P(o|a,s ′ ) means taking joint action a and then transferring to s ′ The probability of obtaining a global observation o when R:S×A→R is the reward function. After taking a in s, the immediate reward is returned. Since the agent actually only observes local state information, the strategy π of each agent i i is the local observation information o i Mapping to actions;

[0022] Step S2.3: The partial state of the system environment observed at each time step t is received as the state information s of the reinforcement learning agent i i In the lower-level strategy, this includes the queue length of vehicles on each lane of intersection i and the current signal phase of the intersection in Is entering the lane i The number of vehicles on phase i is the signal phase of intersection i; in the upper strategy, agent i will be in state s at the end of the macro action i As its macro information i ;

[0023] Step S2.4: In the upper-level strategy, agent i takes each macro action mi At the end, the next macro action needs to be decided. The macro action contains the signal phase and duration information. The optional signal duration includes {10s, 15s, 20s, 25s}, so there are 16 optional macro actions in total. The lower-level strategy of agent i executes action a according to the macro action information. i , that is, the signal phase specified by the macro action, at each time step t, the corresponding signal phase is executed, and each action a i Lasts for 5 seconds, and determines whether the macro action is completed based on the signal duration information;

[0024] Step S2.5: At each time step t, agent i takes macro action m = < τ m ,I m ,π m >, where τ is the time required for the joint macro action m to end, initialization set Depends on the macro action observation history information of agent i denotes the lower-level strategy used to implement macro-action m. Considering the termination condition of the macro-action, the transition probability is redefined as P(s′,τ|s,m), where τ is the time required for the joint macro-action m to end. The end of the joint macro-action means that each agent has completed its own macro-action; Z(z,m,s′)=P(z|m,s′) denotes the likelihood model of the joint macro-action.

[0025] Step S2.6: The goal of the agent is to find a joint upper-level strategy Φ = × i Φ i , make decisions at the macro-action level so that the value of Φ starting from s0 is optimized, and the corresponding action value function is denote the expected cumulative rewards at a specific action and state, respectively;

[0026] Step S2.7: In the lower policy, the reward r of agent i i Defined as the immediate feedback of the current action, that is, the negative value of the average number of vehicles entering the lane in Is entering the lane i The number of vehicles on the upper layer; in the upper layer strategy, agent i takes macro action m i r at the end i Considered as the reward information for this macro action.

[0027] Preferably, step S3 includes:

[0028] Step S3.1: Use the macro-action concurrent experience replay trajectory to filter the macro-action experience data. The agent saves its action, observation, and reward information in the experience replay pool at each time step t. When sampling experience data, only the experience data at the end of the macro action is selected as valid data, and the rest of the data is discarded.

[0029] Step S3.2: In a macro action observation history h i In the example, each agent independently selects a macro action m i , save reward information, At the end of the macro action, the agent obtains a new macro measurement s′ and a new macro action observation history information h′ = <h m ,m i ,s′>; Accordingly, the experience tuple collected by agent i is expressed as where s i It is used to select macro action m i Macro information;

[0030] Step S3.3: Update the QSS network. In each training, Update in a way that minimizes the following TD error:

[0031]

[0032]

[0033] Step S3.4: Update f i Network, in each training, in order to obtain the static ideal probability transfer function Use a neural network f(s,m i ) to predict s′ * , update f(s,m i ) to maximize the following loss function:

[0034]

[0035] The first item in square brackets requires that the next state be The second requirement is to limit the predicted next state to N(s,m i ) set, the hyperparameter λ is based on the transfer model f i (s,m i ) coefficient;

[0036] Step S3.5: When training the f network, the output uses a technique similar to the residual network. The output is the difference between the predicted state and the input state Δ = s′ - s. Therefore, the predicted next state is s′ f =s+f(s,m);

[0037] Step S3.6: Update Q i network, using QSS network and f i Network, Q i The update method is:

[0038]

[0039] The next state predicted by the second item is closer to the Q of the next state I2Q that appears frequently in the data. o The value is closer to the true value.

[0040] Preferably, the step S4 includes:

[0041] Step S4.1: Initialize the Q network, QSS network, and state transition model f of each agent, as well as the corresponding target Q network and target QSS network. Then initialize the CityFlow traffic environment, load the road network data and traffic flow data, and start simulation training.

[0042] Step S4.2: The CityFlow simulator receives road network data and vehicle flow data as input, constructs a simulated traffic network based on it, and automatically generates vehicles to enter the network during the simulation, driving them along predetermined routes to their destinations. At each moment, the simulator accurately simulates the behavior of each vehicle, providing detailed traffic flow information, and effectively accelerates the simulation process through multi-threading technology.

[0043] Step S4.3: The simulator obtains external control action information to control the signal phase and signal duration of the intersection. After each green light signal ends, the traffic light flashes yellow for 3 seconds, then the intersection remains in a full red state for 2 seconds before switching to the next traffic signal.

[0044] Step S4.4: Each agent stores the experience data generated by the interaction with the environment into its own concurrent experience replay buffer In this process, each agent first obtains the local state information of the environment, and then determines whether the macro action is completed. If the macro action is completed, the upper-level strategy of the agent uses the ∈-greedy method to decide the macro action, and then determines the signal phase to be output according to the macro action. i Sum signal duration d i , the lower-level strategy resets the signal duration and outputs the signal phase to the environment i If the macro action is not completed, the lower-level strategy executes the lower-level action according to the signal phase specified by the unfinished macro action and updates the remaining signal duration.

[0045] Step S4.5: At each time step t, the agent saves local state information, macro actions, lower-level actions, and reward information to the concurrent experience replay cache;

[0046] Step S4.6: During the agent model parameter update process, each agent first samples data from its own experience replay buffer and filters out the experience data at the end of the macro action (s i ,m i ,s i ′ ,r i ), then update them separately And update the target network every 5 training rounds

[0047] The multi-intersection asynchronous traffic signal control system provided by the present invention in a non-communication environment includes:

[0048] Module M1: Design a traffic signal control system to control traffic signals and duration at intersections, and design system optimization objectives, including average vehicle travel time and throughput;

[0049] Module M2: Model the multi-intersection traffic signal control problem as a partially observable Markov decision process based on macro-actions, and design the state, action, reward, and goal of the intelligent agent;

[0050] Module M3: An independent ideal Q algorithm based on macro actions to achieve phase and duration control of traffic signals;

[0051] Module M4 deploys the Mac-I2Q algorithm to the intersection agent and trains it in a simulated traffic environment. The agent interacts with the environment, stores and transfers data, and updates network parameters. This process is repeated until the entire simulation process is completed.

[0052] Preferably, the module M1 includes:

[0053] Module M1.1: Vehicles in the system enter the road network. The road network is the environment in which motor vehicles move. It contains N intersections and roads connecting the intersections. The road network is described as a directed graph. in represents an intersection, and ij represents the road leading to intersection j; ji represents the road leading to intersection i from intersection j;

[0054] Module M1.2: Vehicles enter intersections in the system. For an intersection i, the exit road refers to the road leading from intersection i to a certain intersection, that is, ij∈ε. In contrast, ij is the entry road of intersection j. An intersection connects 8 roads, 4 of which are entry roads and 4 are exit roads.

[0055] Module M1.3: When entering an intersection, vehicles select lanes based on their travel direction. A road has three lanes, corresponding to the three travel directions of vehicles at the intersection: left turn, straight ahead, and right turn. Lanes define the travel direction of vehicles within them. Assume that vehicles in each lane do not change lanes within the lane. The lane entering a road is called the entry lane, and the lane exiting a road is called the exit lane. The eight roads at an intersection contain 24 lanes.

[0056] Module M1.4: A vehicle enters an intersection from an entry lane and then leaves from an exit lane. This behavior is called a traffic movement. Traffic movements are divided into left turns, straight ahead, and right turns based on the direction of the vehicle passing through the intersection. Each lane of a road corresponds to a traffic movement.

[0057] Module M1.5: The intersection decides the next signal phase when switching signals. A signal phase is a combination of traffic signals that lasts for a period of time. There are four signal phases available at an intersection: east-west straight ahead, north-south straight ahead, east-west left turn, and north-south left turn. When the intersection makes a traffic signal decision, it makes the decision among the available signal phases.

[0058] Module M1.6: The system records the driving trajectory information of each vehicle entering the road network and calculates the average vehicle driving time and throughput. The average vehicle driving time is the average length of the driving time of all vehicles. The throughput refers to the number of vehicles that complete the driving process in the road network per unit time. The driving time of each vehicle is calculated from the starting position of the vehicle to the time it takes to reach the destination. When a vehicle enters an intersection, its movement is affected by traffic signals. If it encounters a red light, it waits in the intersection until the corresponding green light comes on before continuing to drive.

[0059] Preferably, the module M2 includes:

[0060] Module M2.1: Modeling a multi-agent system as a distributed partially observable Markov decision process based on macro-actions, written as a tuple in It's a picture Every intelligent agent, is the communication link and the road between the agents at adjacent intersections. The number of agents is is the global state space, is the global action space, each agent can only observe the local information o of the global state s i , S i is the local state space of agent i, and The joint observation space of the agent is Accordingly, is the local action space of agent i, and is a collection of joint macro actions, is a finite set of macro-actions for each agent i; is the joint macro-measurement set, is the macroscopic measurement of agent i, setting

[0061] Module M2.2: At each time step t, the environment state starts to transfer from s, and the system takes joint action a and transfers to the new state s ′ , the transfer process is based on the state transfer function P(s ′ |s,a), O(o,a,s ′ )=P(o|a,s ′ ) means taking joint action a and then transferring to s ′ The probability of obtaining a global observation o when R:S×A→R is the reward function. After taking a in s, the immediate reward is returned. Since the agent actually only observes local state information, the strategy π of each agent i i is the local observation information o i Mapping to actions;

[0062] Module M2.3: The partial state of the system environment observed at each time step t is received as the state information s of the reinforcement learning agent i i In the lower-level strategy, this includes the queue length of vehicles on each lane of intersection i and the current signal phase of the intersection in Is entering the lane i The number of vehicles on phase i is the signal phase of intersection i; in the upper strategy, agent i will be in state s at the end of the macro action i As its macro information i ;

[0063] Module M2.4: In the upper-level strategy, agent i performs a macro-action m at each i At the end, the next macro action needs to be decided. The macro action contains the signal phase and duration information. The optional signal duration includes {10s, 15s, 20s, 25s}, so there are 16 optional macro actions in total. The lower-level strategy of agent i executes action a according to the macro action information. i , that is, the signal phase specified by the macro action, at each time step t, the corresponding signal phase is executed, and each action a i Lasts for 5 seconds, and determines whether the macro action is completed based on the signal duration information;

[0064] Module M2.5: At each time step t, agent i takes a macro action m = < τ m ,I m ,π m >, where τ is the time required for the joint macro action m to end, initialization set Depends on the macro action observation history information of agent i denotes the lower-level strategy used to implement macro-action m. Considering the termination condition of the macro-action, the transition probability is redefined as P(s′,τ|s,m), where τ is the time required for the joint macro-action m to end. The end of the joint macro-action means that each agent has completed its own macro-action; Z(z,m,s′)=P(z|m,s′) denotes the likelihood model of the joint macro-action.

[0065] Module M2.6: The goal of the agent is to find a joint upper-level strategy Φ = × i Φ i , make decisions at the macro-action level so that the value of Φ starting from s0 is optimized, and the corresponding action value function is denote the expected cumulative rewards at a specific action and state, respectively;

[0066] Module M2.7: In the lower-level policy, the reward r of agent i i Defined as the immediate feedback of the current action, that is, the negative value of the average number of vehicles entering the lane in Is entering the lane i The number of vehicles on the upper layer; in the upper layer strategy, agent i takes macro action m i r at the end i Considered as the reward information for this macro action.

[0067] Preferably, the module M3 includes:

[0068] Module M3.1: Use macro-action concurrent experience replay trajectories to filter macro-action experience data. The agent saves its action, observation, and reward information in the experience replay pool at each time step t. When sampling experience data, only the experience data at the end of the macro action is selected as valid data, and the rest of the data is discarded.

[0069] Module M3.2: Observing history in a macro action i In the example, each agent independently selects a macro action m i , save reward information, At the end of the macro action, the agent obtains a new macro measurement s′ and a new macro action observation history information h′ = <h m ,m i ,s′>; Accordingly, the experience tuple collected by agent i is expressed as where s i It is used to select macro action m i Macro information;

[0070] Module M3.3: Update the QSS network. In each training, Update in a way that minimizes the following TD error:

[0071]

[0072]

[0073] Module M3.4: Update f i Network, in each training, in order to obtain the static ideal probability transfer function Use a neural network f(s,m i ) to predict s′ * , update f(s,m i ) to maximize the following loss function:

[0074]

[0075] The first item in square brackets requires that the next state be The second requirement is to limit the predicted next state to N(s,m i ) set, the hyperparameter λ is based on the transfer model f i (s,m i ) coefficient;

[0076] Module M3.5: When training the f network, the output uses a technique similar to the residual network, and the output is the difference Δ=s between the predicted state and the input state ′ -s, so the predicted next state is s f ′ =s+f(s,m);

[0077] Module M3.6: Update Q i network, using QSS network and f i Network, Q i The update method is:

[0078]

[0079] The next state predicted by the second item is closer to the Q of the next state I2Q that appears frequently in the data. o The value is closer to the true value.

[0080] Preferably, the module M4 includes:

[0081] Module M4.1: Initializes the Q network, QSS network, and state transition model f of each agent, as well as the corresponding target Q network and target QSS network. Then, initializes the CityFlow traffic environment, loads road network data and traffic flow data, and starts simulation training.

[0082] Module M4.2: The CityFlow simulator receives road network data and vehicle flow data as input, constructs a simulated traffic network based on this data, and automatically generates vehicles to enter the network during the simulation, driving them along predetermined routes to their destinations. At each moment, the simulator accurately simulates the behavior of each vehicle, providing detailed traffic flow information and effectively accelerating the simulation process through multi-threading technology.

[0083] Module M4.3: The simulator obtains external control action information to control the signal phase and signal duration at the intersection. After each green light signal ends, the traffic light flashes yellow for 3 seconds, then the intersection remains in a full red state for 2 seconds before switching to the next traffic signal.

[0084] Module M4.4: Each agent stores the experience data generated by its interaction with the environment into its own concurrent experience replay buffer In this process, each agent first obtains the local state information of the environment, and then determines whether the macro action is completed. If the macro action is completed, the upper-level strategy of the agent uses the ∈-greedy method to decide the macro action, and then determines the signal phase to be output according to the macro action. i Sum signal duration d i , the lower-level strategy resets the signal duration and outputs the signal phase to the environment i If the macro action is not completed, the lower-level strategy executes the lower-level action according to the signal phase specified by the unfinished macro action and updates the remaining signal duration.

[0085] Module M4.5: At each time step t, the agent saves local state information, macro actions, lower-level actions, and reward information to the concurrent experience replay cache;

[0086] Module M4.6: The model parameter update process of the agent. Each agent first samples data from its own experience replay buffer and filters out the experience data at the end of the macro action (s i ,m i ,s i ′ ,r i ), then update them separately And update the target network every 5 training rounds

[0087] Compared with the prior art, the present invention has the following beneficial effects:

[0088] (1) The MacDec-POMDP framework adopted in the present invention is a two-layer architecture. The upper-layer strategy can control the duration of macro actions, thereby meeting the asynchronous signal control requirements of multiple intersections;

[0089] (2) The Mac-I2Q algorithm proposed in this paper solves the environmental non-steady-state problem that occurs in independent Q learning strategies in distributed multi-agent environments, making the performance of the agents closer to the theoretical optimal value and accelerating the algorithm convergence speed, shortening the average vehicle travel time and improving vehicle throughput;

[0090] (3) The present invention can provide a more stable traffic control signal for the traffic light system, which helps to improve the efficiency of transportation management; the variable duration signal strategy proposed by the present invention can provide a more stable signal, which helps to maintain the stability of traffic order; the present invention has a reasonable structure and is easy to use, and can overcome the defects of the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0092] Figure 1 This is an example diagram of the Mac-I2Q algorithm framework in an embodiment of the present invention;

[0093] Figure 2 This is an example diagram of the asynchronous multi-agent reinforcement learning framework in an embodiment of the present invention. DETAILED DESCRIPTION

[0094] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0095] Example 1

[0096] The present invention provides a method for swarm intelligence perception and scheduling of UAVs, referring to Figure 1 and Figure 2 As shown, the method specifically includes the following contents:

[0097] Step S1: Design a traffic signal control system to control the traffic signals and duration at the intersection, and design system optimization objectives such as average vehicle travel time and throughput;

[0098] Step S2: Model the multi-intersection traffic signal control problem as a partially observable Markov decision process based on macro-actions, and design the state, action, reward, and goal of the agent;

[0099] Step S3: Propose an independent ideal Q algorithm based on macro-actions to achieve phase and duration control of traffic signals and solve the problem of environmental non-steady state in the learning of independent intelligent agents;

[0100] In step S4, the Mac-I2Q algorithm is deployed to the intersection agent and trained in a simulated traffic environment. The agent interacts with the environment, stores and transfers data, and updates network parameters. This process is repeated until the entire simulation process is completed.

[0101] Specifically, step S1 includes:

[0102] Step S1.1: The vehicles in the system enter the road network. The road network is the environment in which motor vehicles move, which includes N intersections and roads connecting the intersections. The road network is described as a directed graph in represents an intersection, and ij represents the road leading to intersection j; conversely, ji represents the road leading to intersection i from intersection j;

[0103] Step S1.2: A vehicle enters an intersection in the system. For an intersection i, the exit road is the road leading from intersection i to a certain intersection (for example, j), that is, ij∈ε. Conversely, ij is the entrance road to intersection j. Generally speaking, an intersection connects 8 roads, 4 of which are entrance roads and 4 are exit roads.

[0104] Step S1.3: When a vehicle enters an intersection, it selects a lane based on its direction of travel. A road has three lanes, corresponding to the three possible directions of travel for vehicles at the intersection: left turn, straight ahead, and right turn. A lane specifies the direction of travel for vehicles within it. It is assumed that vehicles in each lane do not change lanes within the lane. The lane entering a road is called the entry lane; the lane exiting a road is called the exit lane. The eight roads at an intersection contain 24 lanes.

[0105] Step S1.4: A vehicle enters an intersection from an entry lane and then exits from an exit lane. This behavior is called a traffic movement. Traffic movements can be divided into left turns, straight ahead, and right turns based on the direction of the vehicle passing through the intersection. Each lane of a road corresponds to a traffic movement.

[0106] Step S1.5: The intersection decides the next signal phase when switching signals. A signal phase is a combination of traffic signals that lasts for a period of time. There are four signal phases available at an intersection: east-west straight ahead, north-south straight ahead, east-west left turn, and north-south left turn. When the intersection makes a traffic signal decision, it makes the decision among the available signal phases.

[0107] Step S1.6: The system records the trajectory of each vehicle entering the road network and calculates the average vehicle travel time and throughput. The average vehicle travel time is the average length of time all vehicles travel. Throughput refers to the number of vehicles that complete a journey on the road network per unit time. The travel time for each vehicle is calculated from the vehicle's starting position until it reaches its destination. When a vehicle enters an intersection, its movement is affected by traffic signals. If it encounters a red light, it waits in the intersection until the corresponding green light comes on before continuing.

[0108] Specifically, step S2 includes:

[0109] Step S2.1: Model the multi-agent system as a distributed partially observable Markov decision process based on macro-actions, which can be written as a tuple in It's a picture is each agent, ij∈ε is the communication link, and is also the road between agents at adjacent intersections. The number of agents is in, is the global state space, and is the global action space. Each agent can only observe local information o of the global state s i , S i is the local state space of agent i, and The joint observation space of the agent is Accordingly, is the local action space of agent i, and is a collection of joint macro actions, is a finite set of macro-actions for each agent i. is the joint macro-measurement set, is the macroscopic measurement of agent i, setting

[0110] Step S2.2: At each time step t, the environment state starts to transfer from s, and the system takes the joint action a and transfers to the new state s ′ , the transfer process is based on the state transfer function P(s ′ |s,a). O(o,a,s ′ )=P(o|a,s ′ ) means taking joint action a and then transferring to s ′ The probability of obtaining a global observation o when . R:S×A→R is the reward function, which returns the immediate reward after taking a in s. Since the agent can actually only observe local state information, the strategy π of each agent i i is the local observation information oi Mapping to actions;

[0111] Step S2.3: The partial state of the system environment observed at each time step t is received as the state information s of the reinforcement learning agent i i In the lower-level strategy, this includes the queue length of vehicles in each lane of intersection i and the current signal phase of the intersection in Is entering the lane i The number of vehicles on phase i is the signal phase of intersection i. In the upper strategy, agent i will be in the state s at the end of the macro action. i As its macro information i This state information design allows the agent to understand its environment more accurately and make more reasonable decisions;

[0112] Step S2.4: In the context of traffic signal control, the core task of the intersection agent is to select the appropriate signal phase and duration to maximize traffic efficiency. In the upper-level strategy, agent i takes each macro action m i At the end, the next macro action needs to be decided. The macro action contains the signal phase and duration information. The optional signal duration includes {10s, 15s, 20s, 25s}, so there are 16 optional macro actions in total;

[0113] The lower-level strategy of agent i executes action a according to the macro-action information i , that is, the signal phase specified by the macro action. At each time step t, the corresponding signal phase is executed, and each action a i Lasts for 5 seconds, and determines whether the macro action is completed based on the signal duration information;

[0114] Step S2.5: At each time step t, agent i takes macro action m = < τ m ,I m ,π m >, where τ is the time required for the joint macro action m to end, initialization set Depends on the macro action observation history information of agent i represents the lower-level strategy to implement the macro action m. Considering the end condition of the macro action, the transition probability is redefined as P(s ′ ,τ|s,m), where τ is the time required for the joint macro action m to end. The end of the joint macro action means that each agent has completed its own macro action; Z(z,m,s ′ )=P(z|m,s ′ ) represents the likelihood model of joint macro-measurement;

[0115] Step S2.6: The goal of the agent is to find a joint upper-level strategy Φ = × i Φ i , make decisions at the macro-action level so that the value of Φ starting from s0 is optimized, and the corresponding action value function is denote the expected cumulative rewards at a specific action and state, respectively;

[0116] Step S2.7: In the lower policy, the reward r of agent i i Defined as the immediate feedback of the current action, that is, the negative value of the average number of vehicles entering the lane in Is entering the lane i The number of vehicles on.

[0117] In the upper-level strategy, agent i takes macro action m i r at the end i This design allows the agent to evaluate the effectiveness of its decisions at a macro level, thereby better planning future action sequences.

[0118] Specifically, step S3 includes:

[0119] Step S3.1: Use the macro-action concurrent experience replay trajectory to filter the macro-action experience data. At each time step t, the agent saves its action, observation, and reward information in the experience replay pool. However, when sampling experience data, only the experience data at the end of the macro action is selected as valid data, and the rest of the data is discarded.

[0120] Step S3.2: In a macro action observation history h i In the example, each agent independently selects a macro action m i , save reward information, The agent obtains a new macro-measurement at the end of the macro action ′ , and obtain new macro action observation history information h ′ = <h m ,m i ,s ′ >. Accordingly, the experience tuple collected by agent i is expressed as where s i It is used to select macro action m i Macro information;

[0121] Step S3.3: Update the QSS network; in each training, Update in a way that minimizes the following TD-error:

[0122]

[0123]

[0124] Step S3.4: Update f i Network; In each training, in order to transfer the static ideal probability function Now, use a neural network f(s,m i ) to predict s ′* , update f(s,m i ) to maximize the following loss function

[0125]

[0126] The first item in square brackets requires that the next state be The second requirement is to limit the predicted next state to N(s,m i ) set, the hyperparameter λ is based on the transfer model f i (s,m i ) coefficient;

[0127] Step S3.5: When training the f network, the output uses a technique similar to the residual network, and the output is the difference Δ=s between the predicted state and the input state ′ -s, instead of directly predicting the next state. Therefore, the predicted next state is s f ′ =s+f(s,m);

[0128] Step S3.6: Update Q i network; using QSS network and f i Network, Q i The update method is:

[0129]

[0130] The next state predicted by the second item is closer to the next state that appears frequently in the data, which means that the probability of the state that appears in the environmental transition probability is not too small. i The value will be closer to the true value, and the situation of low-probability state will be avoided, which can be applied to random environments.

[0131] Specifically, step S4 includes:

[0132] Step S4.1: Initialize the Q network, QSS network, and state transition model f of each agent, as well as the corresponding target Q network and target QSS network. Then initialize the CityFlow traffic environment, load the road network data and traffic flow data, and start simulation training.

[0133] Step S4.2: The CityFlow simulator receives road network and vehicle flow data as input, constructs a simulated traffic network, and automatically generates vehicles that enter the network and drive to their destinations along predetermined routes. The simulator accurately simulates the behavior of each vehicle at each moment, providing detailed traffic flow information and accelerating the simulation process through multithreading.

[0134] Step S4.3: The simulator obtains the control action information output by the external control method to control the signal phase and signal duration at the intersection. According to the traditional signal setting, after each green light signal ends, the traffic light flashes yellow for 3 seconds, then the intersection continues to have a full red signal state for 2 seconds, and then switches to the next traffic signal.

[0135] Step S4.4: Each agent stores the experience data generated by the interaction with the environment into its own concurrent experience replay buffer In this process, each agent first obtains the local state information of the environment, and then determines whether the macro action is completed. If the macro action is completed, the upper-level strategy of the agent uses the ∈-greedy method to decide the macro action, and then determines the signal phase to be output based on the macro action. i Sum signal duration d i , the lower-level strategy resets the signal duration and outputs the signal phase to the environment i If the macro action is not completed, the lower-level strategy executes the lower-level action according to the signal phase specified by the unfinished macro action and updates the remaining signal duration;

[0136] Step S4.5: At each time step t, the agent saves local state information, macro actions, lower-level actions, and reward information to the concurrent experience replay cache;

[0137] Step S4.6: The model parameter update process of the agent, each agent first samples data from its own experience replay buffer and filters out the experience data at the end of the macro action, (s i ,m i ,s i ′ ,r i ), and then update the parameter update formula given above. And update the target network every 5 training rounds

[0138] This example focuses on the problem of asynchronous traffic signal control at multiple intersections without inter-intersection communication. A macro-action-based multi-agent reinforcement learning ideal independent Q-network algorithm is proposed to design signal decision strategies for agents at intersections. First, a traffic signal control system is designed to control the traffic signals and durations at the intersections, with system optimization objectives such as average vehicle travel time and throughput. Second, the multi-intersection traffic signal control problem is modeled as a partially observable Markov decision process based on macro-actions, designing the states, actions, rewards, and goals of the agents. Finally, a macro-action-based independent ideal Q algorithm is proposed to control the phase and duration of traffic signals, addressing the environmental non-stationary nature of the learning environment for independent agents. Experiments on both synthetic and real-world urban datasets demonstrate that Mac-I2Q is significantly effective in asynchronous traffic signal decision-making, achieving shorter average vehicle travel times and higher throughput than current optimal traffic signal control methods.

[0139] Example 2

[0140] The present invention also provides a multi-intersection asynchronous traffic signal control system in a non-communication environment. The multi-intersection asynchronous traffic signal control system can be implemented by executing the process steps of the multi-intersection asynchronous traffic signal control method. That is, those skilled in the art can understand the multi-intersection asynchronous traffic signal control method as a preferred implementation of the multi-intersection asynchronous traffic signal control system in the non-communication environment.

[0141] The multi-intersection asynchronous traffic signal control system in the non-communication environment includes:

[0142] Module M1: Design traffic signal control systems, control traffic signals and duration at intersections, and design system optimization objectives such as average vehicle travel time and throughput;

[0143] The module M1 includes:

[0144] Module M1.1: Vehicles in the system enter the road network. The road network is the environment in which motor vehicles move, consisting of N intersections and roads connecting them. The road network is described as a directed graph in represents an intersection, and ij represents the road from intersection j; conversely, ji represents the road from intersection j to intersection i.

[0145] Module M1.2: When a vehicle enters an intersection in the system, for an intersection i, the exit road is the road leading from intersection i to a certain intersection (e.g., j), i.e., ij∈ε. Conversely, ij is the entry road to intersection j. Generally speaking, an intersection connects eight roads, four of which are entry roads and four are exit roads.

[0146] Module M1.3: Vehicles entering an intersection select lanes based on their direction of travel. A road has three lanes, corresponding to the three possible directions of travel for vehicles at the intersection: left turn, straight ahead, and right turn. Lanes define the direction of travel for vehicles within them. It is assumed that vehicles in each lane do not change lanes within their lanes. The lane entering a road is called the entry lane; the lane exiting a road is called the exit lane. The eight roads at an intersection contain 24 lanes.

[0147] Module M1.4: A vehicle entering an intersection from an entry lane and then exiting from an exit lane is called a traffic movement. Traffic movements can be categorized as left turns, straight ahead, and right turns based on the direction of travel through the intersection. Each lane on a road corresponds to a traffic movement.

[0148] Module M1.5: The intersection decides the next signal phase when switching signals. A signal phase is a combination of traffic signals that lasts for a period of time. There are four signal phases available at an intersection: east-west straight ahead, north-south straight ahead, east-west left turn, and north-south left turn. When the intersection makes a traffic signal decision, it makes the decision among the available signal phases.

[0149] Module M1.6: The system records the trajectory of each vehicle entering the road network and calculates average vehicle travel time and throughput. Average vehicle travel time is the average length of time all vehicles travel. Throughput refers to the number of vehicles that complete a journey on the road network per unit time. The travel time for each vehicle is calculated from the vehicle's starting position until it reaches its destination. When a vehicle enters an intersection, its movement is affected by traffic signals. If it encounters a red light, it waits in the intersection until the corresponding green light comes on before continuing.

[0150] Module M2: Model the multi-intersection traffic signal control problem as a partially observable Markov decision process based on macro-actions, and design the state, action, reward, and goal of the intelligent agent;

[0151] The module M2 includes:

[0152] Module M2.1: Modeling a multi-agent system as a distributed partially observable Markov decision process based on macro-actions, which can be written as a tuple in It's a picture is each agent, ij∈ε is the communication link, and is also the road between agents at adjacent intersections. The number of agents is in, is the global state space, and is the global action space. Each agent can only observe local information o of the global state s i , S i is the local state space of agent i, and The joint observation space of the agent is Accordingly, is the local action space of agent i, and is a collection of joint macro actions, is a finite set of macro-actions for each agent i. is the joint macro-measurement set, is the macroscopic measurement of agent i, setting

[0153] Module M2.2: At each time step t, the environment state starts to transfer from s, and the system takes a joint action a and transfers to the new state s ′ , the transfer process is based on the state transfer function P(s ′ |s,a). O(o,a,s ′ )=P(o|a,s ′ ) means taking joint action a and then transferring to s ′ The probability of obtaining a global observation o when . R:S×A→R is the reward function, which returns the immediate reward after taking a in s. Since the agent can actually only observe local state information, the strategy π of each agent i i is the local observation information o i Mapping to actions.

[0154] Module M2.3: The partial state of the system environment observed at each time step t is received as the state information s of the reinforcement learning agent i i In the lower-level strategy, this includes the queue length of vehicles in each lane of intersection i and the current signal phase of the intersection in Is entering the lane i The number of vehicles on phase i is the signal phase of intersection i. In the upper strategy, agent i will be in the state s at the end of the macro action. i As its macro information i This design of state information allows the agent to understand its environment more accurately and make more reasonable decisions.

[0155] Module M2.4: In the context of traffic signal control, the core task of the intersection agent is to select the appropriate signal phase and duration to maximize traffic efficiency. In the upper-level strategy, agent i performs each macro action m iAt the end, the next macro action needs to be decided. The macro action contains the signal phase and duration information. The optional signal duration includes {10s, 15s, 20s, 25s}, so there are 16 optional macro actions in total.

[0156] The lower-level strategy of agent i executes action a according to the macro-action information i , that is, the signal phase specified by the macro action. At each time step t, the corresponding signal phase is executed, and each action a i It lasts for 5 seconds, and determines whether the macro action is completed based on the signal duration information.

[0157] At each time step t, agent i takes macro action m = < τ m ,I m ,π m >, where τ is the time required for the joint macro action m to end, initialization set Depends on the macro action observation history information of agent i represents the lower-level strategy to implement the macro action m. Considering the end condition of the macro action, the transition probability is redefined as P(s ′ ,τ|s,m), where τ is the time required for the joint macro action m to end. The end of the joint macro action means that each agent has completed its own macro action; Z(z,m,s ′ )=P(z|m,s ′ ) represents the likelihood model of joint macro-measurement.

[0158] Module M2.5: In Mac-DecPOMDP, the agent’s goal is to find a joint upper-level policy Φ = × i Φ i , make decisions at the macro-action level so that the value of Φ starting from s0 is optimized, and the corresponding action value function is denote the expected cumulative rewards at a specific action and state, respectively.

[0159] Module M2.6: In the lower-level policy, the reward r of agent i i Defined as the immediate feedback of the current action, that is, the negative value of the average number of vehicles entering the lane in Is entering the lane i This reward design motivates the agent to reduce the queue length, thereby improving traffic efficiency.

[0160] In the upper-level strategy, agent i takes macro action m i r at the end iThis design allows the agent to evaluate the effectiveness of its decisions at a macro level, thereby better planning future action sequences.

[0161] Module M3: Proposes an independent ideal Q algorithm based on macro-actions to achieve phase and duration control of traffic signals and solve the problem of environmental non-stationary state in the learning of independent intelligent agents;

[0162] The module M3 includes:

[0163] Module M3.1: Use macro-action concurrent experience replay trajectories to filter macro-action experience data. The agent saves its actions, observations, and rewards in the experience replay pool at each time step. However, when sampling experience data, only the experience data at the end of the macro action is selected as valid data, and the rest of the data is discarded.

[0164] Module M3.2: Observing history in a macro action i In the example, each agent independently selects a macro action m i , save reward information, The agent obtains a new macro-measurement at the end of the macro action ′ , and obtain new macro action observation history information h ′ = <h m ,m i ,s ′ >. Accordingly, the experience tuple collected by agent $i$ is expressed as where s i It is used to select macro action m i Macro measurement information.

[0165] Module M3.3: Update the QSS network. In each training, Update in a way that minimizes the following TD error:

[0166]

[0167]

[0168] Module M3.4: Update f i Network. In each training, in order to transfer the static ideal probability function Now, use a neural network f(s,m i ) to predict s′ * , update f(s,m i ) to maximize the following loss function

[0169]

[0170] The first item in square brackets requires that the next state be The second requirement is to limit the predicted next state to N(s,m i ) set, the hyperparameter λ is based on the transfer model f i (s,m i ) coefficient.

[0171] Module M3.5: When training the f network, the output uses a technique similar to the residual network, which outputs the difference between the predicted state and the input state Δ = s′ - s, rather than directly predicting the next state. Therefore, the predicted next state is s′ f =s+f(s,m).

[0172] Module M3.6: Update Q i Network. Using QSS network and f i Network, Q i The update method is:

[0173]

[0174] The next state predicted by the second item is closer to the next state that appears frequently in the data, which means that the probability of the state that appears in the environmental transition probability is not too small. i The value will be closer to the true value, and the situation of low-probability state will be avoided, which can be applied to random environments.

[0175] Module M4: Deploy the Mac-I2Q algorithm to the intersection agent and train it in a simulated traffic environment. The agent interacts with the environment, stores and transfers data, and updates network parameters. This process is repeated until the entire simulation process is completed.

[0176] The module M4 includes:

[0177] Module M4.1: Initializes the Q network, QSS network, and state transition model f of each agent, as well as the corresponding target Q network and target QSS network. Then, initializes the CityFlow traffic environment, loads road network data and traffic flow data, and starts simulation training.

[0178] Module M4.2: The CityFlow simulator receives road network and vehicle flow data as input, constructs a simulated traffic network, and automatically generates vehicles that enter the network and drive to their destinations along predetermined routes. The simulator accurately simulates the behavior of each vehicle at each moment, providing detailed traffic flow information and accelerating the simulation process through multi-threading technology.

[0179] Module M4.3: The simulator acquires the control action information output by the external control method to control the signal phase and signal duration at the intersection. According to the traditional signal setting, after each green signal, the traffic light flashes yellow for 3 seconds, then the intersection continues to have a full red signal state for 2 seconds, and then switches to the next traffic signal.

[0180] Module M4.4: Each agent stores the experience data generated by its interaction with the environment into its own concurrent experience replay buffer In this process, each agent first obtains the local state information of the environment, and then determines whether the macro action is completed. If the macro action is completed, the upper-level strategy of the agent uses the ∈-greedy method to decide the macro action, and then determines the signal phase to be output based on the macro action. i Sum signal duration d i , the lower-level strategy resets the signal duration and outputs the signal phase to the environment i If the macro action is not completed, the lower-level strategy executes the lower-level action according to the signal phase specified by the unfinished macro action and updates the remaining signal duration;

[0181] Module M4.5: At each time step t, the agent saves local state information, macro actions, lower-level actions, and reward information to the concurrent experience replay cache;

[0182] Module M4.6: The model parameter update process of the agent. Each agent first samples data from its own experience replay buffer and filters out the experience data at the end of the macro action. (s i ,m i ,s i ′ ,r i ), and then update the parameter update formula given above. And update the target network every 5 training rounds

[0183] Those skilled in the art will appreciate that, in addition to implementing the system, device, and various modules provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, and the like by logically programming the method steps. Therefore, the system, device, and various modules provided by the present invention can be considered a hardware component, and the modules included therein for implementing various programs can also be considered structures within the hardware component; the modules for implementing various functions can also be considered both software programs for implementing the method and structures within the hardware component.

[0184] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. A method for controlling asynchronous traffic signals at multiple intersections in a non-communication environment, characterized in that: include: Step S1: Design a traffic signal control system to control the traffic signals and duration at the intersection, and design system optimization objectives, including average vehicle travel time and throughput; Step S2: Model the multi-intersection traffic signal control problem as a partially observable Markov decision process based on macro-actions, and design the state, action, reward, and goal of the agent; Step S3: Implementing phase and duration control of traffic signals based on an independent ideal Q algorithm based on macro actions; In step S4, the Mac-I2Q algorithm is deployed to the intersection agent and trained in a simulated traffic environment. The agent interacts with the environment, stores and transfers data, and updates network parameters. This process is repeated until the entire simulation process is completed.

2. The method for controlling asynchronous traffic signals at multiple intersections in a non-communication environment according to claim 1, characterized in that: The step S1 comprises: Step S1.1: The vehicle in the system enters the road network. The road network is the environment in which motor vehicles move, which includes N intersections and roads connecting the intersections. The road network is described as a directed graph. in represents an intersection, and ij represents the road leading to intersection j; ji represents the road leading to intersection i from intersection j; Step S1.2: A vehicle enters an intersection in the system. For an intersection i, the exit road is the road leading from intersection i to a certain intersection, i.e., ij∈ε. Conversely, ij is the entrance road to intersection j. An intersection is connected to 8 roads, 4 of which are entrance roads and 4 are exit roads. Step S1.3: When a vehicle enters an intersection, it selects a lane based on its direction of travel. A road has three lanes, corresponding to the three possible directions of travel for vehicles at the intersection: left turn, straight ahead, and right turn. A lane specifies the direction of travel for vehicles within it. Assume that vehicles in each lane do not change lanes within the lane. The lane entering a road is called the entry lane, and the lane exiting a road is called the exit lane. The eight roads at an intersection contain 24 lanes. Step S1.4: A vehicle enters an intersection from an entry lane and then leaves from an exit lane. This behavior is called a traffic movement. Traffic movements are classified into left turns, straight ahead, and right turns based on the direction of the vehicle passing through the intersection. Each lane of a road corresponds to a traffic movement. Step S1.5: The intersection decides the next signal phase when switching signals. A signal phase is a combination of traffic signals that lasts for a period of time. There are four signal phases available at an intersection: east-west straight ahead, north-south straight ahead, east-west left turn, and north-south left turn. When the intersection makes a traffic signal decision, it makes the decision among the available signal phases. Step S1.6: The system records the driving trajectory information of each vehicle entering the road network and calculates the average vehicle driving time and throughput. The average vehicle driving time is the average length of the driving time of all vehicles. The throughput refers to the number of vehicles that complete the driving process in the road network per unit time. The driving time of each vehicle is calculated from the starting position of the vehicle to the time it takes to reach the destination. When a vehicle enters an intersection, its movement is affected by traffic signals. When it encounters a red light, it waits in the intersection until the corresponding green light comes on and then continues to drive.

3. The method for controlling asynchronous traffic signals at multiple intersections in a non-communication environment according to claim 1, characterized in that: The step S2 comprises: Step S2.1: Model the multi-agent system as a distributed partially observable Markov decision process based on macro-actions, written as a tuple in It's a picture Every intelligent agent, is the communication link, and also the road between the agents at adjacent intersections. The number of agents is is the global state space, is the global action space, each agent can only observe the local information o of the global state s i , S i is the local state space of agent i, and The joint observation space of the agent is Accordingly, is the local action space of agent i, and is a collection of joint macro actions, is a finite set of macro-actions for each agent i; is the joint macro-measurement set, is the macroscopic measurement of agent i, setting Step S2.2: At each time step t, the environment state starts to transfer from s. The system takes the joint action a and transfers to the new state s′. The transfer process is carried out according to the state transfer function P(s′|s,a). O(o,a,s′)=P(o|a,s′) represents the probability of obtaining the global observation o when transferring to s′ after taking the joint action a. R:S×A→R is the reward function. After taking a in s, the immediate reward is returned. Since the agent actually only observes local state information, the strategy π of each agent i i is the local observation information o i Mapping to actions; Step S2.3: The partial state of the system environment observed at each time step t is received as the state information s of the reinforcement learning agent i i In the lower-level strategy, this includes the queue length of vehicles on each lane of intersection i and the current signal phase of the intersection in Is entering the lane i The number of vehicles on phase i is the signal phase of intersection i; in the upper strategy, agent i will be in state s at the end of the macro action i As its macro information i ; Step S2.4: In the upper-level strategy, agent i takes each macro action m i At the end, the next macro action needs to be decided. The macro action contains the signal phase and duration information. The optional signal duration includes {10s, 15s, 20s, 25s}, so there are 16 optional macro actions in total. The lower-level strategy of agent i executes action a according to the macro action information. i , that is, the signal phase specified by the macro action, at each time step t, the corresponding signal phase is executed, and each action a i Lasts for 5 seconds, and determines whether the macro action is completed based on the signal duration information; Step S2.5: At each time step t, agent i takes macro action m = < τ m ,I m ,π m >, where τ is the time required for the joint macro action m to end, initialization set Depends on the macro action observation history information of agent i denotes the lower-level strategy used to implement macro-action m. Considering the termination condition of the macro-action, the transition probability is redefined as P(s′,τ|s,m), where τ is the time required for the joint macro-action m to end. The end of the joint macro-action means that each agent has completed its own macro-action; Z(z,m,s′)=P(z|m,s′) denotes the likelihood model of the joint macro-action. Step S2.6: The goal of the agent is to find a joint upper-level strategy Φ = × i Φ i , make decisions at the macro-action level so that the value of Φ starting from s0 is optimized, and the corresponding action value function is denote the expected cumulative rewards at a specific action and state, respectively; Step S2.7: In the lower policy, the reward r of agent i i Defined as the immediate feedback of the current action, that is, the negative value of the average number of vehicles entering the lane in Is entering the lane i The number of vehicles on the upper layer; in the upper layer strategy, agent i takes macro action m i r at the end i Considered as the reward information for this macro action.

4. The method for controlling asynchronous traffic signals at multiple intersections in a non-communication environment according to claim 1, characterized in that: The step S3 comprises: Step S3.1: Use the macro-action concurrent experience replay trajectory to filter the macro-action experience data. The agent saves its action, observation, and reward information in the experience replay pool at each time step t. When sampling experience data, only the experience data at the end of the macro action is selected as valid data, and the rest of the data is discarded. Step S3.2: In a macro action observation history h i In the example, each agent independently selects a macro action m i , save reward information, At the end of the macro action, the agent obtains a new macro measurement s′ and a new macro action observation history information h′ = <h m ,m i ,s′>; Accordingly, the experience tuple collected by agent i is expressed as where s i It is used to select macro action m i Macro information; Step S3.3: Update the QSS network. In each training, Update in a way that minimizes the following TD error: Step S3.4: Update f i Network, in each training, in order to obtain the static ideal probability transfer function Use a neural network f(s,m i ) to predict s′ * , update f(s,m i ) to maximize the following loss function: The first item in square brackets requires that the next state be The second requirement is to limit the predicted next state to N(s,m i ) set, the hyperparameter λ is based on the transfer model f i (s,m i ) coefficient; Step S3.5: When training the f network, the output uses a technique similar to the residual network. The output is the difference between the predicted state and the input state Δ = s′ - s. Therefore, the predicted next state is s′ f =s+f(s,m); Step S3.6: Update Q i network, using QSS network and f i Network, Q i The update method is: The next state predicted by the second item is closer to the Q of the next state I2Q that appears frequently in the data. i The value is closer to the true value.

5. The method for controlling asynchronous traffic signals at multiple intersections in a non-communication environment according to claim 1, characterized in that: The step S4 comprises: Step S4.1: Initialize the Q network, QSS network, and state transition model f of each agent, as well as the corresponding target Q network and target QSS network. Then initialize the CityFlow traffic environment, load the road network data and traffic flow data, and start simulation training. Step S4.2: The CityFlow simulator receives road network data and vehicle flow data as input, constructs a simulated traffic network based on it, and automatically generates vehicles to enter the network during the simulation, driving them along predetermined routes to their destinations. At each moment, the simulator accurately simulates the behavior of each vehicle, providing detailed traffic flow information, and effectively accelerates the simulation process through multi-threading technology. Step S4.3: The simulator obtains external control action information to control the signal phase and signal duration of the intersection. After each green light signal ends, the traffic light flashes yellow for 3 seconds, then the intersection remains in a full red state for 2 seconds before switching to the next traffic signal. Step S4.4: Each agent stores the experience data generated by the interaction with the environment into its own concurrent experience replay buffer In this process, each agent first obtains the local state information of the environment, and then determines whether the macro action is completed. If the macro action is completed, the upper-level strategy of the agent uses the ∈-greedy method to decide the macro action, and then determines the signal phase to be output according to the macro action. i Sum signal duration d i , the lower-level strategy resets the signal duration and outputs the signal phase to the environment i If the macro action is not completed, the lower-level strategy executes the lower-level action according to the signal phase specified by the unfinished macro action and updates the remaining signal duration. Step S4.5: At each time step t, the agent saves local state information, macro actions, lower-level actions, and reward information to the concurrent experience replay cache; Step S4.6: During the agent model parameter update process, each agent first samples data from its own experience replay buffer and filters out the experience data at the end of the macro action (s i ,m i ,s′ i ,r i ), and then update f respectively i , Q i , and update the target network every 5 training rounds 6. A multi-intersection asynchronous traffic signal control system in a non-communication environment, characterized in that: include: Module M1: Design a traffic signal control system to control traffic signals and duration at intersections, and design system optimization objectives, including average vehicle travel time and throughput; Module M2: Model the multi-intersection traffic signal control problem as a partially observable Markov decision process based on macro-actions, and design the state, action, reward, and goal of the intelligent agent; Module M3: An independent ideal Q algorithm based on macro actions to achieve phase and duration control of traffic signals; Module M4 deploys the Mac-I2Q algorithm to the intersection agent and trains it in a simulated traffic environment. The agent interacts with the environment, stores and transfers data, and updates network parameters. This process is repeated until the entire simulation process is completed.

7. The multi-intersection asynchronous traffic signal control system in a non-communication environment according to claim 6, characterized in that: The module M1 includes: Module M1.1: Vehicles in the system enter the road network. The road network is the environment in which motor vehicles move. It contains N intersections and roads connecting the intersections. The road network is described as a directed graph. in represents an intersection, and ij represents the road leading to intersection j; ji represents the road leading to intersection i from intersection j; Module M1.2: Vehicles enter intersections in the system. For an intersection i, the exit road refers to the road leading from intersection i to a certain intersection, that is, ij∈ε. In contrast, ij is the entry road of intersection j. An intersection connects 8 roads, 4 of which are entry roads and 4 are exit roads. Module M1.3: When entering an intersection, vehicles select lanes based on their travel direction. A road has three lanes, corresponding to the three travel directions of vehicles at the intersection: left turn, straight ahead, and right turn. Lanes define the travel direction of vehicles within them. Assume that vehicles in each lane do not change lanes within the lane. The lane entering a road is called the entry lane, and the lane exiting a road is called the exit lane. The eight roads at an intersection contain 24 lanes. Module M1.4: A vehicle enters an intersection from an entry lane and then leaves from an exit lane. This behavior is called a traffic movement. Traffic movements are divided into left turns, straight ahead, and right turns based on the direction of the vehicle passing through the intersection. Each lane of a road corresponds to a traffic movement. Module M1.5: The intersection decides the next signal phase when switching signals. A signal phase is a combination of traffic signals that lasts for a period of time. There are four signal phases available at an intersection: east-west straight ahead, north-south straight ahead, east-west left turn, and north-south left turn. When the intersection makes a traffic signal decision, it makes the decision among the available signal phases. Module M1.6: The system records the driving trajectory information of each vehicle entering the road network and calculates the average vehicle driving time and throughput. The average vehicle driving time is the average length of the driving time of all vehicles. The throughput refers to the number of vehicles that complete the driving process in the road network per unit time. The driving time of each vehicle is calculated from the starting position of the vehicle to the time it takes to reach the destination. When a vehicle enters an intersection, its movement is affected by traffic signals. If it encounters a red light, it waits in the intersection until the corresponding green light comes on before continuing to drive.

8. The multi-intersection asynchronous traffic signal control system in a non-communication environment according to claim 6, characterized in that: The module M2 includes: Module M2.1: Modeling a multi-agent system as a distributed partially observable Markov decision process based on macro-actions, written as a tuple in It's a picture Every intelligent agent, is the communication link, and also the road between the agents at adjacent intersections. The number of agents is is the global state space, is the global action space, each agent can only observe the local information o of the global state s i , S i is the local state space of agent i, and The joint observation space of the agent is Accordingly, is the local action space of agent i, and is a collection of joint macro actions, is a finite set of macro-actions for each agent i; is the joint macro-measurement set, is the macroscopic measurement of agent i, setting Module M2.2: At each time step t, the environment state starts to transfer from s. The system takes the joint action a and transfers to the new state s′. The transfer process is carried out according to the state transfer function P(s′|s,a). O(o,a,s′)=P(o|a,s′) represents the probability of obtaining the global observation o when transferring to s′ after taking the joint action a. R:S×A→R is the reward function. After taking a in s, the immediate reward is returned. Since the agent actually only observes local state information, the strategy π of each agent i i is the local observation information o i Mapping to actions; Module M2.3: The partial state of the system environment observed at each time step t is received as the state information s of the reinforcement learning agent i i In the lower-level strategy, this includes the queue length of vehicles on each lane of intersection i and the current signal phase of the intersection in Is entering the lane i The number of vehicles on phase i is the signal phase of intersection i; in the upper strategy, agent i will be in state s at the end of the macro action i As its macro information i ; Module M2.4: In the upper-level strategy, agent i performs a macro-action m at each i At the end, the next macro action needs to be decided. The macro action contains the signal phase and duration information. The optional signal duration includes {10s, 15s, 20s, 25s}, so there are 16 optional macro actions in total. The lower-level strategy of agent i executes action a according to the macro action information. i , that is, the signal phase specified by the macro action, at each time step t, the corresponding signal phase is executed, and each action a i Lasts for 5 seconds, and determines whether the macro action is completed based on the signal duration information; Module M2.5: At each time step t, agent i takes a macro action m = < τ m ,I m ,π m >, where τ is the time required for the joint macro action m to end, initialization set Depends on the macro action observation history information of agent i denotes the lower-level strategy used to implement macro-action m. Considering the termination condition of the macro-action, the transition probability is redefined as P(s′,τ|s,m), where τ is the time required for the joint macro-action m to end. The end of the joint macro-action means that each agent has completed its own macro-action; Z(z,m,s′)=P(z|m,s′) denotes the likelihood model of the joint macro-action. Module M2.6: The goal of the agent is to find a joint upper-level strategy Φ = × i Φ i , make decisions at the macro-action level so that the value of Φ starting from s0 is optimized, and the corresponding action value function is denote the expected cumulative rewards at a specific action and state, respectively; Module M2.7: In the lower-level policy, the reward r of agent i i Defined as the immediate feedback of the current action, that is, the negative value of the average number of vehicles entering the lane in Is entering the lane i The number of vehicles on the upper layer; in the upper layer strategy, agent i takes macro action m i r at the end i Considered as the reward information for this macro action.

9. The multi-intersection asynchronous traffic signal control system in a non-communication environment according to claim 6, characterized in that: The module M3 includes: Module M3.1: Use macro-action concurrent experience replay trajectories to filter macro-action experience data. The agent saves its action, observation, and reward information in the experience replay pool at each time step t. When sampling experience data, only the experience data at the end of the macro action is selected as valid data, and the rest of the data is discarded. Module M3.2: Observing history in a macro action i In the example, each agent independently selects a macro action m i , save reward information, At the end of the macro action, the agent obtains a new macro measurement s′ and a new macro action observation history information h′ = <h m ,m i ,s′>; Accordingly, the experience tuple collected by agent i is expressed as where s i It is used to select macro action m i Macro information; Module M3.3: Update the QSS network. In each training, Update in a way that minimizes the following TD error: Module M3.4: Update f i Network, in each training, in order to obtain the static ideal probability transfer function Use a neural network f(s,m i ) to predict s′ * , update f(s,m i ) to maximize the following loss function: The first item in square brackets requires that the next state be The second requirement is to limit the predicted next state to N(s,m i ) set, the hyperparameter λ is based on the transfer model f i (s,m i ) coefficient; Module M3.5: When training the f network, the output uses a technique similar to the residual network. The output is the difference between the predicted state and the input state Δ = s′ - s. Therefore, the predicted next state is s′ f =s+f(s,m); Module M3.6: Update Q i network, using QSS network and f i Network, Q i The update method is: The next state predicted by the second item is closer to the Q of the next state I2Q that appears frequently in the data. i The value is closer to the true value.

10. The multi-intersection asynchronous traffic signal control system in a non-communication environment according to claim 6, characterized in that: The module M4 includes: Module M4.1: Initializes the Q network, QSS network, and state transition model f of each agent, as well as the corresponding target Q network and target QSS network. Then, initializes the CityFlow traffic environment, loads road network data and traffic flow data, and starts simulation training. Module M4.2: The CityFlow simulator receives road network data and vehicle flow data as input, constructs a simulated traffic network based on this data, and automatically generates vehicles to enter the network during the simulation, driving them along predetermined routes to their destinations. At each moment, the simulator accurately simulates the behavior of each vehicle, providing detailed traffic flow information and effectively accelerating the simulation process through multi-threading technology. Module M4.3: The simulator obtains external control action information to control the signal phase and signal duration at the intersection. After each green light signal ends, the traffic light flashes yellow for 3 seconds, then the intersection remains in a full red state for 2 seconds before switching to the next traffic signal. Module M4.4: Each agent stores the experience data generated by its interaction with the environment into its own concurrent experience replay buffer In this process, each agent first obtains the local state information of the environment, and then determines whether the macro action is completed. If the macro action is completed, the upper-level strategy of the agent uses the ∈-greedy method to decide the macro action, and then determines the signal phase to be output according to the macro action. i Sum signal duration d i , the lower-level strategy resets the signal duration and outputs the signal phase to the environment i If the macro action is not completed, the lower-level strategy executes the lower-level action according to the signal phase specified by the unfinished macro action and updates the remaining signal duration. Module M4.5: At each time step t, the agent saves local state information, macro actions, lower-level actions, and reward information to the concurrent experience replay cache; Module M4.6: The model parameter update process of the agent. Each agent first samples data from its own experience replay buffer and filters out the experience data at the end of the macro action (s i ,m i ,s′ i ,r i ), and then update f respectively i , Q i , and update the target network every 5 training rounds

Citation Information

Patent Citations

  • Signal lamp control method, model training method, system and device and storage medium

    CN113643528A