Urban trunk road adaptive signal control method combining Transform and element reinforcement learning
By combining Transformer and meta-reinforcement learning, the historical state characteristics of the intersection are extracted, multi-objective reward functions are designed, and a two-layer meta-learning Bi-MAML framework is constructed, which solves the problem of difficult adaptability of traffic road coordination methods and multi-agent scalability in the existing technology, and realizes efficient signal control and training processes.
Patent Information
- Application Number
- CN202510223882.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing coordination methods for traffic roads are difficult to track traffic flow changes for adaptive control, and the multi-intersection training method cannot solve the problems of multi-agent expansion and dimensional explosion at the same time, and there is a lack of methods that take into account the benefits of coordination direction and overall benefits of intersections.
Transformer is used to extract the historical state characteristics of the intersection, design multi-objective reward functions, build a learning model for multi-intersection group training and dispersed decision making, propose a learning framework for bi-layer meta-learning Bi-MAML algorithm, and design a multi-agent PPO algorithm training process based on Bi-MAML framework.
A dynamic signal phase scheme is implemented to capture potential green wave requirements, reduce training costs, solve the problem of poor mobility between different intersections, and improve the efficiency and applicability of training.
Smart Images

Figure CN120126311A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic signal control, and in particular, to an adaptive signal control method for urban arterial roads combining Transformer and meta-reinforcement learning. Background Technique
[0002] Arterial green wave coordination is one of the important technical means in urban traffic signal control. By coordinating the signal timing plans between intersections, vehicles can pass through multiple intersections in the coordinated direction of the arterial road without stopping at a certain speed, thereby effectively reducing the number of vehicle stops, lowering fuel consumption and carbon emissions, and improving the traffic efficiency of arterial roads. Traditional arterial green wave control methods include the model method, the numerical solution method, and the graphical method, but they all belong to fixed signal timing plans. When sudden situations occur on the road section, resulting in instantaneous changes in traffic conditions, the coordination ability will be greatly weakened. In recent years, in order to overcome the limitations of traditional green wave control, intelligent traffic control methods based on deep reinforcement learning have gradually become a research hotspot, which can make intelligent decisions in complex environments and achieve adaptive coordinated control.
[0003] In the research of adaptive signal control, retaining the common cycle and phase difference of each coordinated intersection in the traditional green wave theory will make it difficult to synchronously adjust all intersection plans. Moreover, when independent agents are set at all intersections for decentralized training, as the road network scale increases, the training cost will become extremely high. To reduce the training time, some studies consider centralized training for all intersections. The first method is to adopt a joint state space and action space. However, as the number of agents increases, this method will have the problem of dimensional explosion. The second method is to let the agents at each intersection share training samples and network parameters, but it is only applicable to homogeneous intersection groups and not applicable to intersections with large differences in shape characteristics. In terms of traffic benefits, traditional green wave coordination methods only consider the benefits of arterial green waves, without considering the balance between arterial roads and branch roads and the overall benefits of intersections. Therefore, how to break the constraints of the common cycle and phase difference to synchronously update the signal plans of each coordinated intersection, capture the real-time vehicle green wave demand to adjust the plan, pay attention to the fairness of branch roads while pursuing the benefits of the coordinated direction, conduct grouped training and decentralized decision-making for intersections, and establish a scalable and migratory training architecture has important research value and practical significance. Summary of the Invention
[0004] The present invention provides an adaptive signal control method for urban arterial roads combining Transformer and meta-reinforcement learning.
[0005] In view of the deficiencies that the existing traffic artery coordination method cannot achieve adaptive control by tracking the changes in traffic flow, the multi-intersection training method adopted in the existing research cannot simultaneously solve the problems of multi-agent scalability and dimensional explosion, and the existing methods do not take into account the coordination direction benefit and the overall benefit of intersections, etc., this invention uses Transformer to extract historical state features, so as to capture the potential demand for vehicles waiting for green waves, designs a multi-objective reward function considering green wave efficiency and the overall traffic efficiency of intersections, constructs a learning mode of multi-intersection grouped training and decentralized decision-making, proposes a double-layer meta-learning Bi-MAML algorithm learning framework, and designs the training process of a multi-agent PPO algorithm based on the Bi-MAML framework.
[0006] This invention is realized through the following technical solutions:
[0007] An urban arterial adaptive signal control method combining Transformer and meta-reinforcement learning, comprising the following steps:
[0008] S1. Determine the coordination direction of the arterial road, usually the direction with large traffic flow;
[0009] S2. Group the intersections by category;
[0010] S2-1. Judge the shape types of each intersection, and classify the intersections by shape into pedestrian crossing intersections, three-way intersections (T-shaped intersections), four-way intersections (cross intersections), five-way intersections, etc.;
[0011] S2-2. Judge the number and positions of the adjacent intersections of each intersection, and divide the adjacent intersections of each intersection into two categories according to their positions, including adjacent intersections within the coordination range and adjacent intersections outside the coordination range;
[0012] If the adjacent intersection is an intersection on the arterial road coordination direction, it is an adjacent intersection within the coordination range; if the adjacent intersection is not an intersection on the arterial road coordination direction, it is an adjacent intersection outside the coordination range; by clarifying the number and position types of the adjacent intersections of each intersection, the intersections are further classified in detail;
[0013] S2-3. Determine the intersection grouping situation; classify the intersections according to the intersection shape type, the number and position types of the adjacent intersections, and only the intersections with the same intersection shape type, the number and position types of the adjacent intersections can be grouped into the same group;
[0014] S3. Determine the detection range of each intersection. The detection range of an intersection includes the distance from the stop line of the road section to the starting point of the road section in each import direction. For this detection length, there are a minimum value (the length of the solid line segment in the import direction) and a maximum value (the total length of the road section in the import direction); according to the result of grouping the intersections in step S2, set the same detection range l for the intersections in the same group. detect ;
[0015] S4. Define the vehicles waiting for the green wave; the vehicles that go straight through the green light in the coordination direction of the upstream intersection within the coordination range are defined as the vehicles waiting for the green wave when they reach the downstream intersection.
[0016] S5. Design a neural network structure based on Transformer that can extract the historical state features of intersections. The improved Transformer module designed in the present invention includes an encoder, a decoder, and a final output layer; both the encoder and the decoder are composed of n layer layers, and each layer adopts the same network structure; in each layer of the encoder, the data is input into the multi-head attention mechanism, and then undergoes a linear transformation through the feed-forward layer and is output; in each layer of the decoder, the data only needs to undergo a linear transformation through the feed-forward layer; after the data passes through the encoder and the decoder, average pooling is performed on the output sequence data to make the time step dimension become 1, and then it goes through another linear layer to compress the data dimension to d out ;
[0017] Specifically, for the multi-head attention mechanism of each encoder layer, the input data is a matrix X with a fixed dimension; each head is a self-attention mechanism, and the query matrix Q i , the key matrix K i and the value matrix V i are obtained respectively through three linear transformations. Define the weight matrices in the three linear transformations as and The result Z of the multi-head attention mechanism is obtained by concatenating the results Z i of all heads and then performing a linear transformation. The calculation formula is shown in Equation (1);
[0018] Z = Concat(Z 1 , Z 2 ,... Z h )W O (1)
[0019]
[0020] where h is the number of multi-heads; W O is the output linear transformation matrix; d k is the dimension of the key matrix;
[0021] S6. For each intersection, establish a single-agent Markov decision model;
[0022] S6-1. Design the action as selecting the next execution phase of the local intersection. The action set is represented as A = {a 1 , a 2 , ……, a k}, where k represents the number of selectable phases; the selected phase will execute a fixed interval of green light time Δt; within a signal cycle, the total green light time of a single phase is restricted by the minimum green light time t G_min and the maximum green light time t G_max ; when switching between phases, before executing the new phase, it is necessary to execute a yellow light time t Y for transition;
[0023] S6-2. Design the state to include the vehicle density matrix, vehicle speed matrix, historical action matrix, and historical state feature vector of the local intersection;
[0024] Among them, before recording the vehicle density matrix and vehicle speed matrix, divide the lanes of the local intersection into multiple lane groups according to the traffic flow direction; then divide each lane group into blocks, and the block length l cell is set to an integer multiple of the headway h s , and the block area cannot exceed the detection range l detect ; count the number of vehicles and the average vehicle speed in each lane block, and perform normalization processing;
[0025] The historical action matrix consists of the historical actions of adjacent intersections and the local intersection. Define the action vector as a one-hot vector of size k; specifically, assume that the historical data acquisition step length is τ. At time t, obtain the action vectors of the local intersection at times t - Δt·m, m = 0, 1, ……, τ respectively, and obtain the action vectors of adjacent intersections within the coordination range at m = 0, 1, ……, τ respectively, where t o→d represents the driving time required to drive from adjacent intersection I o to local intersection I d without stopping, represents the ceiling value after calculating t o→d / Δt; combine the obtained action vectors of the two adjacent intersections and the local intersection into a matrix. Among them, when there is no adjacent intersection, fill all the action vectors it represents with 0 values;
[0026] The historical state feature vector is obtained based on the historical vehicle density matrix in the coordinated direction of the local intersection and the historical vehicle density matrix in the coordinated direction of the adjacent intersection. Specifically, assuming that the historical data acquisition step size is τ, at time t, the vehicle density matrices in the coordinated direction of the local intersection at times t - Δt·m, m = 0, 1, ……, τ are acquired, and the vehicle density matrices in the coordinated direction of the adjacent intersections within the coordination range at m = 0, 1, ……, τ are acquired, where t o→d represents the travel time required to drive from the adjacent intersection I o to the local intersection I d without stopping, represents the ceiling value after calculating t o→d / Δt; the acquired historical vehicle density matrices in the coordinated direction of the local intersection and the historical vehicle density matrices in the coordinated direction of the adjacent intersection are concatenated together and input into the Transformer module for feature extraction, and the historical state feature vector is output;
[0027] Finally, the vehicle density matrix, the vehicle speed matrix, the historical action matrix, and the historical state feature vector are concatenated into the final input state vector;
[0028] S6-3. For each intersection, the reward function is calculated independently according to the traffic state of the intersection, and the calculation formula of the reward function is shown in Equation (3);
[0029]
[0030] where R t represents the reward value at time t; represents the waiting time reward value of the vehicles waiting for the green wave at time t, and the calculation formula is shown in Equation (4); represents the waiting time reward value of other vehicles at time t, and the calculation formula is shown in Equation (5); ρ is a coefficient used to adjust the ratio of the waiting time reward value of other vehicles to the total reward;
[0031]
[0032] where w m represents the cumulative waiting time of vehicle m, represents the number of vehicles waiting for the green wave at time t, represents the number of other non-green-wave waiting vehicles at time t;
[0033] S7. Construct a multi-agent collaboration structure with classified training and decentralized decision-making. Under this structure, intersections in the same group are trained centrally, and all intersections make decentralized decision-making actions. Specifically, agents among intersections in the same group share experience samples for iterative training and also share network parameters. However, when making decision-making actions, all intersections, based on the current state information they obtain, have agents independently decide their own actions.
[0034] S8. Construct a two-layer meta-learning Bi-MAML framework. Under this framework, set the intersections of a certain group as meta-learners and regard other agents as individual learners. Among them, the meta-learner includes all intersections within the corresponding group, and these intersections share the network parameters of the meta-learner. The first layer of meta-learning includes meta-training and fine-tuning. Among them, both meta-training and fine-tuning include inner-layer updates and outer-layer updates. The second layer of meta-learning only has outer-layer updates.
[0035] When the meta-learner has not reached the set number of meta-training rounds E during training meta perform meta-training, and this training is only for the meta-learner. Regard each intersection corresponding to the meta-learner as an independent task, randomly sample the tasks to obtain different batches of tasks. For each batch of tasks, obtain the corresponding samples and divide them into two categories, named support set samples and query set samples respectively. For each batch of tasks, perform inner-layer updates and outer-layer updates respectively. Copy the model parameters θ of the meta-learner to create a network copy, and the inner-layer update is performed on the network copy. For the first update, calculate the loss function L supp (θ) using the support set samples, and update the parameters of the network copy through formula (6) to obtain θ′. Subsequently, the network copy parameters θ′ will be iteratively updated multiple times on the support set;
[0036]
[0037] Among them, α is the learning rate of the inner-layer update, is the gradient of the parameter θ;
[0038] In the outer-layer update, calculate the loss function L query (θ′) on the network copy after the inner-layer update using the query set samples; calculate the mean of L query (θ′) calculated on the query set samples for each batch of tasks to obtain L(θ′), and use the gradient descent method to train and update the original network, and its parameter update formula is as shown in formula (7);
[0039]
[0040] Among them, β is the learning rate of the outer-layer update;
[0041] When the meta-learner reaches the set number of meta-training rounds E during training metaWhen fine-tuning is performed, for all agents, including the meta-learner and all individual learners, an inner update and an outer update are respectively carried out once.
[0042] When the meta-learner and all individual learners have both undergone fine-tuning once, the second layer of learning is carried out; for all agents, including the meta-learner and all individual learners, samples are respectively sampled from the experience pool without dividing the support set samples and the query set samples, and the gradient descent method is used to update the networks of each learner itself.
[0043] S9. Train the multi-agent PPO algorithm model based on the two-layer meta-learning Bi-MAML framework.
[0044] S9-1. Design the structures of the Actor network and the Critic network, and initialize the network parameters.
[0045] S9-2. Initialize the road network environment and the signal plan.
[0046] S9-3. Obtain the current state s of each intersection t , input it into their respective Actor networks, and then output the actions a they respectively select. t ;
[0047] S9-4. Control the operation of the signal controller according to the actions selected by each intersection.
[0048] S9-5. Each intersection obtains its own reward r t , input the state into the Critic network, and output the value function. Subsequently, store the state, action, value function, and reward in the form of a tuple in the experience pool of each intersection.
[0049] S9-6. Determine whether the number of samples reaches the sampling condition. If it reaches the sampling condition, proceed to step S9-7; if it does not reach the sampling condition, jump to step S9-13.
[0050] S9-7. Determine whether the number of completed training rounds reaches the set meta-training rounds E meta , if it has reached, jump to step S9-9; if it has not reached, proceed to step S9-8.
[0051] S9-8. Perform the first layer of meta-training. The specific steps are as follows:
[0052] S9-8-1. Regard each intersection as a task, and perform multi-batch random sampling on the intersections using the meta-learner. Each batch contains several intersections (i.e., several tasks).
[0053] S9-8-2. Extract samples from the experience pool of each intersection, and divide the samples into two parts, which are respectively used as the support set and the query set of this task.
[0054] S9-8-3. Perform the first-layer meta-training on the meta-learner, where during the outer update, the parameters of the Actor network and the Critic network of the meta-learner are updated;
[0055] S9-8-4. Copy the network parameters of the meta-learner to other agents, and then proceed to step S9-12;
[0056] S9-9. Determine whether fine-tuning has been performed. If fine-tuning has been performed, jump to step S9-11; if fine-tuning has not been performed yet, proceed to step S9-10;
[0057] S9-10. Perform the first-layer fine-tuning. The specific steps are as follows:
[0058] S9-10-1. Each experience pool separately extracts multiple batches of samples as the support set and the query set;
[0059] S9-10-2. Each agent performs the first-layer fine-tuning based on its own samples, and updates the parameters of its own Actor network and Critic network respectively through the outer update, and then proceeds to step S9-12;
[0060] S9-11. Perform the second-layer learning. The specific steps are as follows:
[0061] S9-11-1. Each agent separately extracts multiple batches of samples from its own experience pool, and calculates the loss functions of the Actor network and the Critic network according to equations (8) and (9) respectively;
[0062]
[0063] where, θ 1 represents the Actor network parameters; θ 2 represents the Critic network parameters; p t (θ 1 ) is the ratio of the probabilities of selecting action a t under state s according to the current policy and the old policy t respectively; clip(p t (θ 1 ), 1 - ε, 1 + ε) is a function that limits p t (θ 1 ) within the range [1 - ε, 1 + ε]; represents the current value function under state s t ; V t target represents the target value function; A trepresents the advantage function; γ is the reward discount factor; T is the termination time; represents the estimated value of the old value network for state s T ;
[0064] S9-11-2. Each agent updates the parameters of its own Actor network and Critic network respectively according to the calculated loss function, and then proceeds to step S9-12;
[0065] S9-12. Clear the experience pool and proceed to step S9-13;
[0066] S9-13. Determine whether the current round has reached the termination time. If the termination time has been reached, proceed to step S9-14; if the termination time has not been reached, return to step S9-3;
[0067] S9-14. Determine whether all round iterations have been completed. If all round iterations have been completed, end the training; if not all round iterations have been completed, return to step S9-2;
[0068] S10. Use the trained agent models to achieve adaptive signal control for urban arterial roads.
[0069] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0070] (1) By using Transformer to extract the historical state features of local intersections and adjacent intersections, capture potential green wave demands, and dynamically adjust the signal phase plan.
[0071] (2) Group the intersection clusters according to different characteristics such as intersection shape, size, lane organization, etc., and use the method of sharing samples and network parameters for intersections in the same group to reduce the training cost. At the same time, it avoids the problem of low adaptability and weak pertinence caused by using the same network parameters for non-homogeneous intersections.
[0072] (3) The proposed double-layer meta-learning Bi-MAML framework can solve the problem of poor transferability between different intersections in a large-scale road network. By training the meta-learner separately in the first layer and training each agent independently in the second layer, the convergence speed is accelerated in the early stage of training, and the applicability of each agent is strengthened in the later stage of training.
[0073] The present invention has the following advantages and effects compared with the prior art:
[0074] (1) By using Transformer to perform feature extraction and dimension compression on the historical state information of local intersections and adjacent intersections, potential green wave demands can be captured, providing forward-looking state information for agents.
[0075] (2) Group the intersection clusters according to the shape characteristics of the intersections. For intersections in the same group, share the sample and network parameters to reduce the training cost. At the same time, it avoids the problem of low adaptability and poor pertinence caused by using the same network parameters for non-homogeneous intersections.
[0076] (3) The proposed double-layer meta-learning Bi-MAML framework can solve the problem of poor transferability between different intersections in a large-scale road network. By training the meta-learner separately in the first layer and training each agent independently in the second layer, the convergence speed is accelerated in the early stage of training, and the applicability of each agent is strengthened in the later stage of training. Brief Description of the Drawings
[0077] Figure 1 is the flowchart of the urban arterial road adaptive signal control method combining Transformer and meta-reinforcement learning in the embodiment of the present invention;
[0078] Figure 2 is the schematic diagram of the arterial road in the embodiment of the present invention;
[0079] Figure 3 is the schematic diagram of the selectable signal phase in the embodiment of the present invention;
[0080] Figure 4 is the schematic diagram of the intersection lane grouping and partitioning in the embodiment of the present invention;
[0081] Figure 5 is the schematic diagram of the agent classification in the embodiment of the present invention;
[0082] Figure 6 is the meta-training flowchart in the double-layer meta-learning Bi-MAML of the embodiment of the present invention;
[0083] Figure 7 is the fine-tuning flowchart in the double-layer meta-learning Bi-MAML of the embodiment of the present invention;
[0084] Figure 8 is the training flowchart of the multi-agent PPO algorithm based on the double-layer meta-learning Bi-MAML framework in the embodiment of the present invention. Detailed Embodiment
[0085] Next, the solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0086] Please refer to Figure 1 , this embodiment provides an urban arterial road adaptive signal control method combining Transformer and meta-reinforcement learning, including the following steps:
[0087] S1. Determine the coordinated direction of the arterial road, usually the direction with large traffic flow; in one embodiment, the east-west direction is used as the coordinated direction;
[0088] S2. Group the intersections by category;
[0089] S2-1. Determine the shape type of each intersection, and classify the intersections by shape into pedestrian crossing intersections, three-way intersections (T-shaped intersections), four-way intersections (cross intersections), five-way intersections, etc.;
[0090] S2-2. Determine the number and location of the adjacent intersections of each intersection, and classify the adjacent intersections of each intersection into two categories according to location, including adjacent intersections within the coordination range and adjacent intersections outside the coordination range;
[0091] If the adjacent intersection is an intersection on the main road coordination direction, it is an adjacent intersection within the coordination range; if the adjacent intersection does not belong to the intersection on the main road coordination direction, it is an adjacent intersection outside the coordination range; by clarifying the number and location type of the adjacent intersections of each intersection, the intersections are further classified in detail;
[0092] S2-3. Determine the intersection grouping situation; classify the intersections according to the intersection shape type, the number and location type of the adjacent intersections, and only the intersections with the same intersection shape type, the number and location type of the adjacent intersections can be grouped into the same group;
[0093] Please refer to Figure 2 an example in 1 For example, taking a main road with eight intersections as an example, intersection I 1 as the end T-shaped intersection belongs to agent 2 and I 3 as the middle T-shaped intersection belongs to agent 2 and intersection I 4 、I 6 and I 7 as the middle cross intersection belongs to agent 3 and intersection I 5 as the pedestrian crosswalk belongs to agent 4 and intersection I 8 as the end cross intersection belongs to agent 5 ;
[0094] S3. Determine the detection range of each intersection. The detection range of the intersection includes the distance from the stop line of the section to the starting point of the section in each import direction. There are a minimum value (the length of the solid line segment in the import direction) and a maximum value (the total length of the section in the import direction) for this detection length; according to the result of grouping the intersections in step S2, set the same detection range l for the intersections in the same group detect; In one embodiment, the same detection range l is set for all agents (including intersections in different groups). detect = 100 m;
[0095] S4. Define the vehicles waiting for green wave; The vehicles that go straight through the green light in the coordinated direction of the upstream intersection within the coordinated range are defined as the vehicles waiting for green wave when they reach the downstream intersection; Please refer to Figure 2 for an example. The vehicle goes straight through the west entrance of intersection I 1 and arrives at intersection I 2 is judged as a vehicle waiting for green wave. On the contrary, the vehicle turns right through the south entrance of intersection I 1 and arrives at intersection I 2 is judged as a vehicle not waiting for green wave;
[0096] S5. Design a neural network structure based on Transformer that can extract the historical state features of intersections. The improved Transformer module designed in the present invention includes an encoder, a decoder, and a final output layer; Both the encoder and the decoder are composed of n layer layers, and each layer adopts the same network structure; In each layer of the encoder, the data is input into the multi-head attention mechanism, and then output after linear transformation through the feed-forward layer; In each layer of the decoder, the data only needs to go through the feed-forward layer for linear transformation; After the data passes through the encoder and the decoder, average pooling is performed on the output sequence data to make the time step dimension become 1, and then it goes through another linear layer to compress the data dimension to d out ;
[0097] Specifically, for the multi-head attention mechanism of each encoder layer, the input data is a matrix X with a fixed dimension; Each head is a self-attention mechanism, and the query matrix Q i , key matrix K i , and value matrix V i are obtained respectively through three linear transformations. Define the weight matrices in the three linear transformations as and The result Z of the multi-head attention mechanism is obtained by concatenating the results Z i of all heads and then through linear transformation. The calculation formula is shown in Equation (1);
[0098] Z = Concat(Z 1 , Z 2 ,... Z h )W O (1)
[0099]
[0100] where h is the number of multi-heads; WO is the linear transformation matrix of the output; d k is the dimension of the key matrix;
[0101] In one embodiment, set n layer = 6, that is, the data needs to pass through 6 layers of encoders and decoders; set d out = 16, that is, the size of the vector finally output from the Transformer module is 16; set h = 8, that is, the self-attention mechanism with 8 heads is adopted; set d k = d v = 100, that is, the dimensions of both the key matrix and the value matrix are 100;
[0102] S6. For each intersection, establish a single-agent Markov decision model;
[0103] S6-1. Design the action as selecting the next execution phase of the local intersection, and the action set is represented as A = {a 1 , a 2 , ……, a k}, where k represents the number of selectable phases; the selected phase will execute a green light time Δt of a fixed interval; within a signal cycle, the total green light time of a single phase is restricted by the minimum green light time t G_min and the maximum green light time t G_max ; when switching between phases, before executing the new phase, a yellow light time t Y needs to be executed for transition; please refer to Figure 3 for an example, the total number of selectable phases is k = 8, and each action represents a selectable phase. For different intersections, the selectable phases are different. When making a decision on the action, invalid phases need to be masked, but the action space size of all agents still needs to be kept consistent; in one embodiment, set the fixed-interval green light time Δt = 5s, the minimum green light time t G_min = 15s, the maximum green light time t G_max = 60s, and the yellow light time t Y = 5s;
[0104] S6-2. Design the state to include the vehicle density matrix, vehicle speed matrix, historical action matrix, and historical state feature vector of the local intersection;
[0105] Among them, before recording the vehicle density matrix and vehicle speed matrix, the lanes of the local intersection are divided into multiple lane groups according to the traffic flow direction; then each lane group is divided into blocks, and the block length l cell is set to an integer multiple of the headway h s , and the divided area cannot exceed the detection range l detect; Count the number of vehicles and the average vehicle speed in each lane block and perform normalization; see Figure 4 For an example in Figure 4 , the lanes are divided into 8 lane groups according to the traffic flow direction, such as east import left turn, east import straight, west import left turn, etc. The figure shows an example of the lane group division of west import left turn and west import straight. When dividing the blocks, set l cell = h s = 7.5 m;
[0106] The historical action matrix consists of the historical actions of adjacent intersections and the local intersection. Define the action vector as a one-hot vector of size k; specifically, assume that the historical data acquisition step size is τ. At time t, obtain the action vectors of the local intersection at times t - Δt·m, m = 0, 1, ……, τ, and obtain the action vectors of the adjacent intersections within the coordination range at m = 0, 1, ……, τ, where t o→d represents the driving time required to drive from adjacent intersection I o to local intersection I d without stopping, represents rounding up after calculating t o→d / Δt; Combine the obtained action vectors of the two adjacent intersections and the local intersection into a matrix. When there is no adjacent intersection, fill all the action vectors it represents with 0 values;
[0107] The historical state feature vector is obtained from the historical vehicle density matrix in the coordination direction of the local intersection and the historical vehicle density matrix in the coordination direction of the adjacent intersections; specifically, assume that the historical data acquisition step size is τ. At time t, obtain the vehicle density matrices in the coordination direction of the local intersection at times t - Δt·m, m = 0, 1, ……, τ, and obtain the vehicle density matrices in the coordination direction of the adjacent intersections within the coordination range at m = 0, 1, ……, τ, where t o→d represents the driving time required to drive from adjacent intersection I o to local intersection I d without stopping, represents rounding up after calculating t o→d / Δt; Concatenate the obtained historical vehicle density matrices in the coordination direction of the local intersection and the historical vehicle density matrices in the coordination direction of the adjacent intersections, input them into the Transformer module for feature extraction, and output the historical state feature vector;
[0108] See Figure 2 For an example in Figure 2 , set τ = 5 and Δt = 5 s, with intersection I 1 as the local intersection and the upstream intersection I2 to the local intersection I 1 The required travel time t 2→1 = 28 s, then Obtain I 1 The historical action vectors and the historical state matrices in the east - west direction at the state update times of t, t - 5, t - 10, t - 15, and t - 20 respectively, I 2 The historical action training and the historical state matrices in the east - west direction at the state update times of t - 30, t - 35, t - 40, t - 45, and t - 50 respectively; Since I 1 There is only one adjacent intersection, fill the relevant data of other non - existent adjacent intersections with 0 values; Concatenate all historical action vectors into a matrix, and concatenate all historical state matrices into a new matrix and input it into the Transformer to be transformed into a historical state feature vector with a dimension of d out = 16;
[0109] Finally, concatenate the vehicle density matrix, the vehicle speed matrix, the historical action matrix, and the historical state feature vector into the final input state vector;
[0110] S6 - 3. For each intersection, the reward function is calculated independently according to the traffic state of the intersection, and the calculation formula of its reward function is shown in Equation (3);
[0111]
[0112] Among them, R t represents the reward value at time t; represents the waiting time reward value of the vehicles waiting for the green wave at time t, and the calculation formula is shown in Equation (4); represents the waiting time reward value of other vehicles at time t, and the calculation formula is shown in Equation (5); ρ is a coefficient used to adjust the ratio of the waiting time reward value of other vehicles in the total reward;
[0113]
[0114] Among them, w m represents the cumulative waiting time of vehicle m, represents the number of vehicles waiting for the green wave at time t, represents the number of other non - green - wave - waiting vehicles at time t; In one embodiment, ρ = 0.9 is set;
[0115] S7. Construct a multi-agent collaboration structure with classified training and decentralized decision-making. Under this structure, intersections in the same group are trained centrally, and all intersections make decentralized decision-making actions. Specifically, agents among intersections in the same group share experience samples for iterative training and also share network parameters. However, when making decision-making actions, all intersections, based on the current state information they obtain, have agents independently decide their own actions.
[0116] S8. Construct a two-layer meta-learning Bi-MAML framework. Under this framework, set the intersections of a certain group as meta-learners and regard other agents as individual learners. Among them, the meta-learners include all intersections within the corresponding group, and these intersections share the network parameters of the meta-learners. The first layer of meta-learning includes meta-training and fine-tuning. Among them, both meta-training and fine-tuning include inner-layer updates and outer-layer updates. The second layer of meta-learning only has outer-layer updates. Please refer to Figure 5 for an example, and set Agent 1 corresponding intersections I 4 , I 6 and I 7 as meta-learners, and share samples and network parameters among them; set other agents as individual learners; set the number of meta-training rounds E meta = 50;
[0117] When the meta-learner has not reached the set number of meta-training rounds E meta during training, conduct meta-training, and this training is only for the meta-learner; please refer to Figure 6 , which shows the meta-training process; regard each intersection corresponding to the meta-learner as an independent task, randomly sample the tasks to obtain different batches of tasks; for each batch of tasks, obtain the corresponding samples and divide them into two categories, named support set samples and query set samples respectively; for each batch of tasks, conduct inner-layer updates and outer-layer updates respectively; copy the model parameters θ of the meta-learner to create a network copy, and the inner-layer update is performed on the network copy. For the first update, calculate the loss function L supp (θ) using the support set samples, and update the parameters of the network copy through formula (6) to obtain θ′. Subsequently, the network copy parameters θ′ will be iteratively updated multiple times on the support set;
[0118]
[0119] Among them, α is the learning rate of the inner-layer update, is the gradient of the parameter θ; in one embodiment, set α = 0.001;
[0120] In the outer-layer update, calculate the loss function L query (θ′) on the network copy after the inner-layer update using the query set samples; calculate the L for each batch of tasks on the query set samplesquery The mean of L(θ′) is obtained from L(θ′), and the original network is trained and updated using the gradient descent method. The parameter update formula is shown in Equation (7);
[0121]
[0122] where β is the learning rate for the outer layer update; in one embodiment, β is set to 0.0001;
[0123] When the meta-learner reaches the set number of meta-training epochs E meta during training, fine-tuning is performed; see Figure 7 , which shows the fine-tuning process; during fine-tuning, for all agents, including the meta-learner and all individual learners, an inner layer update and an outer layer update are each performed;
[0124] When the meta-learner and all individual learners have each undergone one round of fine-tuning, second-layer learning is performed; for all agents, including the meta-learner and all individual learners, samples are each drawn from the experience pool without dividing the support set samples and query set samples, and the networks of each learner are updated using the gradient descent method;
[0125] S9. Train a multi-agent PPO algorithm model based on the two-layer meta-learning Bi-MAML framework; see Figure 8 , and the training process is as follows:
[0126] S9-1. Design the structures of the Actor network and the Critic network, and initialize the network parameters;
[0127] S9-2. Initialize the road network environment and the signal plan;
[0128] S9-3. Obtain the current state s t of each intersection, input it into its respective Actor network, and then output the action a t selected by each;
[0129] S9-4. Control the operation of the signal controller according to the actions selected by each intersection;
[0130] S9-5. Each intersection obtains its own reward r t , inputs the state into the Critic network, and outputs the value function. Subsequently, the state, action, value function, and reward are stored in the experience pool of each intersection in the form of a tuple;
[0131] S9-6. Determine whether the number of samples reaches the sampling condition. If it reaches the sampling condition, proceed to step S9-7; if it does not reach the sampling condition, jump to step S9-13;
[0132] S9-7. Determine whether the number of completed training rounds has reached the set meta-training round E meta , if it has reached, jump to step S9-9; if not, proceed to step S9-8;
[0133] S9-8. Conduct the first-layer meta-training, and the specific steps are as follows:
[0134] S9-8-1. Treat each intersection as a task, and conduct multi-batch random sampling on the intersections using the meta-learner. Each batch contains several intersections (i.e., several tasks);
[0135] S9-8-2. Extract samples from the experience pool of each intersection, and divide the samples into two parts, which are used as the support set and query set for this task respectively;
[0136] S9-8-3. Conduct the first-layer meta-training on the meta-learner. Among them, when updating the outer layer, update the parameters of the Actor network and Critic network of the meta-learner;
[0137] S9-8-4. Copy the network parameters of the meta-learner to other agents, and then proceed to step S9-12;
[0138] S9-9. Determine whether fine-tuning has been performed. If fine-tuning has been performed, jump to step S9-11; if fine-tuning has not been performed yet, proceed to step S9-10;
[0139] S9-10. Conduct the first-layer fine-tuning, and the specific steps are as follows:
[0140] S9-10-1. Extract multi-batch samples from each experience pool as the support set and query set respectively;
[0141] S9-10-2. Each agent conducts the first-layer fine-tuning according to its own samples, and updates the parameters of its own Actor network and Critic network respectively through outer-layer update, and then proceeds to step S9-12;
[0142] S9-11. Conduct the second-layer learning, and the specific steps are as follows:
[0143] S9-11-1. Each agent extracts multi-batch samples from its own experience pool, and calculates the loss functions of the Actor network and Critic network respectively according to formulas (8) and (9);
[0144]
[0145] Among them, θ 1 represents the Actor network parameters; θ 2 represents the Critic network parameters; p t (θ1 ) is the ratio of the probabilities of selecting action a t under state s according to the current policy and the old policy respectively; clip(p t (θ t ), 1 - ε, 1 + ε) is a function that limits p 1 (θ t ) within the range [1 - ε, 1 + ε]; 1 ) represents the current value function of state s ; V t t target represents the target value function; A t T represents the advantage function; γ is the reward discount factor; T is the termination time; represents the estimated value of the old value network for state s T ; In one embodiment, ε = 0.2, γ = 0.98, and T = 3600s are set;
[0146] S9 - 11 - 2. Each agent updates the parameters of its own Actor network and Critic network respectively according to the calculated loss function, and then enters step S9 - 12;
[0147] S9 - 12. Clear the experience pool and enter step S9 - 13;
[0148] S9 - 13. Determine whether the current round has reached the termination time. If it has reached the termination time, enter step S9 - 14; if it has not reached the termination time, return to step S9 - 3;
[0149] S9 - 14. Determine whether all rounds of iteration have been completed. If all rounds of iteration have been completed, end the training; if not all rounds of iteration have been completed, return to step S9 - 2;
[0150] S10. Use the trained agent models of each agent to implement adaptive signal control for urban arterial roads.
[0151] When facing the time - varying traffic flow in the real traffic environment, vehicle position, vehicle speed, and signal light status and other information are obtained through Internet of Things technology, and this information is converted into state data and input to the trained agents of each agent and the decision actions are output to implement adaptive signal control for urban arterial roads.
[0152] As described above, the present invention can be preferably implemented.
[0153] The above embodiments are the preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An adaptive signal control method for urban arterial roads combining Transformer and meta-reinforcement learning, characterized in that: The steps include: S1. Determine the coordination direction of the main roads; S2, grouping intersections by category; S3, determining the detection range of each intersection, the detection range of the intersection includes the distance from the stop line of the section to the starting point of the section in each entrance direction; S4. Define green wave waiting vehicles: vehicles that go straight with a green light in the coordination direction from the upstream intersection within the coordination range are defined as green wave waiting vehicles when they arrive at the downstream intersection; S5. Design a neural network structure based on Transformer that can extract the historical state characteristics of the intersection. The designed improved Transformer module includes an encoder, a decoder and a final output layer; S6. For each intersection, a single-agent Markov decision model is established; S7, construct a multi-agent collaborative structure for classification training and decentralized decision-making, under which intersections of the same group are centrally trained and all intersections are decentralized decision-making actions; S8. Construct a two-layer meta-learning Bi-MAML framework, in which a group of intersections is set as a meta-learner and other agents are regarded as individual learners; the first layer of meta-learning includes meta-training and fine-tuning, and both meta-training and fine-tuning include inner layer updates and outer layer updates; the second layer of meta-learning only has outer layer updates; S9, training the multi-agent PPO algorithm model based on the bi-layer meta-learning Bi-MAML framework; S10. Utilize the trained intelligent agent models to realize adaptive signal control on urban main roads.
2. The method for adaptive signal control of urban trunk roads combining Transformer and meta-reinforcement learning according to claim 1, characterized in that: In step S2, grouping intersections by category specifically includes the following sub-steps: S2-1. Determine the shape type of each intersection, and classify the intersection into pedestrian crossing intersection, three-way intersection, four-way intersection, and five-way intersection according to the shape; S2-2, determining the number and positions of adjacent intersections of each intersection, and dividing the adjacent intersections of each intersection into two categories according to their positions, including adjacent intersections within the coordination range and adjacent intersections outside the coordination range; S2-3. Determine the intersection grouping; classify the intersections according to the intersection shape type, the number of adjacent intersections and the location type. Only intersections with the same intersection shape type, the number of adjacent intersections and the location type can be classified into the same group.
3. The method for adaptive signal control of urban trunk roads combining Transformer and meta-reinforcement learning according to claim 1, characterized in that: In step S5, both the encoder and the decoder are composed of n layer The network structure is composed of 2 layers, each layer uses the same network structure; in each layer of the encoder, the data is input into the multi-head attention mechanism, and then it is output after being linearly transformed by the feedforward layer; in each layer of the decoder, the data only needs to be linearly transformed by the feedforward layer; after the data passes through the encoder and decoder, the output sequence data is mean pooled so that the time step dimension is normalized to 1, and then it passes through another linear layer to compress the data dimension to d out .
4. The method for adaptive signal control of urban trunk roads combining Transformer and meta-reinforcement learning according to claim 1, characterized in that: In step S6, establishing a single-agent Markov decision model specifically includes the following sub-steps: S6-1, the design action is to select the next execution phase of the local intersection; S6-2, the design state includes the vehicle density matrix, vehicle speed matrix, historical action matrix and historical state feature vector of the local intersection; S6-3. For each intersection, the reward function is independently calculated according to the traffic status of the intersection. The calculation formula of the reward function is shown in formula (1); Among them, R t represents the reward value at time t; It represents the waiting time reward value of the green wave waiting vehicle at time t, and the calculation formula is shown in formula (2); represents the waiting time reward value of other vehicles at time t, and the calculation formula is shown in formula (3); ρ is a coefficient used to adjust the ratio of the waiting time reward value of other vehicles to the total reward; Among them, w m represents the cumulative waiting time of vehicle m, represents the number of vehicles waiting for the green wave at time t, Indicates the number of other non-green wave waiting vehicles at time t.
5. The method for adaptive signal control of urban trunk roads combining Transformer and meta-reinforcement learning according to claim 4, characterized in that: In step S6-2, the design state includes the vehicle density matrix, vehicle speed matrix, historical action matrix and historical state feature vector of the local intersection, specifically: Before recording the vehicle density matrix and the vehicle speed matrix, the lanes of the local intersection are divided into multiple lane groups according to the direction of traffic flow; then each lane group is divided into blocks with a block length of l cell Set as the headway h s The block area cannot exceed the detection range l detect ; Count the number of vehicles and average speed in each lane block and perform normalization; The historical action matrix consists of the historical actions of adjacent intersections and local intersections, and the action vector is defined as a one-hot vector of size k; The historical state feature vector is obtained based on the historical vehicle density matrix of the local intersection coordination direction and the historical vehicle density matrix of the adjacent intersection coordination direction; Finally, the vehicle density matrix, vehicle speed matrix, historical action matrix and historical state feature vector are concatenated into the final input state vector.
6. The method for adaptive signal control of urban arterial roads combining Transformer and meta-reinforcement learning according to claim 1, characterized in that: Step S8: When the meta-learner has not reached the set meta-training round E in training meta When , meta-training is performed, and this training is only for the meta-learner; each intersection corresponding to the meta-learner is regarded as an independent task, and the tasks are randomly sampled to obtain different batches of tasks; for each batch of tasks, the corresponding samples are obtained and divided into two categories, named support set samples and query set samples respectively; for each batch of tasks, inner layer updates and outer layer updates are performed respectively; the model parameters θ of the meta-learner are copied to create a network copy, and the inner layer updates are performed on the network copy. For the first update, the support set samples are used to calculate the loss function L supp (θ), and update the parameters of the network copy to obtain θ′ through formula (4), and then the network copy parameters θ′ will be iteratively updated multiple times on the support set; Among them, α is the learning rate of the inner layer update, is the gradient of the parameter θ; In the outer layer update, the query set samples are used to calculate the loss function L on the updated network copy in the inner layer. query (θ′); L calculated on the query set samples for each batch task query (θ′) is averaged to obtain L(θ′), and the original network is trained and updated using the gradient descent method. The parameter update formula is shown in formula (5); Among them, β is the learning rate of the outer layer update; When the meta-learner reaches the set meta-training round E in training meta When , fine-tuning is performed; in fine-tuning, all agents, including meta-learners and all individual learners, need to perform an inner layer update and an outer layer update respectively; After the meta-learner and all individual learners have been fine-tuned, the second layer of learning is carried out; all agents, including the meta-learner and all individual learners, are sampled from the experience pool without dividing the support set samples and the query set samples, and the gradient descent method is used to update the network of each learner.
7. The method for adaptive signal control of urban arterial roads combining Transformer and meta-reinforcement learning according to claim 1, characterized in that: In step S9, training the multi-agent PPO algorithm model based on the bi-layer meta-learning Bi-MAML framework specifically includes the following sub-steps: S9-1. Design the Actor network and Critic network structures, and initialize the network parameters; S9-2, initialize the road network environment and signal scheme; S9-3. Get the current status s of each intersection t , input to their respective Actor networks, and then output their respective selected actions a t ; S9-4, control the operation of the signal according to the action selected at each intersection; S9-5. Each intersection obtains its own reward r t , input the state to the Critic network and output the value function, then store the state, action, value function and reward in the form of tuples in the experience pool of each intersection; S9-6, determine whether the number of samples meets the sampling conditions, if so, proceed to step S9-7; if not, jump to step S9-13; S9-7, determine whether the number of completed training rounds has reached the set meta-training round E meta If it has been reached, jump to step S9-9; if it has not been reached, go to step S9-8; S9-8, perform first layer meta-training S9-9, determine whether fine adjustment has been performed, if fine adjustment has been performed, jump to step S9-11; if fine adjustment has not been performed, go to step S9-10; S9-10, perform first layer fine-tuning; S9-11, conduct second level learning; S9-12, clear the experience pool and proceed to step S9-13; S9-13, determine whether the current round has reached the end time, if it has reached the end time, proceed to step S9-14; if it has not reached the end time, return to step S9-3; S9-14, determine whether all round iterations are completed. If all round iterations are completed, end the training; if not, return to step S9-2.
8. The method for adaptive signal control of urban arterial roads combining Transformer and meta-reinforcement learning according to claim 7, characterized in that: In step S9-8, the first layer element training specifically includes the following sub-steps: S9-8-1. Treat each intersection as a task and randomly sample multiple batches of intersections using the meta-learner, with each batch containing several intersections. S9-8-2. Extract samples from the experience pool of each intersection and divide the samples into two parts, which serve as the support set and query set of the task respectively; S9-8-3, perform the first layer of meta-training on the meta-learner, wherein the outer layer updates the parameters of the Actor network and Critic network of the meta-learner; S9-8-4. Copy the network parameters of the meta-learner to other agents, and then proceed to step S9-12.
9. The method for adaptive signal control of urban trunk roads combining Transformer and meta-reinforcement learning according to claim 7, characterized in that: In step S9-10, the first layer fine-tuning specifically includes the following sub-steps: S9-10-1. Multiple batches of samples are extracted from each experience pool as support sets and query sets; S9-10-2. Each agent performs first-layer fine-tuning based on its own samples, and updates its own Actor network and Critic network parameters through outer layer updates, and then enters step S9-12.
10. The method for adaptive signal control of urban arterial roads combining Transformer and meta-reinforcement learning according to claim 7, characterized in that: In step S9-11, the second layer learning specifically includes the following sub-steps: S9-11-1. Each agent extracts multiple batches of samples from its own experience pool, and calculates the loss functions of the Actor network and the Critic network according to equations (6) and (7) respectively; Among them, θ1 represents the Actor network parameters; θ2 represents the Critic network parameters; p t (θ1) is in state s t Under the current strategy With the old strategy Select action a respectively t The ratio of the probability of t (θ1),1-ε,1+ε) is a t (θ1) is a function limited to the range [1-ε,1+ε]; Indicates the state s t The current value function of V t target represents the target value function; A t represents the advantage function; γ is the reward reduction coefficient; T is the termination time; Represents the old value network for state s T An estimated value of S9-11-2. Each agent updates its own Actor network and Critic network parameters according to the calculated loss function, and then enters step S9-12.
Citation Information
Patent Citations
Urban area traffic signal control method based on short-time traffic flow prediction
CN113327416A
Large-scale traffic light signal control method based on primitive learning and deep reinforcement learning
CN116137103A
System and method for task control based on bayesian meta-reinforcement learning
US20220180744A1
Method of performing green wave coordination control, electronic device and storage medium
US20230334986A1
Cited By
Traffic signal cooperative control method and system based on multi-scale space attention
CN121545371A
Intersection queuing rapid evacuation control method based on meta reinforcement learning
CN121661822A