Traffic signal lamp induction control method and system based on deep reinforcement learning
By modeling traffic light control as a Markov decision-making process and using deep reinforcement learning transformer network model, real-time response and versatility of existing traffic light control methods are solved, and more efficient traffic flow management is achieved.
Patent Information
- Application Number
- CN202510493049.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing traffic light control methods cannot respond to sudden traffic flow changes in real time, lack global traffic flow coordination, high sensor dependence, difficulty in adapting to complex scenarios, and low versatility and optimization dimensions.
The traffic light control problem is modeled as a Markov decision-making process, the observation space, action space and reward functions are defined, and the transformer network model based on deep reinforcement learning is used for optimization. The optimal action strategy is obtained through discrete timing trajectory data training, considering phase saturation, green wave coordination, emergency vehicle priority and carbon emission punishment.
It improves the effect and versatility of traffic light control, can dynamically adapt to different intersection topology, reduce the limitations of traditional methods, and achieve more efficient traffic flow management.
Smart Images

Figure CN120279738A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic Internet of Things control, and in particular to a traffic signal induction control method and system based on deep reinforcement learning. Background Art
[0002] Traffic signals are important facilities for road traffic management, and indicate the passage or stop of vehicles and pedestrians through the color and combination of lights to ensure traffic safety and order.
[0003] Urban traffic congestion has long been a global problem, having a negative impact on both the economy and the environment. The energy consumed during traffic congestion leads to the emission of greenhouse gases such as carbon dioxide, thus exacerbating the greenhouse effect. There is an urgent need to improve the environment through technological innovation to reduce the pollution emissions caused by traffic congestion.
[0004] Currently, the control of urban traffic signals mainly relies on three methods: one is fixed-time control, the principle of which is to preset the signal cycle, phase duration, and phase sequence according to historical traffic flow statistics data. For example, longer green light times are set for the main roads during the morning and evening rush hours. However, this method cannot respond to sudden traffic flow changes in real time, such as accidents and a sharp increase in traffic flow during holidays, and requires regular data collection to adjust the timing, resulting in high maintenance costs. In addition, the same scheme is difficult to adapt to different intersection topologies, such as T-shaped intersections and crossroads. The second is inductive control, which uses sensors such as geomagnetic coils and cameras to detect the presence of vehicles in the lanes in real time, and dynamically extends the current green light phase or switches phases. For example, when no vehicles are detected in a certain direction, its traffic light is skipped. However, this method only focuses on the vehicles in the current phase, lacking global traffic flow coordination. The sensors need to accumulate a certain amount of time after detecting a vehicle before triggering an action, which is likely to form queues during peak hours, and sensor failures are likely to cause control failures. The third is rule-based adaptive control, such as SCOOT / SCATS, which adjusts the signal cycle through a preset rule library. For example, the SCATS system dynamically allocates green light time according to the saturation. However, this method relies on limited rules summarized by expert experience, making it difficult to cover complex scenarios, such as multi-modal traffic flow mixtures, and traditional algorithms are also unable to process high-dimensional data, such as the vehicle trajectories of the entire road network, with low optimization dimensions and rule updates requiring manual intervention.
[0005] The inventors of the present application found during the implementation of the present invention that the above-mentioned solutions of the prior art have the defects of poor traffic signal control effect and poor versatility. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a traffic signal induction control method and system based on deep reinforcement learning, which has the functions of good traffic signal control effect and good versatility.
[0007] To achieve the above object, on the one hand, an embodiment of the present invention provides a traffic signal sensing control method based on deep reinforcement learning, including: Model the traffic signal control problem at the intersection as a Markov decision process, and define the observation space, action space, and reward function, where the reward function includes: Obtain the reward function according to formula (1), , (1) where, is the reward function, is the th phase efficiency factor, is the th basic reward item, is an integer number, and , is the green wave weight, is the green wave coordination reward item, is the emergency weight, is the emergency vehicle priority item, is the emission weight, is the carbon emission penalty item; Obtain the discrete time-series trajectory data of the traffic signal; Construct and train a transformer network model based on proximal policy optimization according to the discrete time-series trajectory data; Obtain the optimal action policy of the current traffic signal according to the transformer network model based on proximal policy optimization and execute it.
[0008] Optionally, defining the observation space, action space, and reward function includes: Define the observation space according to formula (2): , , (2) where, is the observation space, is the total number of lanes at the intersection, is the observation data of lane , is the number of vehicles in lane , is the average waiting time of low-speed vehicles in lane , is the queue length of lane , is the average vehicle speed of lane , is the emergency vehicle at The waiting time of the lane, is the phase difference between adjacent intersections, is the number of starts and stops of vehicles in the lane and is an integer number; Define the action space according to formula (3): , (3) where, is the action space, is the th discrete traffic signal phase, is the total number of traffic signal phases.
[0009] Optionally, defining the observation space, action space, and reward function further includes: Obtain the phase efficiency factor according to formula (4), , (4) where, is the th basic reward weight, is the saturation of the lane , is the saturation threshold, and , is the traffic capacity; Obtain the green wave coordination reward term according to formula (5), , (5) where, is the average speed of the coordinated lane, is the number of coordinated lanes; Obtain the emergency vehicle priority term according to formula (6), , (6) where, is the set of lanes containing emergency vehicles, is the emergency vehicle waiting time threshold; Obtain the carbon emission penalty term according to formula (7), , (7) where, is the emission coefficient, and , is the lane corresponding maximum allowable speed of the road section.
[0010] Optionally, the Markov decision process includes: Obtain the observation data of the current intersection; Obtain the action strategy of the current traffic signal and execute the action strategy; Obtain the reward function of the current traffic signal; Generate the next observation data according to the transition probability function; Return to the step of obtaining the action strategy of the current traffic signal and executing the action strategy.
[0011] Optionally, obtaining the discrete time series trajectory data of the traffic signal includes: Preset the number of sequences in one period; Judge whether the serial number of the current traffic signal is greater than or equal to the number of sequences; In the case where it is judged that the serial number of the current traffic signal is greater than or equal to the number of sequences, summarize the sequences of the number of sequences to form discrete time series trajectory data; In the case where it is judged that the serial number of the current traffic signal is less than the number of sequences, obtain the sequence of the traffic signal at the next moment; Return to the step of judging whether the serial number of the current traffic signal is greater than or equal to the number of sequences.
[0012] Optionally, constructing and training a transformer network model based on proximal policy optimization according to the discrete time series trajectory data includes: Vectorize the discrete time series trajectory data and input it into the transformer network model based on proximal policy optimization; Obtain the query vector, key vector, and value vector; Obtain the attention function according to formula (8), , (8) where, is the attention function, is the softmax function, is the query vector, is the key vector, is the value vector, is the dimension of the vector; Obtain the encoded output and input it into a multi-layer perceptron to obtain the encoded observation; Input the encoded observation into the next multi-layer perceptron to obtain the value estimate.
[0013] Optionally, constructing and training a transformer network model based on proximal policy optimization according to the discrete time series trajectory data further includes: Input the encoded observation into the decoder; Obtain the phase switching penalty term and the queue balance term; Obtain the objective function of the decoder according to formula (9). , (9) where is the objective function of the decoder, are network parameters, is the policy ratio, is the discrete traffic signal phase at time step is the time step, is the number of training cycles, is the clipping mechanism, is a constant, is the phase weight, is the equilibrium weight, is the phase switching penalty term, is the queue equilibrium term; Obtain the next action policy of the traffic signal according to the objective function of the decoder.
[0014] Optionally, obtaining the phase switching penalty term and the queue equilibrium term includes:[[]] Obtain the phase switching penalty term according to formula (10). , (10) where is the indicator function, which is 1 when the phase switches and 0 otherwise, is the action at time step is the decay coefficient, is the current phase duration, is the minimum green light duration threshold; Obtain the queue equilibrium term according to formula (11). , (11) where is the variance function, is the lane queue length; Obtain the phase weight and the equilibrium weight according to formula (12). , , (12) where is the base weight, is the vehicle density passing through the intersection at time step is the maximum traffic capacity of the section.
[0015] On the other hand, the present invention also provides a traffic signal induction control system based on deep reinforcement learning, including: Traffic signals; A controller, connected to the traffic signals, for executing any one of the above traffic signal induction control methods.
[0016] On yet another aspect, the present invention also provides a computer-readable storage medium storing instructions for being read by a machine to cause the machine to execute any one of the above traffic signal induction control methods.
[0017] Through the above technical solutions, the traffic signal induction control method and system based on deep reinforcement learning provided by the present invention model the control problem of traffic signals at intersections as a Markov decision process, define the observation space, action space, and reward function, and pay more attention to the global state. According to the Markov decision process, the vehicle states at intersections are collected in real time, actions are executed, corresponding rewards are obtained, and discrete time-series trajectory data is acquired. Then, the discrete time-series trajectory data is used to train and optimize a transformer network model based on proximal policy optimization to achieve the optimal action prediction for the next cycle of traffic signals. The transformer network model based on proximal policy optimization can effectively prevent oscillations caused by sudden policy changes during the training process, reduce the dependence on prior knowledge, and can dynamically focus on key lane information and be transferable to different intersection topologies, thereby effectively improving the control effect and generality of traffic signals. In addition, the reward function fully considers multi-dimensional state features such as phase saturation, green wave coordination reward, emergency vehicle priority, and carbon emission penalty, which can further improve the accuracy and effect of obtaining the subsequent optimal policy.
[0018] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification, and are used to explain the embodiments of the present invention together with the following specific implementation, but do not constitute a limitation to the embodiments of the present invention. In the drawings: Figure 1 is a flowchart of a traffic signal induction control method based on deep reinforcement learning according to an embodiment of the present invention; Figure 2 is a flowchart of defining the observation space, action space, and reward function in a traffic signal induction control method based on deep reinforcement learning according to an embodiment of the present invention; Figure 3It is a flowchart of real-time observation of the Markov decision process in the traffic signal sensing control method based on deep reinforcement learning according to an embodiment of the present invention; Figure 4 It is a flowchart of obtaining discrete time-series trajectory data in the traffic signal sensing control method based on deep reinforcement learning according to an embodiment of the present invention; Figure 5 It is a flowchart of training the encoding of the transformer network model based on proximal policy optimization in the traffic signal sensing control method based on deep reinforcement learning according to an embodiment of the present invention; Figure 6 It is a flowchart of training the decoding of the transformer network model based on proximal policy optimization in the traffic signal sensing control method based on deep reinforcement learning according to an embodiment of the present invention; Figure 7 It is a flowchart of obtaining the phase switching penalty term and the queue balance term in the traffic signal sensing control method based on deep reinforcement learning according to an embodiment of the present invention. Detailed implementation manners
[0020] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only for explaining and illustrating the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0021] It should be noted that the acquisition, transmission, storage, use, processing, etc. of data in the technical solution of this application all comply with the relevant regulations of laws and regulations. In the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned, and they should be regarded as exemplary. Their purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0022] Figure 1 It is a flowchart of the traffic signal sensing control method based on deep reinforcement learning according to an embodiment of the present invention. In Figure 1 it, the traffic signal sensing control method may include: In step S1, the traffic signal control problem at the intersection is modeled as a Markov decision process, and the observation space, action space, and reward function are defined. Among them, for the vehicle states at the intersection / central intersection and the action changes of the traffic signals, they can be modeled as a Markov decision process to clearly describe the dynamic changes at the intersection and optimize the decisions of the traffic signals.
[0023] In step S2, discrete temporal trajectory data of traffic lights is obtained. Among them, in the Markov decision process, observations and actions are continuously repeated at every moment, and rewards are obtained. In this process, the preset numbers of observations and actions can be preset. After recording the preset numbers of observations and actions, they can be used as sequence data for one cycle, that is, discrete temporal trajectory data. Specifically, this discrete temporal data can be used as data for subsequent network model training.
[0024] In step S3, a transformer network model based on proximal policy optimization is constructed and trained according to the discrete temporal trajectory data. Among them, after obtaining the discrete temporal trajectory data, a transformer network model based on PPO proximal policy optimization can be constructed, and the discrete temporal trajectory data is used to train this transformer network model. Specifically, after the Markov model continuously obtains sequence data for one cycle, it can be continuously input into this transformer network model for continuous training and optimization.
[0025] In step S4, the optimal action strategy of the current traffic light is obtained according to the transformer network model based on proximal policy optimization and executed. Among them, after training and optimizing the transformer network model with sequence data for one cycle of discrete temporal trajectories, the next action of the traffic light can be predicted according to the state of the current intersection and the action data of the traffic light, so as to obtain the optimal action strategy of the traffic light and execute it.
[0026] In steps S1 to S4, first, the control problem of the traffic lights at the intersection is modeled as a Markov decision process, and at the same time, the observation space / observation space, action space, and reward function of the Markov decision process are defined. The Markov decision process collects the vehicle states at the intersection in real time and executes actions, and then obtains the corresponding rewards. After statistically analyzing the above data, after the preset number of data is statistically obtained, it can be used as sequence data for one cycle, that is, discrete temporal trajectory data. At the same time, a transformer network model based on proximal policy optimization is constructed, and the discrete temporal trajectory data is used to train and optimize the transformer network model based on proximal policy optimization. According to the current state and action data, the action prediction for the next cycle can be obtained, that is, the optimal action strategy of the traffic light, and this optimal action strategy is executed.
[0027] Traditional traffic lights are generally controlled in three ways. Specifically, the first way is fixed-time control, whose principle is to preset the signal light cycle, phase duration, and phase sequence based on historical traffic flow statistics. For example, longer green light time is set for the main road during the morning rush hour. However, the limitations of the first way include: inability to respond to sudden traffic flow changes in real time (such as accidents, sudden increase in traffic flow during holidays); relying on manual parameter adjustment, requiring regular re-collection of data to adjust the timing, with high maintenance costs; poor scene generalization, and the same scheme is difficult to adapt to different intersection topologies (such as T-shaped intersections and crossroads). The second way is inductive control, whose principle is to detect the presence of vehicles in the lane in real time through sensors such as geomagnetic coils and cameras, and dynamically extend the current green light phase or switch phases. For example, when no vehicles are detected in a certain direction, its green light is skipped. However, the limitations of the second way include: local optimization, only focusing on the vehicles in the current phase, lacking global traffic flow coordination; response delay, it takes a certain amount of time to accumulate after the sensor detects the vehicle before triggering an action, and queues are likely to form during the peak period; high equipment dependence, and sensor failures will cause the control to fail. The third way is rule-based adaptive control, such as SCOOT / SCATS, whose principle is to adjust the signal cycle through a preset rule library. For example, the SCATS system dynamically allocates green light time according to the saturation. However, the limitations of the third way include: rigid rules, relying on limited rules summarized by expert experience, and it is difficult to cover complex scenarios (such as the mixture of multi-modal traffic flows); limited computing power, traditional algorithms cannot process high-dimensional data (such as the vehicle trajectories of the entire road network), and the optimization dimension is low; slow iteration speed, rule updates require manual intervention and cannot learn online. In this embodiment of the present invention, the first advantage lies in the dynamic environment adaptation ability. The self-attention mechanism of the transformer is adopted, and by parallel processing multi-lane time series data (such as the number of vehicles, speed, queue length), the spatio-temporal dependence relationship between lanes is captured. Compared with traditional methods: fixed timing and rule libraries cannot model such long-range dependencies. The Markov decision process (MDP) formalizes traffic states, actions (phase switching), and rewards (traffic efficiency) as sequential decision problems, supporting the model to learn the optimal strategy through historical trajectories. Compared with inductive control: MDP models the global state rather than the local vehicle presence, avoiding short-term decisions. The second advantage lies in efficient policy optimization and training stability. The proximal policy optimization algorithm balances between exploration and exploitation by clipping the policy update amplitude (Clipping Mechanism), preventing oscillations caused by sudden policy mutations during the training process. Compared with traditional Q-learning: PPO supports the reuse of small batches of data in multiple rounds, improving the sample efficiency. Directly using the original observation data (such as the number of vehicles, speed, etc.) as input, without the need for manual design of features (such as the saturation calculation formula in traditional methods), reducing the dependence on prior knowledge.The third advantage lies in the generalization ability in complex scenarios. The Transformer model dynamically focuses on key lane information through attention weights and can be transferred to different intersection topologies. The weighted reward function (number of vehicles, speed, queue, waiting time) supports flexible adjustment of the optimization direction. For example, in the scenario of environmental protection priority, the weight of the average speed is increased to reduce idle emissions. Generally speaking, the traffic signal induction control method adopted in the present invention has better control effect and stronger versatility compared with the traditional control method.
[0028] In this embodiment of the present invention, after modeling the control problem of the traffic signal at the intersection as a Markov decision process, it is also necessary to define the observation space, action space and reward function of the Markov decision process. The specific definition steps can be as Figure 2 shown. Specifically, in Figure 2 , the traffic signal induction control method may further include: In step S10, the observation space is defined according to formula (2): , , (2) wherein, is the observation space, is a set, is the total number of lanes at the intersection, is the lane 's observation data, is the lane 's number of vehicles, is the lane 's average waiting time of low-speed vehicles (speed < 0.1 m / s), is the lane 's queue length (low-speed vehicles), is the lane 's average vehicle speed; is the waiting time of the emergency vehicle in the lane. Here, the identification of the emergency vehicle, such as an ambulance, can also be defined to locate its lane; is the phase difference between adjacent intersections, is the lane 's start-stop times of all vehicles, is an integer number.
[0029] In step S11, the action space is defined according to formula (3): , (3) wherein, is the action space, is the th discrete traffic signal phase, is the total number of traffic signal phases. Specifically, this action space is also the discrete traffic signal phase set, which may include red, yellow, and green light phases or phase switching actions.
[0030] In step S12, the phase efficiency factor, filtering coordination reward term, emergency vehicle priority term, and carbon emission penalty term are obtained. Among them, the acquisition of the phase efficiency factor can be calculated according to formula (4), , (4) where is the th basic reward weight, is the saturation of lane , is the saturation threshold, , is the traffic capacity. Specifically, , when , the congestion penalty is strengthened. Specifically, the traffic capacity refers to the maximum number of vehicles that can pass through a certain point (such as a certain phase or lane of an intersection) per unit time (usually 1 hour) under certain road, traffic, and control conditions. It is the core index to measure the efficiency of the traffic system and directly determines whether the signal control strategy can effectively relieve congestion. in , when taking values from 1 to 4, respectively correspond to the dimensional weights of the number of vehicles, average speed, average queue length, and average waiting time of vehicles at the intersection, which are used to balance the dimensional differences of indicators and avoid a single indicator dominating the optimization process.
[0031] In a signalized intersection, the traffic capacity of a certain phase can usually be obtained according to formula (13), , (13) where is the saturation flow rate, which is the maximum flow rate of vehicles continuously passing through the stop line during the green light period, with the unit of vehicles per hour per lane. Typical values: about 1800 - 2000 vehicles per hour for straight lanes, and about 1600 - 1800 vehicles per hour for left-turn lanes. Specifically, the same phase may include multiple lanes, so can also refer to the traffic capacity of a single or multiple lanes of the same phase . is the effective green time, which is the actual available green time of the phase / lane , with the unit of seconds, and the yellow light, all-red time, and vehicle start-up loss time need to be deducted. is the signal cycle length, which is the total time for the signal light to complete all phase rotations, with the unit of seconds.
[0032] By adopting the method of dynamically adjusting the basic reward weight according to the phase saturation degree, the congestion situation of the road can be effectively considered to improve the accuracy of the reward function.
[0033] The green wave coordination reward item can be obtained by calculating according to formula (5). , (5) Among them, is the average vehicle speed of the coordinated lane. is the number of coordinated lanes. That is, the green wave coordination reward item is the sum of the ratios of the average vehicle speed of each coordinated lane that needs to coordinate the green wave passing to the phase difference between adjacent intersections. Specifically, when the reward is the largest. is the intersection spacing. is the phase difference between adjacent intersections, which refers to the time difference of the starting time of the same phase (such as the straight green light on the main road) in the signal light cycles of two adjacent intersections. Its core purpose is to realize the continuous passage of vehicle flows (such as the "green wave band" effect) by coordinating the switching timings of signal lights at multiple intersections, thereby reducing the number of stops and improving the traffic efficiency.
[0034] The emergency vehicle priority item can be obtained by calculating according to formula (6). , (6) Among them, is the set of lanes containing emergency vehicles. is the waiting time threshold for emergency vehicles, that is, the maximum tolerance time threshold for emergency vehicles to wait for the green light at intersections. When the waiting time of an emergency vehicle (such as an ambulance or a fire truck) exceeds , the system will forcibly adjust the signal phase to give priority to releasing the vehicle to avoid delaying rescue tasks. Specifically, this emergency vehicle priority item can detect special vehicles such as ambulances / fire trucks and dynamically increase the weight of the lane where they are located.
[0035] The carbon emission penalty item can be obtained by calculating according to formula (7). , (7) Among them, is the emission coefficient, and (kg / time). is the lane corresponding to the maximum allowable speed of the section.
[0036] In step S13, the reward function is obtained according to formula (1). , (1) Among them, is the reward function, is the th phase efficiency factor, is the th basic reward item, is an integer number, and , is the green wave weight, is the green wave coordination reward item, is the emergency weight, is the emergency vehicle priority item, is the emission weight, is the carbon emission penalty item. Specifically, for the basic reward item it can include as shown in formula (14): , (14) Among them, is the number of vehicles, and , is the average speed, and , is the average queue length, and , is the average waiting time, and . Specifically, the above summation formulas are all the summations of the corresponding parameters for each lane in the intersection .
[0037] In steps S10 to S13, first, the observation space in the Markov decision process of the traffic lights at the intersection is defined and observed. At the same time, the action space is defined as the set of phases of discrete traffic signals, that is, the red light, yellow light, and green light phases. Finally, according to the observed data, the current phase efficiency factor, green wave coordination reward item, emergency vehicle priority item, and carbon emission penalty item are calculated respectively, and the above parameters are summarized to obtain the final reward function. Specifically, this reward function fully considers the multi-dimensional state characteristics of phase saturation, green wave coordination reward, emergency vehicle priority, and carbon emission penalty, and can effectively improve the accuracy and effect of obtaining the subsequent optimal strategy.
[0038] In this embodiment of the present invention, after defining the observation space, action space, and reward function in the Markov decision process, the state of the intersection can be repeatedly observed in real time and the actions of the traffic lights can be executed to obtain the corresponding reward function. Specifically, the steps can be Figure 3 as shown. Specifically, in Figure 3 , this traffic signal induction control method may further include: In step S14, the observation data of the current intersection is obtained. Among them, the observation data of the current intersection is also the parameter data defined in the above observation space, generally represented by to represent the current observation data.
[0039] In step S15, the action strategy of the current traffic signal is obtained and the action strategy is executed. Among them, the action strategy of the current traffic signal can include the initial action strategy or the optimized action strategy, and only the current action strategy needs to be executed. Generally represented by to represent the current action, and .
[0040] In step S16, the reward function of the current traffic signal is obtained. Among them, according to the current observation data and the action strategy, the current immediate reward function can be obtained according to formula (3), and this reward function can quantify the quality of the current decision. Specifically, generally represented by or to represent the current reward function.
[0041] In step S17, the next observation data is generated according to the transition probability function. Among them, the transition probability function is directly sampled through the interaction between the agent / traffic signal and the environment. The agent / traffic signal collects empirical data through trial and error and learns the strategy and value function from it. When the agent / traffic signal selects to execute the action , the next observation state is automatically calculated according to the environmental state (speed, queue, etc.). Therefore, the transition probability is determined by the internal mechanism.
[0042] In step S18, return to the step of obtaining the action strategy of the current traffic signal and executing the action strategy. Among them, Markov continuously repeats observing and executing actions and obtaining rewards at every moment. Therefore, after obtaining the next observation data, the next action strategy will also be executed, and then the corresponding reward or reward function will be obtained.
[0043] In steps S14 to S18, first obtain the observation data of the current intersection, and at the same time obtain the action strategy of the current traffic signal and execute the action strategy. The current reward function can be obtained according to the current observation data and the executed action. Obtain the next observation data according to the transition probability function and execute the next action, and loop in this way to describe and solve the control problem of the traffic signal.
[0044] In this embodiment of the present invention, in order to obtain the discrete time-series trajectory data of the traffic signal, the number of observation and action moments can be preset, and the corresponding data can be statistically summarized according to this number of moments. The specific steps can be as Figure 4 shown. Specifically, in Figure 4In this case, the traffic signal induction control method may further include: In step S20, preset the number of sequences in one cycle. Among them, the number of sequences in one cycle may include pieces.
[0045] In step S21, determine whether the serial number of the current traffic signal is greater than or equal to the number of sequences. Among them, the sequence of the traffic signal may include triple data streams , and the serial number of the traffic signal is also the moment value.
[0046] In step S22, when it is determined that the serial number of the current traffic signal is greater than or equal to the number of sequences, summarize the sequences with the number of sequences to form discrete time series trajectory data. Among them, if the serial number of the current traffic signal is greater than or equal to the number of sequences, it means that a sequence of a preset cycle has been collected. Summarize and statistically analyze these sequences to form discrete time series trajectory data .
[0047] In step S23, when it is determined that the serial number of the current traffic signal is less than the number of sequences, obtain the sequence of the traffic signal at the next moment. Among them, if the serial number of the current traffic signal is less than the number of sequences, it means that a sequence of a preset cycle has not been collected fully and the next sequence needs to be obtained continuously.
[0048] In step S24, return to the step of determining whether the serial number of the current traffic signal is greater than or equal to the number of sequences.
[0049] In steps S20 to S24, after presetting the number of sequences in one cycle, judge the serial number of the sequence data of the traffic signal collected in real time during the Markov decision process. If the serial number is greater than or equal to the preset number of sequences, it means that the collection of sequence data for one cycle is completed, and discrete time series trajectory data can be formed by summarization. Otherwise, continuous collection is required.
[0050] In this embodiment of the present invention, after obtaining the discrete time series trajectory data for one cycle, the discrete time series trajectory data can be used to train the transformer network model based on proximal policy optimization. Specifically, the training steps can be as Figure 5 shown. Specifically, in Figure 5 In this case, the traffic signal induction control method may further include: In step S30, vectorize the discrete time series trajectory data and input it into the transformer network model based on proximal policy optimization. Among them, vectorize the observation data in the discrete time series trajectory data and use it as the input of the network model.
[0051] In step S31, a query vector, a key vector and a value vector are obtained. , the key vector , value vector ,in, is the input sequence, , , is the learnable weight matrix.
[0052] In step S32, the attention function is obtained according to formula (8): , (8) in, is the attention function, is the softmax function, is the query vector, is the key vector, is a value vector, is the dimension of the vector, including the dimension of the query and key vectors, used to scale the dot product to prevent gradient explosion, is the matrix transpose, the matrix transpose The cycle of the following They are not the same and need to be distinguished. Specifically, in the attention mechanism, Sharing the same set of parameters, the context-awareness of sequential decisions is enhanced by global dependency modeling.
[0053] In step S33, the encoded output is obtained and input into a multi-layer perceptron to obtain a coded observation. The encoded output is input into a multi-layer perceptron MLP to obtain a new feature coded observation.
[0054] In step S34, the coded observation is input into the next multi-layer perceptron to obtain a value estimate. After the coded observation is obtained, it can be input into the next multi-layer perceptron, and the value estimate can be used to optimize the coding network parameters in the transformer network.
[0055] In step S30 to step S34, when the network model needs to be trained and encoded, the discrete time series trajectory data needs to be vectorized, which may include vectorization of observation data and actions. Obtain the query vector, key vector, and value vector of the vectorized sequence, calculate the attention function, and encode the output. Input the encoded output into a multi-layer perceptron to obtain the encoded observation, and further input the encoded observation into the multi-layer perceptron to obtain the value estimate.
[0056] In this embodiment of the present invention, after the discrete time-series trajectory data is vectorized and encoded, the decoding operation can be performed. Specifically, the decoding process can be as Figure 6 shown. Specifically, in Figure 6 , the traffic signal sensing control method may further include: In step S35, the encoded observation is input into the decoder.
[0057] In step S36, obtain the phase switching penalty term and the queue balance term. Among them, the obtaining of the phase switching penalty term and the queue balance term may include the steps as Figure 7 shown. Specifically, in Figure 7 , the obtaining step may include: In step S360, obtain the phase switching penalty term according to formula (10), , (10) where is the indicator function, which is 1 when the phase switches and 0 otherwise, is the action at time step is the attenuation coefficient to ensure a greater short-term switching penalty, is the current phase duration, is the minimum green light duration threshold. Specifically, is the minimum green light duration set by the traffic signal control system to ensure the safe and orderly passage of vehicles through the intersection. For example, if is 10 seconds, the signal must maintain the current green light state for at least 10 seconds before switching to the next phase. If the green light time is too short, such as 5 seconds, vehicles may not have enough time to pass through the intersection, resulting in an increase in queue length, a decrease in traffic efficiency, and even potential safety hazards. Frequent phase switching can lead to an increase in vehicle starts and stops, fuel waste, and driver discomfort. Therefore, the number of phase switches between adjacent actions is counted, and a penalty is imposed if the current phase is different from the previous phase.
[0058] In step S361, obtain the queue balance term according to formula (11), , (11) where is the variance function, is the queue length of lane . An overly long single-lane queue will cause the "green wave band" to break and reduce the overall efficiency of the road network. Therefore, the variance of the queue lengths of each lane is calculated to encourage an even distribution.
[0059] In step S362, obtain the phase weight and the balance weight according to formula (12), , , (12) Among them, is the basic weight, such as initially 0.5, is the real-time traffic flow, that is, the vehicle density passing through the intersection at time step , usually measured by the number of vehicles per minute or vehicle density (such as lane occupancy), is the maximum traffic capacity of the road section. Specifically, congestion needs to be alleviated first during peak hours, and the traffic speed can be optimized during non-peak hours. During peak hours is increased to emphasize queue balance.
[0060] In step S37, the objective function of the decoder is obtained according to formula (9), , (9) Among them, is the objective function of the decoder, are the network parameters, is the policy ratio, is the discrete traffic signal phase at time step, is the time step, is the number of training cycles in one training period, is the clipping mechanism that limits the network parameters to , preventing the policy update step size from being too large and ensuring training stability, is a very small constant as the limit range, is the phase weight, is the balance weight, is the phase switching penalty term, is the queue balance term. Specifically, the policy ratio is the probability ratio of the new policy (current parameters ) to the old policy (parameters before update ) under the same observation state and action, and can be specifically as shown in formula (15), , (15) Among them, is the encoded observation, is the new policy, is the old policy.
[0061] In step S38, the next action policy of the traffic signal is obtained according to the objective function of the decoder. Among them, the next action policy can include the action policy of the traffic signal in the next cycle.
[0062] In steps S35 to S38, during the decoding process of the network model, a phase switching penalty term and a queue balance term are introduced based on the clipping objective function of proximal policy optimization. By penalizing frequent phase switching, the invalid signal cycles can be effectively reduced, improving the driver experience. By means of the variance term, the model is forced to focus on the bottleneck lanes, avoiding the spread of local congestion. In addition, the dynamic weights of the phase switching penalty term and the queue balance term enable the algorithm to automatically switch the optimization objective under different traffic flows, taking into account both throughput and efficiency.
[0063] On the other hand, the present invention also provides a traffic signal sensing control system based on deep reinforcement learning. Specifically, the traffic signal sensing control system may include traffic signals and a controller. The controller is connected to the traffic signals and is configured to execute any one of the above traffic signal sensing control methods.
[0064] On yet another aspect, the present invention also provides a computer-readable storage medium. Specifically, the computer-readable storage medium stores instructions that are used to be read by a machine so that the machine executes any one of the above traffic signal sensing control methods.
[0065] Through the above technical solutions, a traffic signal sensing control method and system provided by the present invention model the control problem of traffic signals at intersections as a Markov decision process, define the observation space, action space, and reward function, and pay more attention to the global state. According to the Markov decision process, the vehicle states at intersections are collected in real time, actions are executed, corresponding rewards are obtained, and discrete time-series trajectory data is acquired. Then, the discrete time-series trajectory data is used to train and optimize a transformer network model based on proximal policy optimization to achieve the optimal action prediction for the next cycle of traffic signals. The transformer network model based on proximal policy optimization can effectively prevent oscillations caused by sudden changes in policies during the training process, reduce the dependence on prior knowledge, and at the same time can dynamically focus on key lane information and can be migrated to different intersection topologies, thereby effectively improving the effect and generality of traffic signal control. In addition, the reward function fully considers multi-dimensional state features such as phase saturation, green wave coordination reward, emergency vehicle priority, and carbon emission penalty, which can further improve the accuracy and effect of obtaining the subsequent optimal policy.
[0066] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0067] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0068] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0069] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0070] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0071] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0072] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0073] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0074] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A traffic signal induction control method based on deep reinforcement learning, characterized in that, Including: Model the traffic signal control problem at intersections as a Markov decision process, and define the observation space, action space, and reward function, where the reward function includes: Obtain the reward function according to formula (1). ,(1) Among them, is the reward function, is the th phase efficiency factor, is the th basic reward item, is an integer number, and , is the green wave weight, is the green wave coordination reward item, is the emergency weight, is the emergency vehicle priority item, is the emission weight, is the carbon emission penalty item; Obtain the discrete time-series trajectory data of the traffic signal. Construct and train a Transformer network model based on proximal policy optimization according to the discrete time-series trajectory data. Obtain the optimal action policy of the current traffic signal according to the Transformer network model based on proximal policy optimization and execute it.
2. The traffic signal induction control method according to claim 1, characterized in that, Defining the observation space, action space, and reward function includes: Define the observation space according to formula (2). , ,(2) Among them, is the observation space, is the total number of lanes at the intersection, is the lane observation data, is the number of vehicles in the lane , is the average waiting time of low-speed vehicles in the lane , is the queue length of the lane , is the average vehicle speed in the lane , is the waiting time of the emergency vehicle in the lane, is the phase difference between adjacent intersections, is the lane start-stop times of the vehicles in, is an integer number; Define the action space according to formula (3). ,(3) Among them, is the action space, is the th discrete traffic signal phase, is the total number of traffic signal phases.
3. The traffic signal sensing control method according to claim 2, wherein Defining the observation space, action space, and reward function also includes: Obtain the phase efficiency factor according to formula (4). ,(4) Among them, is the th basic reward weight, is the saturation of lane , is the saturation threshold, and , is the traffic capacity; Obtain the green wave coordination reward term according to formula (5). ,(5) Among them, To coordinate the average vehicle speed of lanes, is the number of lanes to be coordinated; Obtain the emergency vehicle priority term according to formula (6). ,(6) Among them, is the set of lanes containing emergency vehicles, is the waiting time threshold for emergency vehicles; Obtain the carbon emission penalty term according to formula (7). ,(7) Among them, is the emission coefficient, and , is the maximum allowable speed of the corresponding section of the lane.
4. The traffic signal induction control method according to claim 1, characterized in that, The Markov decision process includes: Obtain the observation data of the current intersection. Obtain the action policy of the current traffic signal and execute the action policy. Obtain the reward function of the current traffic signal. Generate the next observation data according to the transition probability function. Return to the step of obtaining the action policy of the current traffic signal and executing the action policy.
5. The traffic signal induction control method according to claim 4, wherein Obtaining the discrete time-series trajectory data of the traffic signal includes: Preset the number of sequences in a cycle. Judge whether the serial number of the current traffic signal is greater than or equal to the number of sequences. In the case where it is judged that the serial number of the current traffic signal is greater than or equal to the number of sequences, summarize the sequences of the number of sequences to form discrete time-series trajectory data. In the case where it is judged that the serial number of the current traffic signal is less than the number of sequences, obtain the sequence of the traffic signal at the next moment. Return to the step of judging whether the serial number of the current traffic signal is greater than or equal to the number of sequences.
6. The traffic signal induction control method according to claim 2, wherein Constructing and training a Transformer network model based on proximal policy optimization according to the discrete time-series trajectory data includes: Vectorize the discrete time-series trajectory data and input it into the Transformer network model based on proximal policy optimization. Obtain the query vector, key vector, and value vector. Obtain the attention function according to formula (8). ,(8) Among them, is the attention function, is the softmax function, is the query vector, is the key vector, is the value vector, is the dimension of the vector; Obtain the encoded output and input it into a multi-layer perceptron to obtain the encoded observation. Input the encoded observation into the next multi-layer perceptron to obtain the value estimate.
7. The traffic signal induction control method according to claim 6, wherein Constructing and training a Transformer network model based on proximal policy optimization according to the discrete time-series trajectory data also includes: Input the encoded observation into the decoder. Obtain the phase switching penalty term and the queue balance term. Obtain the objective function of the decoder according to formula (9). ,(9) Among them, is the objective function of the decoder, is the network parameter, is the policy ratio, is the discrete traffic signal phase at the time step, is the time step, is the number of training cycles, is the clipping mechanism, is a constant, is the phase weight, is the equilibrium weight, is the phase switching penalty term, is the queue equilibrium term; Obtain the next action policy of the traffic signal according to the objective function of the decoder.
8. The traffic signal induction control method according to claim 7, wherein, Obtaining the phase switching penalty term and the queue balance term includes: Obtain the phase switching penalty term according to formula (10). ,(10) wherein, is an indicator function, which is 1 when the phase switches and 0 otherwise, is the action at the time step, is the attenuation coefficient, is the current phase duration, is the minimum green light duration threshold; Obtain the queue balance term according to formula (11). ,(11) Among them, is the variance function, is the queue length of the lane; Obtain the phase weight and the equalization weight according to formula (12). , ,(12) Among them, is the basic weight, is the time step when the vehicle density passing through the intersection, is the maximum traffic capacity of the road section.
9. A traffic signal induction control system based on deep reinforcement learning, characterized in that, Including: Traffic lights; A controller, connected to the traffic lights, for executing the traffic light sensing control method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions for being read by a machine to cause the machine to execute the traffic light sensing control method according to any one of claims 1-8.
Citation Information
Patent Citations
Traffic light control method and system based on multi-agent reinforcement learning in control area
CN115631638A
Transform adaptive deep reinforcement learning distributed flexible load intelligent regulation and control method
CN117728428A
Adaptive traffic signal control method based on reinforcement learning and self-attention mechanism
CN118942261A
Multi-objective optimization for real time traffic light control and navigation systems for urban saturated networks
US20080094250A1
Data processing method and apparatus, device, and computer-readable storage medium
WO2020224444A1