Traffic signal lamp adaptive timing control method and system based on deep reinforcement learning

By applying deep reinforcement learning technology in traffic signal control to generate and optimize the movement of traffic lights, the problem of limited adaptability and effectiveness in complex traffic environments is solved, and flexible and precise control of traffic lights is achieved, reducing the waste of green light signals.

CN119992846APending Publication Date: 2025-05-13HENAN UNIV OF SCI & TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510151101.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Traditional adaptive traffic signal control methods are limited in adaptability and effectiveness in complex traffic environments, and cannot effectively reduce the waste of green light signals.

Method used

Adaptive time-matching control method for traffic lights based on deep reinforcement learning is adopted to obtain real-time traffic states through the pre-constructed urban intersection model, and theoretical optimal actions are generated using the depth deterministic strategy gradient algorithm, and the actual optimal actions are obtained to control traffic lights through comparison with the standard action interval.

Benefits of technology

It realizes flexible and precise control of traffic lights, minimizes the waste of green light signals, and improves the efficiency of traffic flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992846A_ABST
    Figure CN119992846A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic signal lamp adaptive timing control method based on deep reinforcement learning. The method comprises the following steps: obtaining a real-time traffic state based on a pre-constructed urban intersection model; based on the real-time traffic state, a theoretical optimal action is generated through a pre-trained traffic signal control model based on a depth deterministic strategy gradient algorithm, the theoretical optimal action comprises a plurality of single actions in one-to-one correspondence with the traffic lights, and each single action comprises an action state and an action duration; comparing the theoretical optimal action with a predetermined standard action interval, and correcting the theoretical optimal action according to a comparison result to obtain an actual optimal action; and controlling the traffic signal lamp based on the actual optimal action. The invention provides a traffic signal lamp adaptive timing control method and system based on deep reinforcement learning. The traffic signal lamp adaptive timing control method and system can control traffic signal lamps more flexibly and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic light control, and in particular to a method and system for adaptive timing control of traffic lights based on deep reinforcement learning. Background Art

[0002] With the acceleration of urbanization and the increase in the number of cars, urban traffic congestion is becoming increasingly serious. Traditional fixed-time control strategies are inefficient because they cannot adapt to real-time traffic conditions. Adaptive traffic signal control uses real-time data and algorithms to flexibly adjust signal timing, becoming a more effective solution. However, in the face of complex traffic environments, the adaptability and effectiveness of traditional adaptive traffic signal control methods are limited. The development of deep reinforcement learning provides a new way to solve this problem, learning the optimal control strategy through continuous interaction with the traffic environment. Strategies based on deep reinforcement learning are divided into two categories: signal phase control and signal phase duration control. The former is flexible but may cause the driver to be unable to predict the status of the signal light, while the latter adjusts the phase duration in real time according to traffic conditions, but the phase switching order is fixed, which may lead to waste of green light time. Summary of the invention

[0003] In order to address the deficiencies in the prior art, the present invention provides a traffic light adaptive timing control method and system based on deep reinforcement learning, which can control traffic lights more flexibly and accurately.

[0004] In order to achieve the above object, the specific scheme adopted by the present invention is: a traffic light adaptive timing control method based on deep reinforcement learning, comprising the following steps: Get real-time traffic status based on pre-built urban intersection models; Based on the real-time traffic status, a traffic signal control model based on a deep deterministic policy gradient algorithm that is pre-trained generates a theoretical optimal action. The theoretical optimal action includes multiple individual actions that correspond to traffic lights one by one. The individual actions include the action state and the duration of the action. Compare the theoretical optimal action with the predetermined standard action interval, and modify the theoretical optimal action according to the comparison result to obtain the actual optimal action; Traffic lights are controlled based on actual optimal actions.

[0005] As a further optimization of the above-mentioned traffic light adaptive timing control method based on deep reinforcement learning: the urban intersection model includes multiple roads and multiple intersections, where the roads are in the north-south direction or the east-west direction, and each road includes at least one left-turn lane, at least one through lane and at least one right-turn lane, and all lanes in the same direction correspond to one traffic light.

[0006] As a further optimization of the above-mentioned traffic light adaptive timing control method based on deep reinforcement learning: in the urban intersection model, the road is divided into multiple cells distributed along the direction of travel, and the length of the cells gradually increases in the direction away from the intersection.

[0007] As a further optimization of the above-mentioned traffic light adaptive timing control method based on deep reinforcement learning: after obtaining the real-time traffic status, the real-time traffic status is converted into a state information matrix S = {P, V, T, D}, where P is the position of the vehicle, V is the vehicle speed, T is the phase switching order of the traffic light, and D is the phase duration of the traffic light.

[0008] As a further optimization of the above-mentioned traffic light adaptive timing control method based on deep reinforcement learning: the traffic signal control model includes an Actor network based on a policy function and a Critic network based on a value function, the Actor network is used to generate alternative actions based on real-time traffic status, and the Actor network includes a first main network and a first target network, the Critic network is used to perform value evaluation on the alternative actions, and the Critic network includes a second main network and a second target network.

[0009] As a further optimization of the above-mentioned traffic signal adaptive timing control method based on deep reinforcement learning: the method of training the traffic signal control model includes: Collect multiple real traffic conditions to form a data set and initialize the model parameters; Randomly select training actions based on the actual traffic state and the phase switching order of traffic lights, and calculate the reward value after executing the training action; The actual traffic status, training actions, reward values, and the changes in the actual traffic status are stored as experience in the experience pool; a number of samples are extracted from the summed tree experience pool, and the target value is obtained by solving the objective function based on the samples; the first main network of the Actor network is updated by minimizing the loss, and the second main network of the Critic network is updated by gradient descent; Updating the first target network and the second target network; The training ends when the noise exploration rate is lower than the preset threshold.

[0010] As a further optimization of the above-mentioned traffic light adaptive timing control method based on deep reinforcement learning: give priority to experience when storing it in the experience pool, and extract experience based on the priority when extracting samples from the experience pool.

[0011] As a further optimization of the above traffic light adaptive timing control method based on deep reinforcement learning: the empirical expression is (st ,a t ,r t ,s t+1 ), where r t is the reward value, a t For training action, s t+1 is the change result of the real traffic state after executing the training action, s t It is the real traffic status; The way to prioritize experience is: Among them, Q T is, γ is, μ T is, μ is, θ μ is the weight of the first main network, θ Q is the weight of the second main network, is the weight of the second target network, is the weight of the second target network; When extracting samples from the experience pool, the probability of experience being extracted is: Among them, p i for.

[0012] As a further optimization of the above-mentioned traffic light adaptive timing control method based on deep reinforcement learning: the method of randomly selecting training actions is: select a with probability ε t =μ(s t |θ μ )+N a and choose a with probability ε t =μ(s t |θ μ ), where N a To adaptively change the noise, we have: Among them, ξ∈[0,1] is the noise selection probability, X is, and U(·) is.

[0013] A traffic light adaptive timing control method system based on deep reinforcement learning is used to implement the above-mentioned traffic light adaptive timing control method based on deep reinforcement learning. The system includes: Data acquisition module, used to obtain real-time traffic status based on pre-built urban intersection models; The model operation module is used to generate theoretical optimal actions based on real-time traffic status through a pre-trained traffic signal control model based on a deep deterministic policy gradient algorithm; The action optimization module is used to compare the theoretical optimal action with the predetermined standard action interval, and to modify the theoretical optimal action according to the comparison result to obtain the actual optimal action; Control output module for controlling traffic lights based on actual optimal actions.

[0014] Beneficial effects: The present invention combines signal phase duration control and signal phase conversion mechanism to flexibly control signal phase and phase duration, thereby minimizing the waste of green light signals; the present invention designs a traffic signal control model based on the improved DDPG algorithm, and in order to avoid the network training falling into the local optimum, introduces an exploration noise that adaptively changes according to the network output action value, thereby improving the exploration performance of the algorithm; the present invention adopts a priority experience playback mechanism based on a sum tree structure, thereby improving the sample learning efficiency, being able to obtain high-quality experience during the model training process, and better improving the learning performance of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flow chart of the control method of the present invention; Figure 2 It is a schematic diagram of the training process of the traffic signal control model; Figure 3 It is a schematic diagram of an intersection in the urban intersection model. DETAILED DESCRIPTION

[0016] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0017] like Figure 1 As shown, the traffic light adaptive timing control method based on deep reinforcement learning includes S1 to S4.

[0018] S1. Obtain real-time traffic status based on a pre-built urban intersection model. The urban intersection model includes multiple roads and multiple intersections, where the roads are north-south or east-west, and each road includes at least one left-turn lane, at least one through lane, and at least one right-turn lane. All lanes in the same direction correspond to one traffic light. Figure 3As shown, a city intersection model can be constructed in existing urban traffic simulation software, for example, existing software such as SUMO, MOSAIC and PHABMICS can be used. The number of lanes in the road can be increased or decreased according to actual conditions. For example, in one embodiment of the present invention, a road includes a left-turn lane, two through lanes and a right-turn lane.

[0019] Furthermore, in the urban intersection model, the road is divided into multiple cells distributed along the direction of travel, and the length of the cell gradually increases in the direction away from the intersection, because the farther the vehicle is from the intersection, the smaller the impact on the traffic light. On this basis, after obtaining the real-time traffic status, the real-time traffic status is converted into a state information matrix S = {P, V, T, D}, where P is the position of the vehicle, V is the vehicle speed, T is the phase switching order of the traffic light, and D is the phase duration of the traffic light. Among them, the positions of all vehicles can form a position matrix, the speeds of all vehicles can form a speed matrix, and in the speed matrix, the vehicle speed is divided by the speed limit of the road to obtain the normalized speed. The phase conversion rule of the traffic light is set to L = {l NS-S ,l NS-L ,l EW-S ,l EW-R}, where l NS-S Green light for north-south straight driving, l NS-L The green light for north-south left turn is l EW-S Green light for east-west straight travel, l EW-R The phase switching order of traffic lights is used to represent the order of each element in the phase conversion rule.

[0020] S2. Generate theoretical optimal actions based on real-time traffic status through a pre-trained traffic signal control model based on a deep deterministic policy gradient algorithm. The theoretical optimal actions include multiple individual actions corresponding to traffic lights. The individual actions include action states and action durations. The action states are used to control the color of traffic lights, and the action durations are used to control the duration of traffic lights in a certain color. The traffic signal control model is constructed based on a deep deterministic policy gradient algorithm (DDPG), which includes an Actor network based on a policy function and a Critic network based on a value function. The Actor network is used to generate alternative actions based on real-time traffic status, and the Actor network includes a first main network and a first target network. The Critic network is used to evaluate the value of alternative actions, and the Critic network includes a second main network and a second target network. The specific structure and working principle of the Actor network and the Critic network are prior art and will not be repeated here.

[0021] Further, such as Figure 2 As shown, the method for training the traffic signal control model includes T1 to T7.

[0022] T1. Collect multiple real traffic states to form a data set and initialize the model parameters. The method of initializing the model parameters includes: setting the capacity of the experience pool R to D, setting the capacity of extracting samples from the experience pool to N, setting the noise exploration rate to ε, and setting the weight of the first main network to θ μ , the weights of the second main network are set to θ Q , the weights of the second target network are set to The weights of the second target network are set to

[0023] T2. Randomly select training actions based on the actual traffic status and the phase switching order of traffic lights, and calculate the reward value after executing the training action. The method of randomly selecting training actions is: select a with probability ε t =μ(s t |θ μ )+N a and choose a with probability ε t =μ(s t |θ μ ), where N a To adaptively change the noise, we have: Among them, ξ∈[0,1] is the noise selection probability, which means that according to the current action value μ(s t |θ μ ), the noise is 1-μ(s t |θ μ ) has a probability in [0,1-μ(s t |θ μ )] space upward exploration, there is μ(s t |θ μ ) has a probability of [-μ(s t |θ μ ),0] to explore downward, so as to realize the adaptive adjustment selection based on the current action value, X is, U(·) is.

[0024] The calculation method of reward value is R = k1 * W t +k2*Q l +k3V t , where k1, k2, k3 are weight factors for each evaluation index. The meaning and calculation method of each evaluation index are as follows: Average vehicle waiting time Vehicle queue length Average vehicle speed Among them, V min is the minimum threshold of driving speed, N car is the number of vehicles entering the intersection in a training cycle, V i is the speed of the i-th vehicle, N l is the number of lanes in the intersection, S tl is the state of the signal light, S tl 0 means the signal light is green, V n is the average speed of the nth vehicle passing through the intersection.

[0025] T3. Store the actual traffic status, training actions, reward values, and changes in the actual traffic status into the experience pool as experience. Experience is represented by (s t ,a t ,r t ,s t+1 ), where r t is the reward value, a t For training action, s t+1 is the change result of the real traffic state after executing the training action, s t The real traffic status is given priority when storing experience in the experience pool, and experience is extracted based on the priority when extracting samples from the experience pool.

[0026] The way to prioritize experience is: Among them, Q T is, γ is, μ T is, μ is, θ μ is the weight of the first main network, θ Q is the weight of the second main network, is the weight of the second target network, is the weight of the second target network. When the |TD-error| value is large, it means that the current Q value is much different from the target Q value. At this time, this experience value should be selected to update the network.

[0027] When extracting samples from the experience pool, the probability of experience being extracted is: Among them, p i for.

[0028] T4, extract several samples from the sum tree experience pool, and obtain the target value by solving the target function based on the samples. The probability of sampling from a mini-batch in R Furthermore, the method to solve the objective function is

[0029] T5. Update the first main network of the Actor network by minimizing the loss, and update the second main network of the Critic network by gradient descent. The method for updating the first main network is: The method for updating the second main network is:

[0030] T6. Update the first target network and the second target network. The specific method is:

[0031] T7. When the noise exploration rate is lower than a preset threshold, the training is terminated. Specifically, when the noise exploration rate ε gradually decreases from 1 to 0.1 when it exceeds 60% of the total rounds M, the training is terminated.

[0032] S3. Compare the theoretical optimal action with the predetermined standard action interval, and modify the theoretical optimal action according to the comparison result to obtain the actual optimal action. After obtaining the theoretical optimal action, the theoretical optimal action may not meet the actual needs, or exceed the reasonable range. For example, the green light time is too short, resulting in vehicles being unable to pass smoothly, or the green light time is too long, resulting in traffic jams in other directions. In order to solve this problem, the present invention uses the standard action interval to modify the theoretical optimal action. In one embodiment of the present invention, the standard action interval is used to limit the green light duration, where the lower limit is D min The upper limit is D max When the green light duration S of the theoretical optimal action is less than D min When S>D max When S = D max In one embodiment of the present invention, D min and D max 10s and 60s respectively.

[0033] S4. Control the traffic lights based on the actual optimal action.

[0034] On the other hand, considering the different traffic demands and traffic pressures on different roads, the theoretical optimal action can also set appropriate traffic light phase conversion rules for each intersection based on the actual situation of each intersection to control vehicles to pass in a reasonable order. For example, at some intersections, straight-going vehicles can be given priority, followed by left-turning vehicles, while at other intersections, left-turning vehicles can be given priority, followed by straight-going vehicles.

[0035] The present invention combines signal phase duration control and signal phase conversion mechanism to flexibly control signal phase and phase duration, thereby minimizing the waste of green light signals; the present invention designs a traffic signal control model based on an improved DDPG algorithm, and in order to prevent network training from falling into a local optimum, introduces an exploration noise that adaptively changes according to the network output action value, thereby improving the exploration performance of the algorithm; the present invention adopts a priority experience replay mechanism based on a sum tree structure, thereby improving sample learning efficiency, being able to obtain high-quality experience during the model training process, and better improving the learning performance of the network.

[0036] The present invention also provides a traffic light adaptive timing control method system based on deep reinforcement learning, which is used to implement the above-mentioned traffic light adaptive timing control method based on deep reinforcement learning. The system includes a data acquisition module, a model operation module, an action optimization module and a control output module.

[0037] The data acquisition module is used to obtain real-time traffic status based on the pre-built urban intersection model.

[0038] The model running module is used to generate theoretical optimal actions based on real-time traffic status through a pre-trained traffic signal control model based on a deep deterministic policy gradient algorithm.

[0039] The action optimization module is used to compare the theoretical optimal action with the predetermined standard action interval, and to modify the theoretical optimal action according to the comparison result to obtain the actual optimal action.

[0040] Control output module for controlling traffic lights based on actual optimal actions.

[0001] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0041] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0042] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A traffic light adaptive timing control method based on deep reinforcement learning, characterized in that: The steps include: Get real-time traffic status based on pre-built urban intersection models; Based on the real-time traffic status, a traffic signal control model based on a deep deterministic policy gradient algorithm that is pre-trained generates a theoretical optimal action. The theoretical optimal action includes multiple individual actions that correspond to traffic lights one by one. The individual actions include the action state and the duration of the action. Compare the theoretical optimal action with the predetermined standard action interval, and modify the theoretical optimal action according to the comparison result to obtain the actual optimal action; Traffic lights are controlled based on actual optimal actions.

2. The method for adaptive timing control of traffic lights based on deep reinforcement learning according to claim 1, characterized in that: The urban intersection model includes multiple roads and multiple intersections, wherein the roads are in the north-south direction or the east-west direction, and each road includes at least one left-turn lane, at least one through lane and at least one right-turn lane, and all lanes in the same direction correspond to one traffic light.

3. The method for adaptive timing control of traffic lights based on deep reinforcement learning as claimed in claim 2, characterized in that: In the urban intersection model, the road is divided into multiple cells distributed along the travel direction, and the length of the cells gradually increases in the direction away from the intersection.

4. The method for adaptive timing control of traffic lights based on deep reinforcement learning as claimed in claim 3, characterized in that: After acquiring the real-time traffic status, the real-time traffic status is converted into a state information matrix S = {P, V, T, D}, where P is the position of the vehicle, V is the vehicle speed, T is the phase switching sequence of the traffic light, and D is the phase duration of the traffic light.

5. The method for adaptive timing control of traffic lights based on deep reinforcement learning according to claim 1, characterized in that: The traffic signal control model includes an Actor network based on a policy function and a Critic network based on a value function. The Actor network is used to generate alternative actions based on real-time traffic status, and the Actor network includes a first main network and a first target network. The Critic network is used to evaluate the value of the alternative actions, and the Critic network includes a second main network and a second target network.

6. The method for adaptive timing control of traffic lights based on deep reinforcement learning as claimed in claim 5, characterized in that: Methods for training traffic signal control models include: Collect multiple real traffic conditions to form a data set and initialize the model parameters; Randomly select training actions based on the actual traffic state and the phase switching order of traffic lights, and calculate the reward value after executing the training action; The actual traffic status, training actions, reward values, and the changes in the actual traffic status are stored as experience in the experience pool; a number of samples are extracted from the summed tree experience pool, and the target value is obtained by solving the objective function based on the samples; the first main network of the Actor network is updated by minimizing the loss, and the second main network of the Critic network is updated by gradient descent; Updating the first target network and the second target network; The training ends when the noise exploration rate is lower than the preset threshold.

7. The method for adaptive timing control of traffic lights based on deep reinforcement learning according to claim 6, characterized in that: Priorities are assigned to experiences when they are stored in the experience pool, and experiences are extracted based on the priorities when samples are extracted from the experience pool.

8. The method for adaptive timing control of traffic lights based on deep reinforcement learning as claimed in claim 7, characterized in that: Experience is expressed as (s t ,a t ,r t ,s t+1 ), where r t is the reward value, a t For training action, s t+1 is the change result of the real traffic state after executing the training action, s t It is the real traffic status; The way to prioritize experience is: Among them, Q T is, γ is, μ T is, μ is, is the θth μ The weight of a main network, θ Q is the weight of the second main network, is the weight of the second target network, is the weight of the second target network; When extracting samples from the experience pool, the probability of experience being extracted is: Among them, p i for.

9. The method for adaptive timing control of traffic lights based on deep reinforcement learning according to claim 7, characterized in that: The method of randomly selecting training actions is: select a with probability ε t =μ(s t |θ μ )+N a and choose a with probability ε t =μ(s t |θ μ ), where N a To adaptively change the noise, we have: Among them, ξ∈[0,1] is the noise selection probability, X is, and U(·) is.

10. A traffic light adaptive timing control method system based on deep reinforcement learning, characterized in that: The system is used to implement the traffic light adaptive timing control method based on deep reinforcement learning as described in any one of claims 1 to 9, and the system comprises: Data acquisition module, used to obtain real-time traffic status based on pre-built urban intersection models; The model operation module is used to generate theoretical optimal actions based on real-time traffic status through a pre-trained traffic signal control model based on a deep deterministic policy gradient algorithm; The action optimization module is used to compare the theoretical optimal action with the predetermined standard action interval and optimize the action according to the comparison results. The theoretical optimal action is corrected to obtain the actual optimal action; Control output module for controlling traffic lights based on actual optimal actions.

Citation Information

Cited By

  • Special line painting control method and system based on intelligent algorithm

    CN120848293A

  • A special line painting control method and system based on intelligent algorithm

    CN120848293B