Asymmetric traffic flow signal control method based on deep reinforcement learning

By applying deep reinforcement learning technology in traffic signal control, using lane occupancy as the state space, dynamically adjusting the signal control solution, the resource utilization imbalance and traffic congestion caused by asymmetric traffic flows are solved, and more efficient traffic flow management is achieved.

CN120220439AInactive Publication Date: 2025-06-27SHANDONG UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510694058.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problems of uneven temporal and spatial resource utilization and traffic congestion at intersections caused by asymmetric traffic flows, especially in environments with strong dynamic changes and irregularities in traffic flows.

Method used

The asymmetric traffic flow signal control method based on deep reinforcement learning is adopted. By constructing a deep reinforcement learning signal control model, lane occupancy is used as the state space of the optimization model, and the signal control scheme is dynamically adjusted according to real-time traffic flow information to improve the space-time resource utilization efficiency of traffic flow.

Benefits of technology

It effectively improves the efficiency of time and space resource utilization at the intersection, reduces traffic congestion and parking delays, and can achieve better signal control effects in complex dynamic traffic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220439A_ABST
    Figure CN120220439A_ABST
Patent Text Reader

Abstract

The invention discloses an asymmetric traffic flow signal control method based on deep reinforcement learning, belongs to the field of asymmetric traffic flow control, and aims to establish a corresponding phase scheme in a targeted manner by analyzing the characteristics and types of asymmetric traffic flows, break through a conventional eight-phase four-stage signal control mode, and improve the control precision of the asymmetric traffic flows. And the new phase scheme is used as the motion selection of a deep reinforcement learning signal control model, the motion selection in deep reinforcement learning is improved, the improvement strategy takes a lane occupancy matrix as an optimization model state space, and the dimension of the state space is reduced. A deep reinforcement learning algorithm is adopted, and a signal control decision is iteratively improved in a complex and dynamic traffic environment through interaction of an intelligent agent and the environment and trial and error learning, so that the traffic signal control method based on the deep reinforcement learning algorithm can obtain a better control effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of asymmetric traffic flow signal control, and particularly relates to an asymmetric traffic flow signal control method based on deep reinforcement learning. Background Art

[0002] In traditional traffic signal control theory, the arrival of vehicle flows at each approach is assumed to be uniform or follow a certain distribution. Under this assumption, the asymmetry of oncoming vehicle flows in each cycle is generally small, so the use of traditional eight phases is effective. However, in the actual road operation scenario, the development of intersections is dynamic and complex, and traffic flows have characteristics such as non-linearity, dynamics, uncertainty, and self-organization. Due to the irregularity of traffic flows and the uneven arrival distribution of vehicles, the traffic flow at intersections will produce an asymmetric phenomenon.

[0003] Facing the non-linear, complex and dynamic traffic system, there are bottlenecks in traditional signal control in adapting to the dynamic changes of traffic flows and implementing feedback control. Pan Yiyong summarized the formation reasons of tidal traffic flows, proposed a new lane scheme design and operation organization method, and verified the superiority of the tidal lane scheme based on vissim simulation. Zhang Taiwen, Zhang Cunbao, etc. proposed to set reverse dynamic variable lanes to alleviate the situation that traditional signal control schemes at intersections are difficult to meet asymmetric traffic flows, and supplemented the judgment conditions for the reverse variable lane switching control method. Jiang et al. analyzed the threshold changes of determined parameters at asymmetric traffic flow intersections and the influence of asymmetric signal cycles on the vehicle operation states at intersections. Kumarasamy et al. introduced an innovative solution by integrating decentralized graph-based multi-agent reinforcement learning (DGMARL) with digital twin to enhance traffic signal optimization.

[0004] However, the existing literature mainly studies the optimization of randomly arriving traffic flows in low traffic volume environments and the optimization of stable asymmetric traffic flows on the oncoming approaches of signal intersections. When the traffic volumes on the oncoming approaches are basically symmetric at a small level, the phase-symmetric signal control scheme can adjust itself to solve the problem. This is mainly based on the assumptions of traffic signal control theory. When the asymmetry is severe and unstable, it will cause unbalanced utilization of intersection spatio-temporal resources and increase stops and delays. In addition, most of the state spaces of current deep reinforcement learning signal control models are high-dimensional microscopic traffic flow parameters - the vehicle positions and speed matrices on the approaches are used as the state space. These high-dimensional data increase the computational complexity, and when the traffic flow fluctuates violently, it will cause the imbalance of intersection spatio-temporal resource utilization.

[0005] Therefore, how to dynamically generate a signal control scheme according to the real-time traffic flow operation state at intersections has important theoretical significance and practical value for improving the utilization efficiency of intersection spatio-temporal resources and alleviating traffic congestion. Summary of the Invention

[0006] In order to solve the defects such as poor control effect for asymmetric traffic flow in the prior art, an asymmetric traffic flow signal control method based on deep reinforcement learning is proposed. A deep reinforcement learning signal control model is constructed, with the lane occupancy rate as the optimization model state space. Signal adjustments are made to the traffic flow operation status according to the real-time collected traffic flow information, effectively improving the actual control efficiency and road traffic efficiency.

[0007] The present invention is implemented by the following technical solutions: An asymmetric traffic flow signal control method based on deep reinforcement learning, comprising the following steps: Step 1: Collect the traffic flow information of each lane at the intersection, and calculate the vehicle arrival rate of each approach and the saturation flow rate in each direction according to the traffic flow data; Step 2: Construct a deep reinforcement learning signal control model and train and optimize it; The deep reinforcement learning algorithm DQN is used as the core framework of the signal control model, and the signal control model takes the occupancy rate as the state space; The signal control model includes a state space selection module, a joint action plan formulation module, and a feedback adjustment module; The state space selection module takes the occupancy rate as the state space, and the occupancy rate refers to the proportion of the range detected by the radar-vision fusion occupied by the vehicles in each approach; The joint action plan formulation module analyzes the asymmetry coefficient of the intersection traffic flow, then divides the traffic flow state, and formulates the corresponding phase diagram and green light release time; The feedback adjustment module guides the signal machine to seek the best action in signal control according to whether the current road traffic condition is improved and the degree of improvement after the joint action plan is executed, so as to achieve the optimal control effect; Step 3: Based on the signal control model, make signal adjustments to the traffic flow operation status according to the real-time collected traffic flow information; The signal control machine judges the traffic flow state category according to the asymmetry coefficient of the intersection traffic flow at the previous moment and selects a suitable joint action for phase switching and adjusting the green light time; at this time, the intersection environment outputs the state at the next moment and the reward value obtained by this action according to this action. On this basis, the signal control machine makes a selection of the action to be output at the next moment, thus forming a cycle.

[0008] Compared with the prior art, the advantages and positive effects of the present invention are as follows: According to the characteristics of the dynamic and irregular nature of the actual traffic flow at intersections, this solution takes the state of asymmetric traffic flow into account in the deep reinforcement learning signal control model, supplements the action selection of the deep reinforcement learning signal control model, enables the model to make appropriate signal adjustments according to the actual traffic flow conditions, avoids the waste of spatio-temporal resources in the actual intersection operation, enables it to adjust the appropriate phase and green light duration according to the traffic state, and also has a good control effect under special states of asymmetric traffic flow, effectively improving the actual control efficiency and road traffic efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is a schematic diagram of the principle of deep reinforcement learning in an embodiment of the present invention; Figure 2 It is a flowchart of the joint action design for asymmetric traffic flow in an embodiment of the present invention; Figure 3 It is an action diagram corresponding to the first type of traffic flow in an embodiment of the present invention, and the east-west left-turn traffic flow and the east-west straight-through traffic flow are released symmetrically respectively; Figure 4 It is an action diagram corresponding to the second type of traffic flow in an embodiment of the present invention; among them, (a) is a phase scheme of symmetrically releasing the east-west straight-through and east-west left-turn traffic flows; (b) is a phase scheme of independently releasing a single inlet in the east-west direction; Figure 5 It is an action diagram corresponding to the third type of traffic flow in an embodiment of the present invention; after the east-west left-turn traffic flow is symmetrically released first, then a single inlet release of east left-turn + east straight-through or west left-turn + west straight-through is respectively adopted, and finally the east-west straight-through traffic flow is symmetrically released; Figure 6 It is a corresponding diagram of the intersection and arrival rate in an embodiment of the present invention; Figure 7 It is a framework diagram of the deep reinforcement learning signal control model in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0010] In order to more clearly understand the above objects, features and advantages of the present invention, the following further describes the present invention with reference to the accompanying drawings and embodiments. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0011] An embodiment is a method for asymmetric traffic flow signal control based on deep reinforcement learning according to the embodiment of the present invention, including the following steps: Step 1, collect traffic flow information of each lane at the intersection; Use a combined radar and vision device to collect the vehicle arrival traffic volume Q, vehicle queue length L, and vehicle queue count N for each lane at the intersection, as well as the vehicle body length l, vehicle speed v, and vehicle lost time T for each lane; calculate the vehicle arrival rate q for each approach and the saturation flow rate s for each direction based on the collected traffic flow data. Step 2: Construct a signal control model based on deep reinforcement learning and train and optimize it. As Figure 7 Shown in the model framework diagram, in this embodiment, the deep reinforcement learning algorithm (DQN) is used as the main algorithm for the signal control model. This model uses the occupancy rate as the state space. The deep reinforcement learning algorithm (DQN) is the main algorithm for the signal control model, and this model uses the occupancy rate as the state space. Deep reinforcement learning is based on the basic framework of reinforcement learning. Through processes such as observing the environmental state space, taking actions, and receiving rewards, it learns the optimal strategy. Deep reinforcement learning adds a neural network to the basic framework of reinforcement learning. Reinforcement learning mainly observes the environmental state space, receives rewards, and makes action selections, while the neural network guides the decision-making of future action selections by summarizing the patterns in existing data and experiences, enabling the model to learn the optimal strategy.

[0012] The signal control model includes a state space selection module, a joint action plan formulation module, and a feedback adjustment module. The state space selection module uses the occupancy rate as the state space, where the occupancy rate refers to the proportion of the radar-vision fusion detection range occupied by vehicles in each approach. The joint action plan formulation module analyzes the asymmetry coefficient of the intersection traffic flow, then divides the traffic flow state, and formulates the corresponding phase diagram and green light release time. The feedback adjustment module guides the signal machine to seek the best action in signal control based on whether the current road traffic condition is improved and the degree of improvement after the joint action plan is executed, in order to achieve the optimal control effect.

[0013] Step 3: Based on the signal control model, make signal adjustments to the vehicle flow operation status according to the real-time collected traffic flow information. Combined with Figure 1 and Figure 2 Shown, the signal controller determines the traffic flow state category based on the intersection traffic flow asymmetry coefficient at the previous moment and selects an appropriate joint action for phase switching and green light time adjustment; at this time, the intersection environment outputs the state at the next moment and the reward value obtained by this action. Based on this, the signal controller makes an action selection for the output at the next moment. In this way, a cycle is formed. Specifically: Step 31: State space selection.

[0014] Taking the proportion (occupancy rate) of the detection range of the radar-vision fusion occupied by vehicles in each approach as the state space of the signal control model, the occupancy rate of a certain lane is expressed as follows: ; where, o n represents the occupancy rate of lane n, l j represents the vehicle length (m) of the j-th vehicle, represents the total number of vehicles on lane n, L d represents the length (m) of the radar-vision detection area.

[0015] The state space represented by the radar-vision integrated machine detection is: ; where, O is the state space matrix, and the elements in the matrix are the occupancy rate values of each lane.

[0016] Step 32: Develop a joint action plan; In view of the dynamic and irregular characteristics of the intersection traffic flow, this embodiment makes supplements to the actions of the deep reinforcement learning. First, analyze the asymmetry coefficient of the intersection traffic flow, and then divide the traffic flow state to formulate the corresponding phase diagram and green light release time. The flow chart is as Figure 2 shown.

[0017] 1. Calculate the traffic flow asymmetry coefficient: The phase traffic flow asymmetry coefficient is defined as: (3); where, ac i is the traffic flow asymmetry coefficient of the i-th phase; Q i1 is the traffic volume (pcu / cycle) arriving at a single lane of the first lane group in the i-th phase; Q i2 is the traffic volume (pcu / cycle) arriving at a single lane of the second lane group in the i-th phase.

[0018] Whether the phase traffic flow is symmetric depends on whether the phase traffic flow asymmetry coefficient exceeds the threshold, that is (4); where, ps i is the symmetric state of the traffic flow in the i-th phase, ps i =0 indicates that the phase traffic flow is basically symmetric, ps i ≠0 indicates that the phase traffic flow is asymmetric, and the positive and negative indicate the direction of traffic flow asymmetry. Let ps1, ps2, ps3, ps4 be the symmetric states of the traffic flows in the east-west left-turn, east-west straight, north-south left-turn, and north-south straight phases respectively; ac i0 is the threshold of the traffic flow asymmetry in the i-th phase, which is determined according to experience or experiments.

[0019] 2. Classify the types of combined actions and formulate a combined action plan by judging the traffic flow state; In specific implementation, the intersection can be divided into two cases: east-west left turn - east-west straight and north-south left turn - north-south straight for analysis (since the right-turn traffic flow does not conflict with other traffic flows, this plan does not consider the right-turn traffic flow). The following takes the east-west left turn and east-west straight as an example to illustrate: According to formula (4), the following states can be calculated and analyzed for the oncoming traffic flow in the east-west direction of the intersection: The first type of traffic flow: When ps1 = 0 and ps2 = 0, the east-west left-turn traffic flow and the east-west straight traffic flow are released symmetrically respectively, as Figure 3 shown. The combined action set is A1 = [(p1, g1), (p2, g2)], where p1 is the east-west left-turn phase, p2 is the east-west straight phase, and the corresponding release times t are g1 and g2 respectively. g1 and g2 are the initially set effective green light times under normal circumstances.

[0020] The second type of traffic flow: When ps1 × ps2 = 0, but ps1 ≠ 0 or ps2 ≠ 0; Figure 4 As shown, if [max(q 11 , q 12 ) / s1 + max(q 21 , q 22 ) / s2] < [max(q 11 / s1, q 21 / s2) + max(q 12 / s1, q 22 / s2)], where q 11 , q 12 represent the vehicle arrival rates (pcu / s) in two directions of the east-west left-turn phase respectively; q 21 , q 22 represent the vehicle arrival rates (pcu / s) in two directions of the east-west straight phase respectively; s1 and s2 are the saturation flows (pcu / s) of the east-west left-turn and east-west straight phases respectively. Then the left-turn traffic flow and the straight traffic flow respectively adopt the symmetric release phase plan shown in (a), and the release times are g1 and g2 respectively (the same as the first type of traffic flow). Otherwise, adopt the single-entrance independent release phase plan shown in (b). The combined action set of the second type of traffic flow is A2 = [(p1, g1), (p2, g2)] or A2 = [(p5, t 东 ), (p6, t 西 )], p5 is the east-left-straight phase, p6 is the west-left-straight phase, and the release times are t 东 = N 东 / s1′, t 西 = N西 / s2′, s1′, and s2′ are the saturation flow rates (pcu / s) when the east and west approach roads are released separately, and N 东 is the number of queuing vehicles at the east approach, and N 西 is the number of queuing vehicles at the west approach.

[0021] The third type of traffic flow: When ps1 × ps2 = 1, as Figure 5 shown, first release the left-turn traffic symmetrically until the saturation flow rate of the approach with less traffic volume disappears, and then turn off the left-turn traffic signal of that approach. The east left-turn release time t e-0 = N e-0 / s e-0 , and the west left-turn release time t w-0 = N w-0 / s w-0 , and the east-west straight-through release time t = N 直 / s2, where N e-0 is the number of queuing vehicles for east left-turn, s e-0 is the saturation flow rate of east left-turn, N w-0 is the number of queuing vehicles for west left-turn, s w-0 is the saturation flow rate of west left-turn, N 直 is the number of queuing vehicles for east-west straight-through; then release a single approach with left-turn + straight-through, and its release time Δt1 is to maximize the comprehensive traffic benefit of the intersection: The calculation process of the release time Δt1 is as follows: (1) After implementing the Figure 5 shown combined phase plan, the reduction in vehicle delay for the east-west straight-through phase is: (5); Since the traffic signal of the direction with a small left-turn arrival rate is turned off in advance, the increase in vehicle delay for the subsequent arriving vehicles is: (6); The total change in delay is ; (2) The comprehensive benefit calculation model for extending the current phase by Δt1 seconds at the intersection is: ; ; ; ; ; Where: R1 is the green light time of the east-west left-turn phase if the east-west left-turn phase is cut off, and the next red light display time of the east-west left-turn phase (s); K1 is the number of vehicles (pcu) that can pass through the stop line at the approach with a large traffic volume in the east-west left-turn phase during the time interval Δt1; f1 is the equivalent cost of a single vehicle stop; N is the current number of queuing vehicles (pcu) at all approach lanes of other phases; M is the number of vehicles (pcu) that can decelerate to a stop at all approach lanes of other phases during the time interval Δt1; Y1, Y2, Y3, and Y4 are the green light intervals (s) for the east-west left-turn to east-west straight, east-west straight to north-south left-turn, north-south left-turn to north-south straight, and north-south straight to east-west left-turn, respectively; N ij (0) is the number of queuing vehicles (pcu) at each approach lane of other phases when calculating ΔD12. q 31 ,q 32 represent the vehicle arrival rates (pcu / s) in two directions of the east-west left-turn phase, respectively; q 41 ,q 42 represent the vehicle arrival rates (pcu / s) in two directions of the east-west straight phase, respectively; s3 and s4 are the saturation flows (pcu / s) of the north-south left-turn and north-south straight phases, respectively; Then the total benefit of extending the intersection by Δt1 seconds is: ; This extension is meaningful only when D1≥0. That is, find Δt1 such that D1≥0, and at this time, Δt1 is the extended time for a single release of left-turn + straight.

[0022] Finally, the straight traffic flow is released symmetrically until the saturation flow rate of the straight traffic flow disappears, and then the straight phase ends.

[0023] As Figure 5 shown, the combined action set is A3 = [(p1, t e-0 = N e-0 / s e-0 )-(p6, Δt1)-(p2, t = N 直 / s2)), (p1, t w-0 = N w-0 / s w-0 )-(p5, Δt1)-(p2, t = N 直 / s2)].

[0024] Fourth type of traffic flow: When ps1×ps2 = -1, a phase plan of releasing the left-turn traffic flow and the straight traffic flow symmetrically is adopted (the plan is the same as Figure 3 , but the corresponding release times are different). In the case of the fourth type of traffic flow, the required green light time for the left-turn phase is t 左 = [min(q11 , q 12 ) + |q 11 -q 12 | / 2], the green light time required for the straight-through phase is t 直 = [min(q 21 , q 22 ) + |q 21 -q 22 | / 2], the combined action set is A4 = [(p1, t 左 ), (p2, t 直 )].

[0025] The combined action situation and principle in the north-south direction are the same as those discussed in the above east-west direction, and will not be elaborated here.

[0026] Step 33: Select a reward function to conduct feedback verification on the formulated combined action plan; The reward function is the feedback of the traffic signal controller on the current road traffic conditions after selecting a combined action. It can reflect whether the traffic state has improved and the degree of improvement after implementing the phase plan of the combined action in the model, and guide the signal controller to seek the best action in signal control, so as to achieve the optimal control effect.

[0027] In this embodiment, the standardized average vehicle speed, the standardized average lost time, and the standardized maximum average queue length are selected as the components of the reward function.

[0028] The calculation process of the components of the reward function for each import lane in the radar-vision fusion detection area is as follows: (1) Standardized average vehicle speed ; In the formula, is the standardized average vehicle speed of lane l; is the average speed of all vehicles currently in lane l (m·s -1 ); is the maximum speed of all vehicles in lane l (m·s -1 ); is the minimum speed of all vehicles in lane l (m·s -1 ).

[0029] (2) Standardized maximum average queue length ; In the formula, is the standardized maximum average queue length of lane l; is the average value of the maximum queue length in lane l during the detection interval (m); is the maximum value (m) of all the maximum queue lengths of lane l within the detection interval; is the minimum value (m) of all the maximum queue lengths of lane l within the detection interval.

[0030] (3) Normalized average lost time ; In the formula, r3 l is the normalized average lost time of lane l; is the average lost time (s) of all the vehicles passing through lane l within the detection interval, where the lost time is defined as the difference between the actual passing time of the vehicle through the intersection and the time for the vehicle to pass through the intersection without stopping at the speed when it just enters the detection area; is the maximum lost time (s) of all the vehicles passing through lane l within the detection interval; is the minimum lost time (s) of all the vehicles passing through lane l within the detection interval.

[0031] Let the set of straight and left-turn lane numbers in the four directions of the intersection be Lane, then the reward function of the whole intersection is: ; In the formula, F is the reward function of the agent during training, R t+1 is the reward value of the agent at time t + 1, R t is the reward value obtained by the agent at time t.

[0032] By analyzing the characteristics and types of asymmetric traffic flow, the present invention specifically establishes a corresponding phase plan, breaks the conventional eight-phase four-stage signal control mode, and uses the new phase plan as the action selection of the deep reinforcement learning signal control model, thereby improving the action selection in deep reinforcement learning. The improvement strategy uses the lane occupancy matrix as the optimization model state space, reducing the dimension of the state space. By using the deep reinforcement learning algorithm, through the interaction between the agent and the environment, and using trial-and-error learning, the signal control decision is iteratively improved in a complex and dynamic traffic environment, so that the traffic signal control method based on the deep reinforcement learning algorithm can achieve better control effects.

[0033] The above are only the preferred embodiments of the present invention, and the present invention is not limited to other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. An asymmetric traffic flow signal control method based on deep reinforcement learning, characterized in that It includes the following steps: Step 1: Collect the traffic flow information of each lane at the intersection, and calculate the vehicle arrival rate of each approach lane and the saturation flow rate in each direction according to the traffic flow data; Step 2: Construct a deep reinforcement learning signal control model and train and optimize it; Adopt the deep reinforcement learning algorithm DQN as the core framework of the signal control model, and the signal control model takes the occupancy rate as the state space; Step 3: Based on the signal control model, make signal adjustments to the traffic flow operation status according to the real-time collected traffic flow information; The signal controller judges the traffic flow state category according to the traffic flow asymmetry coefficient at the previous moment of the intersection and selects an appropriate joint action to perform phase switching and adjust the green light time; at this time, the intersection environment outputs the state at the next moment and the reward value obtained by this action according to this action. On this basis, the signal controller makes a selection of the action output at the next moment, thus forming a cycle.

2. The asymmetric traffic flow signal control method based on deep reinforcement learning according to claim 1, characterized in that: The traffic flow information includes the vehicle arrival volume Q, the vehicle queue length L and the vehicle queue number N, the vehicle body length l, the vehicle speed v and the lost time T of the vehicle on each lane.

3. The asymmetric traffic flow signal control method based on deep reinforcement learning according to claim 1, characterized in that: In the said Step 2, the signal control model includes a state space selection module, a joint action plan formulation module and a feedback adjustment module; The state space selection module takes the occupancy rate as the state space, and the occupancy rate refers to the proportion of the range detected by the radar-vision fusion occupied by the vehicles on each approach lane; The joint action plan formulation module analyzes the traffic flow asymmetry coefficient of the intersection, then divides the traffic flow state, and formulates the corresponding phase diagram and green light release time; The feedback adjustment module guides the signal machine to seek the best action in signal control according to whether the current road traffic condition is improved and the degree of improvement after the joint action plan is executed, so as to achieve the optimal control effect.

4. The asymmetric traffic flow signal control method based on deep reinforcement learning according to claim 1, wherein: The said Step 3 specifically includes the following steps: Step 31: Select the state space, and the state space is expressed as: ; Among them, O is the state space matrix, and the element o n in the matrix is the occupancy value of each lane; Step 32: Formulate a joint action plan; (1) First, calculate the phase traffic flow asymmetry coefficient and , where ac i is the traffic flow asymmetry coefficient of the i-th phase; Q i1 is the traffic volume arriving at a single lane of the first lane group of the i-th phase during a signal cycle; Q i2 is the traffic volume arriving at a single lane of the second lane group of the i-th phase during a signal cycle; Whether the phase traffic flow is symmetric depends on whether the traffic flow asymmetry coefficient of this phase exceeds the threshold, that is ; where, ps i is the symmetric state of the traffic flow in the i-th phase, and ps i = 0 indicates that the traffic flow of the phase is symmetric, and ps i ≠ 0 indicates that the traffic flow of the phase is asymmetric, and the positive or negative sign indicates the direction of the traffic flow asymmetry, and ac i0 is the threshold of the traffic flow asymmetry in the i-th phase; (2) Divide the types of joint actions, and formulate a joint action plan by judging the traffic flow state; For the case of east-west left turn - east-west straight, the analysis is as follows: The first type of traffic flow: when ps1 = 0 and ps2 = 0; The east-west left-turn traffic flow and the east-west straight traffic flow are released symmetrically respectively, and the joint action set is A1 = [(p1, g1), (p2, g2)], where p1 is the east-west left-turn phase, p2 is the east-west straight phase, and the corresponding release times t are g1 and g2 respectively; The second type of traffic flow: when ps1 × ps2 = 0, but ps1 ≠ 0 or ps2 ≠ 0; If [max(q 11 ,q 12 ) / s1 + max(q 21 ,q 22 ) / s2] < [max(q 11 / s1, q 21 / s2) + max(q 12 / s1, q 22 / s2)], where q 11 , q 12 respectively represent the vehicle arrival rates in two directions of the east-west left-turn phase; q 21 , q 22 respectively represent the vehicle arrival rates in two directions of the east-west straight-ahead phase; s1 and s2 are the saturation flows of the east-west left-turn and east-west straight-ahead phases respectively; then the left-turn traffic flow and the straight-ahead traffic flow adopt a symmetric release phase plan; otherwise, adopt a single-entrance independent release phase plan; The combined action set of the second type of traffic flow is A2 = [(p1, g1), (p2, g2)] or A2 = [(p5, t 东 ), (p6, t 西 )], where p5 is the east straight-through phase, p6 is the west straight-through phase, and the release times are t 东 = N 东 / s1′, t 西 = N 西 / s2′, s1′ and s2′ are the saturation flow rates when the east and west approaches are released separately, N 东 is the number of queued vehicles at the east approach, N 西 is the number of queued vehicles at the west approach; The third type of traffic flow: when ps1 × ps2 = 1; First, release the left-turn traffic flow symmetrically until the saturation flow rate of the approach with less traffic volume disappears, and then turn off the left-turn traffic signal at that approach. The east left-turn release time t e-0 =N e-0 / s e-0 , and the west left-turn release time t w-0 =N w-0 / s w-0 . The east-west straight-through release time t = N 直 / s2. Here, N e-0 is the number of queued vehicles for east left-turn, s e-0 is the saturation flow rate of east left-turn, N w-0 is the number of queued vehicles for west left-turn, s w-0 is the saturation flow rate of west left-turn, N 直 is the number of queued vehicles for east-west straight-through; then release traffic for a single approach with left-turn + straight-through. The release time Δt1 is such that the comprehensive traffic benefit of the intersection is maximized. Finally, release the straight-through traffic flow symmetrically until the saturation flow rate of the straight-through traffic flow disappears, ending the straight-through phase; the combined action set is A3 = [(p1, t e-0 =N e-0 / s e-0 )-(p6, Δt1)-(p2, t = N 直 / s2)),(p1, t w-0 =N w-0 / s w-0 )-(p5, Δt1)-(p2, t = N 直 / s2)]; The fourth type of traffic flow: when ps1 × ps2 = -1; Adopt a phase plan where the left-turning traffic flow and the straight-through traffic flow are released symmetrically respectively, but the corresponding release times are different. In the case of the fourth type of traffic flow, the required green-light time for the left-turn phase is t 左 =[min(q 11 ,q 12 )+|q 11 -q 12 | / 2], and the required green-light time for the straight-through phase is t 直 =[min(q 21 ,q 22 )+|q 21 -q 22 | / 2]. The combined action set is A4 = [(p1, t 左 ), (p2, t 直 )]; The joint action situation and principle of the north-south left turn - north-south straight direction are the same as those of the above east-west direction.

5. The method for asymmetric traffic flow signal control based on deep reinforcement learning according to claim 4, wherein: For the third type of traffic flow, the calculation process of the release time Δt1 is as follows: (1) The reduction in the vehicle delay of the straight phase in the east-west direction is: ; Due to the early closing of the traffic signal in the direction with a small left-turn arrival rate, the increase in the delay of the subsequent arriving vehicles is: ; The total delay change is ; (2)The comprehensive benefit calculation model for extending the current phase by Δt1 seconds at the intersection is as follows: ; ; ; ; ; Where: R1 is the green light time of the east-west left-turn phase when cutting off, and the next red light display time of the east-west left-turn phase; K1 is the number of vehicles that can pass through the stop line at the approach with a large traffic volume in the east-west left-turn phase within the time interval Δt1; f1 is the equivalent cost of one vehicle stop; N is the current number of queuing vehicles in all approach lanes of other phases; M is the number of vehicles that can decelerate to a stop in all approach lanes of other phases within the time interval Δt1; Y1, Y2, Y3, Y4 are the green light intervals when switching from east-west left-turn to east-west straight, from east-west straight to north-south left-turn, from north-south left-turn to north-south straight, and from north-south straight to east-west left-turn respectively; N ij (0) is the number of queuing vehicles in each approach lane of other phases when calculating ΔD12; q 31 ,q 32 respectively represent the vehicle arrival rates in two directions of the east-west left-turn phase; q 41 ,q 42 respectively represent the vehicle arrival rates in two directions of the east-west straight phase; s3 and s4 are the saturation flows of the north-south left-turn and north-south straight phases respectively; Then the total benefit of extending the intersection by Δt1 seconds is: ; Find Δt1 such that D1 ≥ 0. At this time, Δt1 is the extended time for a single release of left turn + straight.

6. The asymmetric traffic flow signal control method based on deep reinforcement learning according to claim 5, wherein: Step 3 described above further includes: Step 33: Select a reward function to conduct feedback verification on the formulated joint action plan. The reward function reflects whether the traffic state has improved and the degree of improvement after implementing the phase plan in the joint action of the signal lights in the traffic signal control model, guiding the signal machine to optimize the feedback of the linkage action in signal control. Among them, the standardized average vehicle speed, the standardized average lost time, and the standardized maximum average queue length are used as the constituent elements of the reward function.

Citation Information

Patent Citations

  • Signal lamp control method and system based on deep intensive learning and storage medium

    CN109472984A

  • Deep reinforcement learning traffic signal control poisoning attack method based on Trojan horse attack

    CN115426150A

  • Traffic signal control method based on reinforcement learning

    CN115705771A

  • Multi-intersection traffic signal control method based on deep reinforcement learning

    CN117012044A