Networked intersection ecological guidance cooperative control method based on deep Q network

By deploying intelligent signal control equipment and deep Q network models at network-connected intersections, dynamically optimizing the phase release strategy, the problems of traffic efficiency and road rights allocation at network-connected intersections in the existing technology are solved, and more efficient traffic flow management and energy consumption reduction are achieved.

CN120199073AActive Publication Date: 2025-06-24JILIN UNIVERSITY

Patent Information

Application Number
CN202510358354.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-24
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The prior art cannot achieve the optimal traffic efficiency and right of road allocation within networked intersections.

Method used

Using the ecological guidance collaborative control method of network-connected intersections based on deep Q network, through the communication between intelligent signal control equipment, roadside units and on-board units, the time when the fleet arrives at the parking line is predicted, the echelons are dynamically divided, the deep reinforcement learning model is constructed, and the phase release strategy is optimized.

Benefits of technology

It improves the traffic efficiency of the Internet Interchange intersection, ensures fair road rights in all directions, and reduces the energy consumption of the vehicle's travel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199073A_ABST
    Figure CN120199073A_ABST
Patent Text Reader

Abstract

The invention discloses a deep Q network-based networked intersection ecological guidance cooperative control method, and relates to the field of traffic information engineering and control. According to the invention, the problem that the optimal passing efficiency and road right distribution in the intersection cannot be obtained in the prior art is solved. The invention provides a method for carrying out active cooperation on a networked vehicle and an intelligent annunciator, carrying out behavior prediction on an upstream vehicle and responding to a passing demand. The annunciator generates a signal control strategy based on a phase sequence and phase coupling method, considers that a signal state in a future time period participates in control optimization of a current stage so as to avoid short view, and uses DQN optimization solution; the minimum energy consumption is used as target optimization control, decentralized control is adopted to improve the control effect and prediction precision, and an energy consumption model is constructed to evaluate and optimize the vehicle trajectory control effect. Compared with a traditional method, the traffic efficiency of the network connection intersection is improved, the road right fairness of all the flow directions is guaranteed, and the vehicle travel energy consumption is reduced. The method can be applied to network connection intersection control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of traffic information engineering and control, and particularly relates to an ecological guidance collaborative control method for connected intersections based on deep Q-network. Background Art

[0002] As a key node of the urban traffic network, intersections have long faced problems such as low traffic efficiency and frequent congestion. With the development of intelligent connected technology, by using V2I communication, a centralized autonomous intersection management controller is installed at intersections to coordinate the movement of CAVs, ensure the efficient and conflict-free passage of CAVs, and optimize the riding comfort (Zhao, Rui; Li, Yun; Wang, Kui; Fan, Yuze; Gao, Fei; Gao, Zhenhai. Centralized Cooperation for Connected Autonomous Vehicles at Intersections by Safe Deep Reinforcement Learning [J]. IEEE Transactions on Mobile Computing, 2024, Vol. 23(12): 1-18). In an ideal fully connected autonomous driving scenario, vehicles can autonomously coordinate the passing order through information interaction and theoretically do not need to rely on traditional signal control. However, in practical applications, signal control is still required to ensure the fair distribution of road rights among multiple traffic participants such as motor vehicles, non-motor vehicles, and pedestrians. Therefore, designing a collaborative guidance strategy for connected vehicles and signal systems is of great practical significance in improving the operation efficiency of intersections.

[0003] CAV trajectory optimization has been widely concerned in recent years. To improve the overall performance of the system, various trajectory optimization methods have been proposed, including the decentralized control model (DM). In DM, each vehicle is regarded as an autonomous agent, which determines its own control strategy based on the information sensed or received from other vehicles and roadside units to maximize its own performance. The literature (Rios-torres, J., Malikopoulos, A.A., 2017. A Survey on the Coordination of Connected and Automated Vehicles at Intersections and Merging at Highway On-Ramps. IEEE Trans. Intell. Transp. Syst. 18, 1066–1077.) shows that sharing intersection information with drivers in advance can reduce vehicle energy consumption by up to 20%. Therefore, if vehicles can plan their speed management in advance for intersections, energy loss can be reduced on the premise of ensuring driving safety.

[0004] In summary, although certain achievements have been made by existing methods, the passing efficiency and right-of-way allocation within intersections of existing methods still cannot achieve the optimal state. Therefore, proposing a new method is an urgent problem to be solved currently. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem that the existing technology cannot obtain the best passing efficiency and right-of-way allocation within intersections, and a cooperative control method for ecological guidance of connected intersections based on a deep Q-network is proposed.

[0006] The technical solution adopted by the present invention to solve the above technical problems is: a cooperative control method for ecological guidance of connected intersections based on a deep Q-network, and the method specifically includes the following steps:

[0007] Step 1: Deploy intelligent signal control devices, roadside units, and in-vehicle units. The roadside units send the collected information to the intelligent signal control devices;

[0008] And establish a communication link between the roadside unit and the in-vehicle unit;

[0009] Step 2: For each approach within the current intersection, the intelligent signal control device predicts the arrival time of the vehicle platoon at the stop line according to the received information;

[0010] Step 3: Dynamically divide the vehicles on each approach of the current intersection according to the prediction result of the arrival time of the vehicle platoon at the stop line, and divide the vehicles into two parts: a release echelon and a detention echelon;

[0011] Step 4: Intensively process the phases of vehicles in each flow direction based on the number of vehicles in the release echelons of each flow direction, construct the state of the deep reinforcement learning model according to the intensive processing, and output the actual phase release sequence and the start time of the green light for each phase through the deep reinforcement learning model;

[0012] Step 5: Based on the prediction result in Step 2 and the output of the deep reinforcement learning model, determine whether to maintain the phase release plan currently output by the deep reinforcement learning model;

[0013] If maintaining the phase release plan currently output by the deep reinforcement learning model, perform vehicle movement planning for each phase according to the start time of the green light for each phase output by the deep reinforcement learning model and release the first phase in the phase release plan, then obtain the prediction result of the arrival time of the vehicle platoon at the stop line in each phase based on the vehicle movement planning result, and then use the prediction result to return to execute Step 3;

[0014] If not maintaining the phase release plan currently output by the deep reinforcement learning model, perform vehicle movement planning for each phase according to the start time of the green light for each phase output by the deep reinforcement learning model, then obtain the prediction result of the arrival time of the vehicle platoon at the stop line in each phase based on the vehicle movement planning result, and then use the prediction result to return to execute Step 3.

[0015] Furthermore, multi-source sensing devices are arranged in the roadside unit, and the multi-source sensing devices are used to collect the instantaneous speed, the distance from the approach stop line, and the running time of each connected vehicle in real time; and send the information collected in real time and the real-time traffic flow data to the intelligent signal control device.

[0016] Furthermore, the specific process of Step 2 is as follows:

[0017] For the vehicles in any one flow direction within the current intersection, take the vehicle exiting the merging area of the previous intersection as the starting point of the control area, allow the vehicle to enter the control area and drive to the maximum speed with the maximum acceleration, then the predicted value of the arrival time of the nth vehicle at the stop line is:

[0018]

[0019] Among them, represents the predicted value of the arrival time of the nth vehicle at the stop line;

[0020] represents the control start time of the nth vehicle, that is, the time when the nth vehicle enters the control area;

[0021] represents the speed of the nth vehicle at moment;

[0022] represents the maximum allowable acceleration of the nth vehicle;

[0023] represents the predicted value of the arrival time of the (n - 1)-th vehicle at the stop line;

[0024] represents the speed of the n-th vehicle at moment;

[0025] represents the speed of the n-th vehicle at moment;

[0026] τ n is the headway of the n-th vehicle and the vehicle in front;

[0027] h n is the distance between the n-th vehicle and the vehicle in front;

[0028] is the body length of the vehicle in front of the n-th vehicle;

[0029] D represents the distance of the n-th vehicle relative to the stop line at the current moment;

[0030] When n = 0, denote When the n-th vehicle travels at the maximum acceleration and accelerates to v n,max at the stop line, the predicted value of the arrival time of the n-th vehicle at the stop line is rewritten as:

[0031]

[0032] where, represents the distance between the n-th vehicle and the stop line at moment;

[0033] represents the speed of the n-th vehicle when entering the control area.

[0034] Furthermore, in the third step, the specific process of dynamic echelon division of vehicles is as follows:

[0035] (1) The spatio-temporal trajectory of the released echelon vehicles satisfies:

[0036] h n ′ < d c and g min ≤ t duration ≤ g max

[0037] where, h n ′ represents the headway between the n-th vehicle and the (n - 1)-th vehicle;

[0038] d c represents the critical spacing;

[0039] gmin Indicates the shortest green light time of the phase;

[0040] g max Indicates the maximum green light time of the phase;

[0041] t duration Indicates the sum of Δt and the driving time of the nth vehicle in the intersection. Δt represents the difference between the time when the nth vehicle leaves the stop line and the time when the first vehicle leaves the stop line;

[0042] (2) The set of vehicles N' in the detention echelon is N' = M - N, where N represents the set of vehicles in the release echelon and M represents the set of all vehicles in the current traffic flow direction.

[0043] Furthermore, the constraints satisfied when calculating the predicted value of the arrival time of the nth vehicle at the stop line are:

[0044] ① Speed constraint:

[0045]

[0046] Among them, Indicates the speed of the nth vehicle at time t;

[0047] ② Acceleration constraint:

[0048]

[0049] Among them, Indicates the maximum acceleration that the vehicle is allowed to adopt;

[0050] Indicates the maximum deceleration that the vehicle is allowed to adopt;

[0051] Indicates the acceleration of the nth vehicle at time t;

[0052] ③ Phase constraint:

[0053]

[0054] Among them, g e = t e t e Indicates the end time of the vehicle phase in the current traffic flow direction;

[0055] Let Indicates the time when the nth vehicle leaves the stop line.

[0056] Furthermore, the calculation method of the end time t e of the vehicle phase in the current traffic flow direction is:

[0057] Take the time for the vehicles in the current flow direction to pass through the intersection as the green light duration G, and the green light duration G is:

[0058] G = t d + t l

[0059] Wherein, t d is the dissipation time of the vehicles in the current flow direction;

[0060] t l is the driving time of the last vehicle in the current flow direction within the intersection during the phase;

[0061] Calculate the end time t e of the vehicle phase in the current flow direction as:

[0062]

[0063] Furthermore, the specific process of Step Four is as follows:

[0064] Step Four One: Construct a phase information matrix Ψ according to the prediction result, and let the i-th row in the phase information matrix Ψ

[0065] Where: n i represents the number of vehicles in the i-th flow release echelon;

[0066] represents the moment when the first vehicle in the i-th flow direction arrives at the stop line;

[0067] represents the moment when the first vehicle in the i-th flow direction leaves the stop line,

[0068] g i,e represents the end time of the vehicle phase in the i-th flow direction;

[0069] G i represents the green light duration in the i-th flow direction;

[0070] Step Four Two: Conduct intensive processing on the phases of vehicles in each flow direction. The specific process is as follows:

[0071] Step Four Two One: Establish a merging judgment matrix C with dimensions I×I, where I represents the total number of vehicle flow directions at the current intersection, and the element C i,j in the i-th row and j-th column of the matrix C ∈ {0, 1}. When and only when the conditions (1) to (3) are met, the phase of the i-th flow direction and the phase of the j-th flow direction meet the merging condition, that is, C i,j = 1;

[0072] Condition (1): There is no spatial conflict between the vehicle trajectories in the $i$-th flow direction and the vehicle trajectories in the $j$-th flow direction;

[0073] Condition (2): $Q$ i $-Q$ j $ / \max(Q$ i , $Q$ j ) $\leq \delta$, where $\delta$ is a threshold, and $Q$ i represents the equivalent flow rate of the $i$-th flow direction phase, and $Q$ j represents the equivalent flow rate of the $j$-th flow direction phase;

[0074] Condition (3): $\Delta t$ i,j $= |T$ i $- T$ j $| \leq \varepsilon$, where $\varepsilon$ is a threshold, and $T$ i represents the phase demand duration of the $i$-th flow direction, and $T$ j represents the phase demand duration of the $j$-th flow direction;

[0075] Step 4.2.2: Perform matrix fusion:

[0076]

[0077] Among them, represents the Hadamard product;

[0078] $\Psi$ represents the transpose of $\Psi$ T ;

[0079] $\Psi'$ represents the merged phase set, that is, the intensive processing result;

[0080] Step 4.3: Use the state space $s = \{o, \Psi'\}$ as the input of the deep reinforcement learning model, where $I_0$ represents the total number of merged phases, represents a phase release order, and the adjusted phase release order and the start time of the green light for each phase are output through the deep reinforcement learning model.

[0081] Furthermore, in the merged phase set $\Psi'$, for the flow directions that meet the phase merging conditions, the start time of the green light $g$ b $' = \min\{g$ i,b $\}$, $i \in I'$, the end time of the green light $g$ e $' = \max\{g$ i,e $\}$, $i \in I'$, the duration of the green light of the merged phase $G' = g$ b $' - g$ e $'$;

[0082] $I'$ represents the set composed of the phases that are allowed to be merged with the $i$-th flow direction phase and the $i$-th flow direction phase;

[0083] n i ′ represents the total number of vehicles in the release echelon of the merged phase after phase merging within the set I′.

[0084] Furthermore, the reward function used during the training of the deep reinforcement learning model is:

[0085]

[0086] where, represents the moment when the leading vehicle of the i-th predicted flow direction arrives at the stop line;

[0087] g i ′ ,b represents the starting moment of the green light released for the i-th flow direction output by the deep reinforcement learning model;

[0088] α1 and α2 represent balance constants;

[0089] R represents the reward function.

[0090] Even further, the specific process of step five is as follows:

[0091] Define the phase adjustment error function E:

[0092] E = ψ p -ψ r

[0093] where, ψ p represents the vector composed of the moments when the first vehicle of each released phase arrives at the stop line predicted;

[0094] ψ r represents the vector composed of the actual green light opening moments within each released phase currently output by the deep reinforcement learning model;

[0095] When E = 0, maintain the phase release plan currently output by the deep reinforcement learning model, perform motion planning for the vehicles in each phase according to the actual green light opening moments within each released phase currently output by the deep reinforcement learning model, and obtain the vehicle motion planning results within each phase;

[0096] The intelligent signal control device sends the motion planning results to the on-vehicle unit through the roadside unit, and the vehicles in the first released phase are released according to the motion planning results; obtain the prediction results of the arrival moments of the vehicle platoons at the stop line within each phase based on the vehicle motion planning results within each phase, and use the obtained prediction results to return to and execute step three;

[0097] The specific process of the motion planning is as follows:

[0098]

[0099] where K is a constant;

[0100] represents the acceleration of the i'-th vehicle;

[0101] represents the time when the i'-th vehicle enters the control area;

[0102] represents the time when the i'-th vehicle leaves the control area;

[0103]

[0104] where f represents a second-order linear dynamics equation;

[0105] d i′ (t) represents the position of the i'-th vehicle at time t;

[0106] represents the speed of the i'-th vehicle at time t;

[0107] represents the control start time of the i'-th vehicle;

[0108] represents the distance from the control start point to the stop line;

[0109] When E≠0, according to the actual green light start times in each release phase output by the deep reinforcement learning model currently, motion planning is performed on the vehicles in each phase to obtain the vehicle motion planning results in each phase;

[0110] And according to the vehicle motion planning results in each phase, prediction results of the arrival times of vehicle platoons at the stop line in each phase are obtained, and the obtained prediction results are used to return to step three for execution.

[0111] The beneficial effects of the present invention are:

[0112] Aiming at the control problem of connected intersections, the present invention proposes a method for actively coordinating connected vehicles and intelligent signal machines, predicting the behaviors of upstream vehicles, and responding to traffic demands. The signal machine generates a signal control strategy based on the method of phase sequence and phase coupling, considering participating the signal states in future time periods in the current stage of control optimization to avoid short-sightedness, and aiming at the problems of complex solution and high dimension in traditional control problems, using DQN for optimization and solution; connected vehicles optimize control with the goal of minimizing energy consumption under constraints, and adopt decentralized control to improve control effects and prediction accuracy, and construct an energy consumption model to evaluate and optimize the vehicle trajectory control effects. The results show that, compared with traditional control methods, the method of the present invention improves the traffic efficiency of connected intersections, guarantees the right-of-way fairness of each flow direction, and reduces the travel energy consumption of vehicles at the same time.

[0113] By designing a DQN network model, accurately calibrating the state input and reward function, and dynamically optimizing the signal phase adjustment strategy, multi-objective collaborative optimization can be achieved while reducing the computational load. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] Figure 1 is a flowchart of a connected intersection ecological guidance collaborative control method based on a deep Q network according to the present invention;

[0115] Figure 2 is a schematic diagram of echelon division;

[0116] Figure 3 is a schematic diagram of intersection conflicts;

[0117] Figure 4 is a schematic diagram of a connected intersection;

[0118] Figure 5 is a schematic diagram of the arrival of vehicle flow in advance of the signal phase;

[0119] Figure 6 is a schematic diagram of the arrival of vehicle flow with a delayed signal phase Figure 1 ;

[0120] Figure 7 is a schematic diagram of the arrival of vehicle flow with a delayed signal phase Figure 2 ;

[0121] Figure 8 is a traffic simulation environment for urban intersections. DETAILED DESCRIPTION OF THE INVENTION

[0122] DETAILED DESCRIPTION OF THE INVENTION I: Combined with Figure 1 This embodiment is described. A connected intersection ecological guidance collaborative control method based on a deep Q network according to this embodiment specifically includes the following steps:

[0123] Step 1: Deploy intelligent signal control devices, roadside units (RSUs) and on-vehicle units (OBUs) with V2X communication capabilities. The roadside units send the collected information to the intelligent signal control devices;

[0124] And establish a communication link between the roadside unit and the on-vehicle unit;

[0125] Step 2: For each approach within the current intersection, the intelligent signal control device predicts the arrival time of the vehicle platoon at the stop line according to the received information;

[0126] Step 3: Dynamically divide the vehicles on each approach of the current intersection according to the prediction result of the arrival time of the vehicle platoon at the stop line, and divide the vehicles into two parts: a release echelon and a retention echelon;

[0127] Step 4: Intensively process the phases of the vehicles in each flow direction according to the number of vehicles in the release echelon of each flow direction, construct the state of the deep reinforcement learning model based on the intensive processing, and output the actual phase release sequence and the start time of the green light for each phase through the deep reinforcement learning model (it should be noted that for the first released phase in the actual phase release sequence, the actual start time of the green light for the first phase is equal to the time when the first vehicle in the first phase arrives at the stop line, and the actual start time of the green light for the second phase is the time when the last vehicle in the release gradient of the first phase exits the intersection, and so on, to obtain the actual start time of the green light for each phase);

[0128] Step 5: According to the prediction result in Step 2 and the output of the deep reinforcement learning model, determine whether to maintain the phase release plan currently output by the deep reinforcement learning model;

[0129] If maintaining the phase release plan currently output by the deep reinforcement learning model, perform vehicle movement planning for each phase according to the start time of the green light for each phase output by the deep reinforcement learning model, release the first phase in the phase release plan, then obtain the prediction result of the arrival time of the vehicle fleet at the stop line in each phase according to the vehicle movement planning result, and then use the prediction result to return to execute Step 3;

[0130] It should be noted that each time when the decision E = 0 is satisfied, only the vehicles in the first phase are released, and then the remaining vehicles in the first phase participate in the phase sorting again according to the predicted arrival time;

[0131] If not maintaining the phase release plan currently output by the deep reinforcement learning model, perform vehicle movement planning for each phase according to the start time of the green light for each phase output by the deep reinforcement learning model, then obtain the prediction result of the arrival time of the vehicle fleet at the stop line in each phase according to the vehicle movement planning result, and then use the prediction result to return to execute Step 3.

[0132] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that a multi-source sensing device is arranged in the roadside unit, and the multi-source sensing device is used to collect the instantaneous speed, the distance from the stop line of the approach lane, and the running time of each connected vehicle in real time; and send the information collected in real time and the real-time traffic flow data to the intelligent signal control device.

[0133] Other steps and parameters are the same as those in Specific Embodiment 1.

[0134] Specific Embodiment 3: The difference between this embodiment and Specific Embodiment 1 or 2 is that the specific process of Step 2 is as follows:

[0135] For any vehicle flow direction within the current intersection, taking the moment when the vehicle exits the merging area of the previous intersection as the starting point of the control area, allowing the vehicle to enter the control area and drive at the maximum acceleration until it reaches the maximum speed, the predicted value of the arrival time of the nth vehicle at the stop line is:

[0136]

[0137] Wherein, represents the predicted value of the arrival time of the nth vehicle at the stop line;

[0138] represents the start time of control of the nth vehicle, that is, the moment when the nth vehicle enters the control area;

[0139] represents the speed of the nth vehicle at moment;

[0140] represents the maximum allowable acceleration of the nth vehicle;

[0141] represents the predicted value of the arrival time of the (n - 1)th vehicle at the stop line;

[0142] represents the speed of the nth vehicle at moment;

[0143] represents the speed of the nth vehicle at moment;

[0144] τ n is the headway between the nth vehicle and the vehicle in front;

[0145] h n is the distance between the nth vehicle and the vehicle in front;

[0146] is the body length of the vehicle in front of the nth vehicle;

[0147] D represents the distance of the nth vehicle from the stop line at the current moment;

[0148] When n = 0, denote When the nth vehicle drives at the maximum acceleration and accelerates to v n,max at the stop line, the predicted value of the arrival time of the nth vehicle at the stop line is rewritten as:

[0149]

[0150] Wherein, represents the distance between the nth vehicle and the stop line at

[0151] represents the speed when the nth vehicle enters the control area.

[0152] Other steps and parameters are the same as those in the first or second specific implementation manner.

[0153] Specific implementation manner four: The difference between this implementation manner and one of the first to third specific implementation manners is that in step three, as Figure 2 shown, the specific process of dynamic echelon division of vehicles is as follows:

[0154] (1) The spatio-temporal trajectory of the released echelon vehicles satisfies:

[0155] h n ′ < d c and g min ≤ t duration ≤ g max

[0156] where h n ′ represents the headway between the nth vehicle and the (n - 1)th vehicle;

[0157] d c represents the critical headway, taking 15 m / s, used to divide the discontinuous traffic flow;

[0158] g min represents the minimum green time of the phase;

[0159] g max represents the maximum green time of the phase;

[0160] t duration represents the sum of Δt and the driving time of the nth vehicle in the intersection (generally taken as a fixed value), and Δt represents the difference between the time when the nth vehicle leaves the stop line and the time when the first vehicle leaves the stop line;

[0161] As Figure 3 shown is an intersection including 8 phases. For the vehicles in each phase, the judgment method of condition (1) is adopted. For any phase, all the vehicles satisfying condition (1) in this phase are divided into the released echelon;

[0162] (2) The set of vehicles N' in the detained echelon = M - N, where N represents the set of vehicles in the released echelon, and M represents the set of all vehicles in the current flow direction.

[0163] Other steps and parameters are the same as those in one of the first to third specific implementation manners.

[0164] Specific implementation manner five: The difference between this implementation manner and one of the first to fourth specific implementation manners is that the constraint satisfied when calculating the predicted value of the arrival time of the nth vehicle at the stop line is:

[0165] ① Speed constraint:

[0166]

[0167] Among them, represents the speed of the nth vehicle at time t;

[0168] ② Acceleration constraint:

[0169]

[0170] Among them, represents the maximum acceleration that the vehicle is allowed to adopt;

[0171] represents the maximum deceleration that the vehicle is allowed to adopt;

[0172] represents the acceleration of the nth vehicle at time t;

[0173] ③ Phase constraint: Since the start and end times of the optimized phase change within the control time domain, phase constraints are imposed on the vehicle platoons corresponding to each phase to ensure safe passage within the intersection merging area;

[0174]

[0175] Among them, g e = t e , t e represents the end time of the vehicle phase in the current flow direction;

[0176] Let represent the time when the nth vehicle leaves the stop line.

[0177] Other steps and parameters are the same as those in any one of the specific embodiments one to four.

[0178] Specific embodiment six: The difference between this embodiment and any one of the specific embodiments one to five is that the calculation method of the end time t e of the vehicle phase in the current flow direction is:

[0179] Take the time for the vehicle in the current flow direction to pass through the intersection as the green light duration G, and the green light duration G is:

[0180] G = t d + t l

[0181] Among them, t d is the dissipation time of the vehicle in the current flow direction;

[0182] t lis the driving time of the last vehicle in the current vehicle flow direction within the intersection during the phase;

[0183] Calculate the end time t of the vehicle phase in the current vehicle flow direction e as:

[0184]

[0185] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.

[0186] Specific Embodiment Seven: The difference between this embodiment and any one of the first to sixth specific embodiments is that the specific process of Step Four is as follows:

[0187] Step Four One. As Figure 4 shown is the arrival schematic diagram of all vehicles within the control area. According to the prediction results, construct the phase information matrix Ψ, and let the i-th row in the phase information matrix Ψ

[0188] where: n i represents the number of vehicles in the i-th vehicle release echelon in the flow direction;

[0189] represents the moment when the first vehicle in the i-th flow direction arrives at the stop line;

[0190] represents the moment when the first vehicle in the i-th flow direction leaves the stop line,

[0191] g i,e represents the end time of the vehicle phase in the i-th flow direction;

[0192] G i represents the green light duration in the i-th flow direction;

[0193] Since the basic phase queue order meets the traffic demands of each flow direction, it ensures that vehicles in each flow direction are in a protected phase to avoid conflicts. To reduce phase switching, phases are merged through intensive processing. During the phase merging process, the following need to be considered: 1. There should be a non-conflicting relationship between vehicles. As Figure 3 shown, that is, vehicle trajectories should not intersect at the same time; 2. The traffic volumes of the phases to be merged should not differ too much, so as to reduce the green light idling of the direction with a smaller traffic volume in the merged phase.

[0194] Step Four Two. Conduct intensive processing on the phases of vehicles in each flow direction. The specific process is as follows:

[0195] Step Four Two One. Establish a merging determination matrix C with dimensions I×I, where I represents the total number of vehicle flow directions at the current intersection, and the element C in the i-th row and j-th column of the matrix C i,j∈ {0, 1}, the phases of the i-th flow and the j-th flow satisfy the merging condition if and only if conditions (1) to (3) are met, that is, C i,j = 1;

[0196] Condition (1): There is no spatial conflict between the vehicle trajectories of the i-th flow and the j-th flow;

[0197] Condition (2): Q i -Q j | / max(Q i , Q j ) ≤ δ, where δ is taken as 0.25, and Q i represents the equivalent flow rate of the phase of the i-th flow, and Q j represents the equivalent flow rate of the phase of the j-th flow. The equivalent flow rate of each flow is calculated based on the number of vehicles in the release echelon of each flow;

[0198] Condition (3): Δt i,j = |T i -T j | ≤ ε (ε is taken as 3 s), and T i represents the phase demand duration of the i-th flow, and T j represents the phase demand duration of the j-th flow;

[0199] Step Four Two Two: Perform matrix fusion:

[0200]

[0201] where, represents the Hadamard product;

[0202] Ψ represents the transpose of Ψ T ;

[0203] Ψ' represents the merged phase set, that is, the intensive processing result;

[0204] Step Four Three: Use the state space s = {o, Ψ′} as the input of the deep reinforcement learning model, where, I0 represents the total number of merged phases, represents a phase release order (obtained randomly), and the adjusted phase release order and the start time of the green light for each phase are output through the deep reinforcement learning model.

[0205] Other steps and parameters are the same as those in any one of the specific embodiments one to six.

[0206] The present invention uses a trained DQN network model to solve problems. During the training process of the DQN network model, the Q-value network is used to calculate the Q-value selected by the current control strategy and update the Q-value iteratively, and the Target network is used to calculate the Q-value of the next state in the temporal difference algorithm. The update of network parameters comes from copying the parameters of the Q-value network.

[0207]

[0208] Update the network parameters through stochastic gradient descent, and update the weights θ, θ of the main network and the target network - 。

[0209]

[0210] θ - ←τθ+(1 - τ)θ -

[0211] In the formula, α is the learning rate, is the loss function, which is usually the mean square error (MSE) between the predicted Q-value and the target Q-value, τ is the soft update parameter.

[0212] Experience replay is used to store historical spatio-temporal states of intersections, signal decisions, benefit values, etc. Generate a control strategy through global information learning to adapt to traffic states.

[0213] T k =(s k ,a k ,R k ,s k+1 )

[0214] Specific Embodiment 8: Different from one of Embodiments 1 to 7, in the merged phase set Ψ', for the flow directions that meet the phase merging conditions, the start time g of the green light of the merged phase b ′=min{g i,b},i∈I′, the end time g of the green light of the merged phase e ′=max{g i,e},i∈I′, I′ represents the set composed of the phases that are allowed to merge with the phase of the i-th flow direction and the phase of the i-th flow direction, n i ′ represents the total number of vehicles in the release echelon of the merged phase after merging the phases in the set I′. The green light duration G′ of the merged phase is G′ = g b ′ - g e ′.

[0215] Other steps and parameters are the same as those in one of Embodiments 1 to 7.

[0216] Embodiment 9: The difference between this embodiment and any one of Embodiments 1 to 8 is that the reward function used during the training of the deep reinforcement learning model is as follows:

[0217]

[0218] where, represents the moment when the leading vehicle in the predicted i-th flow direction arrives at the stop line;

[0219] g i ′ ,b represents the start moment of the green light released for the i-th flow direction output by the deep reinforcement learning model;

[0220] α1 and α2 represent balance constants;

[0221] R represents the reward function.

[0222] Other steps and parameters are the same as those in any one of Embodiments 1 to 8.

[0223] In this embodiment, the design of the reward function takes into account the intersection traffic efficiency and the fairness of road rights. The goal is to maximize the number of vehicles passing through each phase, and further balance the fairness of road rights in each flow direction to minimize the delay of the leading vehicle in each direction.

[0224] Embodiment 10: The difference between this embodiment and any one of Embodiments 1 to 9 is that the specific process of Step 5 is as follows:

[0225] Define the phase adjustment error function E:

[0226] E = ψ p - ψ r

[0227] where, ψ p represents the vector composed of the moments when the first vehicle in each predicted released phase arrives at the stop line;

[0228] ψ r represents the vector composed of the actual green light start moments within each released phase currently output by the deep reinforcement learning model;

[0229] When E = 0, the end moment is equal to the predicted moment, the end state is equal to the predicted state, maintain the phase release plan currently output by the deep reinforcement learning model, and perform motion planning for the vehicles in each phase according to the actual green light start moments within each released phase currently output by the deep reinforcement learning model to obtain the vehicle motion planning results within each phase;

[0230] The intelligent signal control device sends the motion planning results to the on-vehicle unit through the roadside unit, and the vehicles within the first release phase are released according to the motion planning results; the predicted results of the arrival times of the vehicle platoons at the stop line within each phase are obtained based on the motion planning results of the vehicles within each phase, and the obtained predicted results are used to return to step three;

[0231] The specific process of the motion planning is as follows:

[0232] There is a monotonic relationship between the fuel consumption of each vehicle and its control input a i (t). Therefore, the problem of fuel consumption optimization can be transformed into optimizing the vehicle acceleration, and can be further expressed as minimizing the L2 norm of the acceleration.

[0233]

[0234] where K is a constant;

[0235] represents the acceleration of the i'-th vehicle;

[0236] represents the time when the i'-th vehicle enters the control area;

[0237] represents the time when the i'-th vehicle leaves the control area;

[0238]

[0239] Defining the solution of the vehicle speed trajectory within the control area as an optimal control problem, establishing the state equation of the controlled system, and assuming that f satisfies the existence of a locally convergent solution within the constraint conditions of the equation:

[0240]

[0241] where f represents the second-order linear dynamics equation;

[0242] d i′ (t) represents the position of the i'-th vehicle at time t;

[0243] represents the speed of the i'-th vehicle at time t;

[0244] represents the start time of control of the i'-th vehicle;

[0245] represents the distance from the start of control to the stop line;

[0246] When E < 0, since the green light duration is the shortest time to ensure that the divided vehicle platoons completely pass through the intersection and to avoid a green light with no vehicles passing, at this time Due to the limitations of vehicle and road safety, the guiding strategy may have non-optimal solutions, that is it is necessary to dynamically adjust the phase parameters based on the real-time platoon state, such as Figure 5 shown;

[0247] When E > 0, at this time due to the phase delay to turn on the light, the platoon that arrives in advance may queue up at the stop line, that is at the same time, when subsequent vehicles arrive one after another and cause the release echelon to expand, the mismatch between the position of the last vehicle and the passing time window will trigger an increase in the demand for green light duration. It is necessary to dynamically adjust the phase parameters based on the real-time platoon state, such as Figure 6 and Figure 7 shown;

[0248] When E ≠ 0, according to the actual green light opening time in each release phase output by the deep reinforcement learning model currently, perform motion planning on the vehicles in each phase to obtain the vehicle motion planning results in each phase;

[0249] And according to the vehicle motion planning results in each phase, obtain the prediction results of the platoon arrival time at the stop line in each phase, and use the obtained prediction results to return to step three for execution.

[0250] Other steps and parameters are the same as one of the specific embodiments one to nine.

[0251] When the vehicle is not controlled, due to the loss of the right of way, the vehicle will stop and wait at the stop line. In the platoon control system, as Figure 4 shown, the vehicle can communicate with the surrounding environment (including signal machines and other CAVs) through vehicle-to-vehicle (V2V) and internet-to-vehicle (I2V) technologies. All CAVs are equipped with on-vehicle controllers, which can control the CAVs at each time and each location according to the signal strategy and vehicle information. In order to smooth the vehicle trajectory to save energy consumption and provide driving safety, the CAV will be controlled to have a smooth trajectory before approaching the signal intersection, without sudden acceleration and deceleration or excessive start and stop in the section.

[0252] Experimental part

[0253] In view of the complex and diverse current urban road operation scenarios, the road operation conditions and traffic operation states are dynamically changing, and the fixed-time signal system is difficult to adapt to the characteristics of real-time changing traffic flow and frequent start and stop during vehicle driving. The present invention proposes a connected intersection ecological guiding collaborative control method to achieve adaptive control of urban intersections and ecological guidance of vehicles. Use VISSIM simulation software to build a simulated operation scenario of urban intersections, and apply the method of the present invention to traffic flow states with different flow rates and different arrival rates. Now, the rationality of the method of the present invention will be further elaborated through the process of establishing the simulation scenario, and its specific implementation process is as follows:

[0254] Step 1: First, according to the analysis of urban intersections, design a control experiment of the intersection applying the method of the present invention and the conventional intersection under the traditional fixed timing.

[0255] Step 2: Build a traffic environment of the urban intersection on the VISSIM platform, and perform relevant traffic flow parameter settings, signal group arrangements, and detector settings. While making the simulation environment as close as possible to the real environment, obtain risk perception data, such as Figure 8 shown.

[0256] Step 3: Take the traffic flow parameters of driving on urban roads as the input. Design experiments with the one-way traffic flow of the roads in the simulation environment being 800 pcu / h and 1200 pcu / h respectively. Through Python connecting to the com port, the traffic flow data during the operation of the simulation environment can be detected and processed to carry out subsequent model building.

[0257] Step 4: Use the processed data to predict the vehicle arrival behavior and preliminarily arrange the required passing phases for it. Optimize the optimal release order of the required passing phases of the vehicle through the DQN network, and use the on-vehicle controller to solve the optimal control problem online to achieve control, and output the signal control scheme and the vehicle guidance scheme.

[0258] Step 5: Evaluate the operation state of the urban intersection under the method of the present invention. Select the average vehicle delay, travel time, and vehicle output power as the evaluation indicators for the operation effect of the method of the present invention. The traffic evaluation indicators are shown in Table 1. It can be seen from Table 1 that the method of the present invention improves the passing efficiency and average vehicle delay of the intersection, and realizes the reduction of vehicle energy consumption (the energy consumption calculation formula is:

[0259] Table 1

[0260]

[0261]

[0262] Finally, through simulation, it can be basically determined that the simulation results are consistent with the actual situation and have certain reference value and practical significance, verifying the effectiveness of the connected intersection ecological guidance collaborative control method proposed by the present invention.

[0263] The method of the present invention extracts the road information of connected vehicles in real time. To respond to their traffic demands, it first makes a preliminary arrangement for the traffic phases they need, and then when performing real-time control on different urban intersections, it can adjust the right-of-way allocation of each flow direction by different divisions of the released vehicle platoons. Secondly, in the method of the present invention, the DQN network model is first used to optimize the order of the preliminarily arranged phases, and then the on-vehicle controller is used to guide the vehicles in combination with signal strategies and vehicle body conditions. Finally, the non-optimal control behaviors are iteratively optimized to determine the cooperative control method. Since there are various neural network models, different network models can be selected for replacement during actual operation, and the phase sequence optimization task can also be completed. The on-vehicle controller can also select different optimal control objectives to meet different operation requirements.

[0264] The above-mentioned calculation examples of the present invention are only for illustrating in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made on the basis of the above description. It is impossible to enumerate all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. A networked intersection ecological guidance collaborative control method based on deep Q network, characterized in that: The method specifically comprises the following steps: Step 1: deploy intelligent signal control equipment, roadside units and vehicle-mounted units, and the roadside units send the collected information to the intelligent signal control equipment; And establish a communication link between the roadside unit and the vehicle-mounted unit; Step 2: For each entrance lane in the current intersection, the intelligent signal control device predicts the time when the convoy arrives at the stop line based on the received information; Step 3: According to the predicted result of the time when the convoy arrives at the stop line, the vehicles at each entrance of the current intersection are dynamically divided into two echelons: a release echelon and a retention echelon; Step 4: Perform intensive processing on the phases of vehicles in each direction according to the number of vehicles in the release echelon of each direction, build the state of the deep reinforcement learning model based on the intensive processing, and output the actual phase release sequence and the green light start time of each phase through the deep reinforcement learning model; Step 5: According to the prediction result in step 2 and the output of the deep reinforcement learning model, determine whether to maintain the phase release scheme currently output by the deep reinforcement learning model; If the phase release plan currently output by the deep reinforcement learning model is maintained, the vehicle motion planning of each phase is performed according to the green light start time of each phase output by the deep reinforcement learning model, and the first phase in the phase release plan is released. Then, the predicted result of the time when the convoy arrives at the stop line in each phase is obtained based on the vehicle motion planning result, and then the predicted result is used to return to execute step three; If the phase release plan currently output by the deep reinforcement learning model is not maintained, the vehicle motion planning for each phase is performed according to the green light start time of each phase output by the deep reinforcement learning model, and then the arrival time of the fleet at the stop line in each phase is predicted based on the vehicle motion planning result, and then the prediction result is used to return to execute step three.

2. According to claim 1, a networked intersection ecological guidance collaborative control method based on deep Q network is characterized in that: The roadside unit is provided with a multi-source sensing device, which is used to collect the instantaneous speed, distance from the entrance lane stop line and running time of each connected vehicle in real time; and send the real-time collected information and real-time traffic flow data to the intelligent signal control device.

3. According to claim 2, a networked intersection ecological guidance collaborative control method based on deep Q network is characterized in that: The specific process of step 2 is as follows: For any flow of vehicles in the current intersection, the vehicle exiting the merge area of ​​the previous intersection is taken as the starting point of the control area, and the vehicle is allowed to enter the control area and drive at the maximum acceleration to the maximum speed. The predicted value of the time when the nth vehicle reaches the stop line is: in, It represents the predicted value of the time when the nth vehicle reaches the stop line; Indicates the start time of control of the nth vehicle, that is, the moment when the nth vehicle enters the control area; Indicates that the nth car is The speed of the moment; represents the maximum allowed acceleration of the nth vehicle; It represents the predicted value of the time when the n-1th vehicle reaches the stop line; Indicates that the nth car is The speed of the moment; Indicates that the nth car is The speed of the moment; τ n is the headway between the nth vehicle and the preceding vehicle; h n is the inter-vehicle distance between the nth vehicle and the preceding vehicle; is the body length of the preceding vehicle of the nth vehicle; D represents the distance of the nth vehicle relative to the stop line at the current moment; When n = 0, When the nth vehicle is traveling at maximum acceleration and accelerates to v before the stop line n,max When , the predicted value of the time when the nth vehicle reaches the stop line is rewritten as: in, express The distance between the nth vehicle and the stop line at the moment; Indicates the speed of the nth vehicle when it enters the control area.

4. According to claim 3, a networked intersection ecological guidance collaborative control method based on deep Q network is characterized in that: In step 3, the specific process of dividing vehicles into dynamic echelons is as follows: (1) The space-time trajectory of the released echelon vehicles satisfies: h n '<d c And g min ≤t duration ≤g max Among them, h n ′ represents the headway between the nth vehicle and the n-1th vehicle; d c represents the critical spacing; g min Indicates the shortest green light time of the phase; g max Indicates the maximum green time of the phase; t duration represents the sum of Δt and the travel time of the nth vehicle in the intersection, and Δt represents the difference between the time when the nth vehicle leaves the stop line and the time when the first vehicle leaves the stop line; (2) The set of vehicles in the retention echelon is N′=MN, where N represents the set of vehicles in the release echelon and M represents the set of all vehicles in the current flow direction.

5. The method for ecological guidance and collaborative control of networked intersections based on a deep Q network according to claim 4 is characterized in that: The constraints satisfied when calculating the predicted value of the time when the nth vehicle reaches the stop line are: ①Speed ​​constraint: in, represents the speed of the nth vehicle at time t; ② Acceleration constraint: in, Indicates the maximum acceleration the vehicle is allowed to adopt; Indicates the maximum deceleration the vehicle is allowed to adopt; represents the acceleration of the nth vehicle at time t; ③Phase constraint: Among them, g e =t e , t e Indicates the end time of the vehicle phase of the current flow direction; make Indicates the time when the nth vehicle leaves the stop line.

6. The method for ecological guidance and collaborative control of networked intersections based on a deep Q network according to claim 5 is characterized in that: The end time t of the vehicle phase of the current flow direction e The calculation method is: The time it takes for vehicles in the current flow direction to pass through the intersection is taken as the green light duration G, which is: G=t d +t l Among them, t d is the dissipation time of vehicles in the current flow direction; t l is the travel time of the tail vehicle in the current phase of the onward vehicle in the intersection; Calculate the end time t of the vehicle phase in the current flow direction e for:

7. The method for ecological guidance and collaborative control of networked intersections based on a deep Q network according to claim 6 is characterized in that: The specific process of step 4 is as follows: Step 41: Construct a phase information matrix Ψ based on the prediction results, and let the i-th row in the phase information matrix Ψ Where: n i represents the number of vehicles in the i-th flow release echelon; It indicates the time when the first vehicle in the i-th flow direction reaches the stop line; represents the time when the first vehicle in the i-th flow direction leaves the stop line, g i,e represents the end time of the vehicle phase of the i-th flow direction; G i represents the green light duration of the i-th flow direction; Step 42: Intensive processing is performed on the phases of each flow direction vehicle. The specific process is as follows: Step 421: Establish a merge decision matrix C of dimension I×I, where I represents the total number of vehicle flows at the current intersection, and the element C in the i-th row and j-th column of the matrix C i,j ∈{0,1}, the phase of the i-th flow direction and the phase of the j-th flow direction meet the merging condition if and only if conditions (1) to (3) are met, that is, C i,j =1; Condition (1): There is no spatial conflict between the vehicle trajectory of the i-th flow direction and the vehicle trajectory of the j-th flow direction; Condition (2), Q i -Q j | / max(Q i ,Q j )≤δ, δ is the threshold, Q i represents the phase equivalent flow rate of the i-th flow direction, Q j represents the phase equivalent flow in the jth flow direction; Condition (3), Δt i,j =|T i -T j |≤ε, ε is the threshold, T i represents the phase demand duration of the i-th flow direction, T j represents the phase demand duration of the jth flow direction; Step 422: Execute matrix fusion: in, represents the Hadamard product; Ψ stands for Ψ T The transpose of Ψ' represents the combined phase set, i.e., the result of intensive processing; Step 43: Use the state space s = {o, Ψ′} as the input of the deep reinforcement learning model, where I0 represents the total number of phases after merging. Represents a phase release sequence, and outputs the adjusted phase release sequence and the green light start time of each phase through the deep reinforcement learning model.

8. The method for ecological guidance and collaborative control of networked intersections based on a deep Q network according to claim 7 is characterized in that: In the merged phase set Ψ', for the flow direction that meets the phase merging condition, the green light start time g of the merged phase b ′=min{g i,b }, i∈I′, the green light end time g of the merged phase e ′=max{g i,e },i∈I′, The green light duration of the merged phase G′=g b ′-g e ′; I′ represents the set consisting of the phases allowed to merge with the i-th flow direction phase and the i-th flow direction phase; n i ′ represents the total number of vehicles in the released echelon of the merged phase after the phases in the set I′ are merged.

9. The method for ecological guidance and collaborative control of networked intersections based on a deep Q network according to claim 8 is characterized in that: The reward function used in the training of the deep reinforcement learning model is: in, It indicates the predicted time when the first vehicle of the ith flow direction arrives at the stop line; g i ' ,b Indicates the green light start time for the i-th flow release output by the deep reinforcement learning model; α1 and α2 represent equilibrium constants; R represents the reward function.

10. The method for ecological guidance and coordinated control of networked intersections based on a deep Q network according to claim 9, characterized in that: The specific process of step five is as follows: Define the phase adjustment error function E: E=ψ p -ψ r Among them, ψ p A vector representing the predicted time when the first vehicle in each release phase reaches the stop line; ψ r Represents the vector composed of the actual green light on time in each release phase currently output by the deep reinforcement learning model; When E=0, the phase release scheme currently output by the deep reinforcement learning model is maintained, and the motion planning of the vehicles in each phase is performed according to the actual green light on time in each release phase currently output by the deep reinforcement learning model, and the vehicle motion planning results in each phase are obtained; The intelligent signal control device sends the motion planning results to the vehicle-mounted unit through the roadside unit, and the vehicles in the first release phase are released according to the motion planning results; the predicted results of the time when the convoy in each phase arrives at the stop line are obtained according to the vehicle motion planning results in each phase, and the obtained prediction results are used to return to execute step three; The specific process of motion planning is as follows: Where K is a constant; represents the acceleration of the i′th vehicle; represents the time when the i′th vehicle enters the control area; represents the time when the i′th vehicle leaves the control area; Where, f represents the second-order linear dynamic equation; d i′ (t) represents the position of the i′th vehicle at time t; represents the speed of the i′th vehicle at time t; represents the control start time of the i′th vehicle; Indicates the distance from the start of control to the stop line; When E≠0, according to the actual green light on time in each release phase currently output by the deep reinforcement learning model, the vehicle motion planning is performed for each phase to obtain the vehicle motion planning results in each phase; And according to the vehicle motion planning results in each phase, the predicted results of the time when the convoy in each phase arrives at the stop line are obtained, and the obtained predicted results are used to return to execute step three.

Citation Information

Patent Citations

  • Traffic signal lamp control system and method based on reinforcement learning and dynamic timing

    CN113763723A

  • Cooperative guidance control method for mixed vehicle fleet at entrance lane of intersection

    CN119169818A

  • Vehicle-road cooperative control method and system for double-layer ecological city in mixed traffic flow environment

    CN119380561A

  • Hierarchical Optimization-based Coordinated Control of Traffic Rules and Mixed Traffic in Multi-Intersection Environments

    US20240331535A1

Cited By

  • Road train dynamic road right distribution method based on reinforcement learning and rule constraint

    CN121838503A