Hybrid Traffic Flow Cooperative Control Method Based on Double-Layer Parameterized Deep Reinforcement Learning
Through the double-layer parameterized deep reinforcement learning model, combined with convolutional neural network and deep Q network, the lane traffic priority and order are optimized, and the collaborative control problem between CAV and HV in hybrid flow intersections is solved, and the intelligence level and traffic efficiency of the traffic control system are improved.
Patent Information
- Application Number
- CN202411633391.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-11-15
AI Technical Summary
When facing hybrid flow intersections, the existing traffic control methods lack flexibility and intelligence, and cannot effectively coordinate networked autonomous driving vehicles (CAVs) with human-driving vehicles (HVs), resulting in unbalanced traffic flow and low traffic efficiency. The existing deep reinforcement learning algorithms are difficult to comprehensively consider multi-level decision-making and high-dimensional traffic characteristics.
Using a method based on double-layer parameterized deep reinforcement learning, a high-dimensional traffic space feature is extracted through a convolutional neural network, and a two-layer deep reinforcement learning model is constructed. The first layer selects the lane with the highest pass priority and determines the pass duration. The second layer optimizes and coordinates the pass sequence of vehicles on the lane, combines the event-triggered redecision mechanism to optimize model parameters to improve pass efficiency and safety.
Effectively coordinate the passage of CAV and HV, improve the traffic efficiency and safety of hybrid flow intersections, and is especially suitable for hybrid flow intersections without signal control.
Smart Images

Figure CN119479295B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of collaborative control of autonomous driving and human driving, and particularly to a collaborative control method for mixed traffic flow based on double-layer parameterized deep reinforcement learning. Background Art
[0002] With the acceleration of urbanization, traffic congestion has become an important problem faced by many cities. In particular, the traffic efficiency at key traffic nodes such as intersections directly affects the smoothness and safety of the entire traffic network. In recent years, autonomous driving technology has developed rapidly, and connected autonomous vehicles (CAVs) have gradually been applied to the traffic system, bringing new challenges to traditional traffic control methods. Under this background, how to coordinate the passage of CAVs and human-driven vehicles (HVs) at mixed-flow intersections, especially at unsignalized intersections, has become a key issue in traffic system optimization.
[0003] Traditional traffic signal control methods usually adjust the passage priority at intersections based on fixed-cycle signal switching or preset rules based on traffic flow. These methods are suitable for fixed traffic patterns, but in the face of mixed traffic flows (CAVs and HVs existing simultaneously), traditional methods lack flexibility and intelligence and cannot make optimal decisions based on real-time traffic conditions, resulting in unbalanced traffic flows and low traffic efficiency.
[0004] In recent years, deep reinforcement learning (DRL), as an emerging artificial intelligence technology, has achieved remarkable success in many fields, especially in traffic control, autonomous driving and other fields. DRL continuously obtains feedback from the environment and learns how to maximize a certain goal (such as traffic flow or safety) through intelligent decision-making. However, although DRL has made some progress in traffic control, existing research still has limitations: (1) For mixed-flow intersections, existing DRL methods often do not fully consider the synergistic effect between the two, resulting in mutual interference and disharmony between CAVs and HVs, reducing the traffic efficiency at intersections; (2) Existing DRL methods are usually optimized for a single task, such as the order of vehicle passage, and it is difficult to comprehensively consider multiple levels of decision-making. In actual traffic scenarios, not only the lane priority needs to be determined, but also the dynamic adjustment of coordinated and conflicting lanes needs to be considered, and at the same time, the optimization of passage duration is required. These factors require a multi-level reinforcement learning mechanism to coordinate. (3) Traffic intersections usually have highly complex environmental characteristics (such as traffic flow density, traffic state, vehicle speed, etc.), which are key factors affecting traffic efficiency and safety. However, existing reinforcement learning algorithms usually lack effective feature extraction capabilities when facing these high-dimensional and complex traffic characteristics, resulting in unsatisfactory decision-making effects. Summary of the Invention
[0005] To address the deficiencies of the prior art, the present invention proposes a collaborative control method for mixed traffic flow based on double-layer parameterized deep reinforcement learning. By combining a convolutional neural network (CNN), high-dimensional spatial features are extracted, and dynamic adjustment of passing priorities, passing durations, and passing sequences is achieved through two layers of deep reinforcement learning models, thereby enhancing the intelligence level and passing efficiency of the traffic control system.
[0006] The technical solution adopted by the present invention to solve the above tasks is as follows:
[0007] A collaborative control method for mixed traffic flow based on double-layer parameterized deep reinforcement learning, comprising the following steps:
[0008] Construct a road model for a signal-free artificial-autonomous driving mixed flow intersection and divide the intersection area into functional zones;
[0009] Based on the original multi-channel parameterized deep Q-network (MPDQN) algorithm, combine CNN to construct the first-layer deep reinforcement learning algorithm model (CNN-MPDQN), and effectively extract the high-dimensional spatial features of the traffic at the intersection through CNN; the core task of the CNN-MPDQN algorithm is to select a lane with the highest passing priority from multiple lanes and determine the passing duration of that lane;
[0010] According to the lane with the highest passing priority selected by the first-layer reinforcement learning, determine the coordinated lanes and conflicting lanes; within the passing duration, vehicles on the coordinated lanes can pass through the intersection in coordination, while vehicles on the conflicting lanes need to stop and wait for the next round of optimization;
[0011] According to the traffic flow of each coordinated lane, respectively allocate the proportion of the passing duration to maximize the utilization of the passing duration and optimize the traffic flow efficiency;
[0012] Construct the second-layer deep reinforcement learning algorithm model, adopt the deep Q-network (DQN) algorithm, and within the first-layer passing duration, reinforcement learning determines the passing sequence of vehicles on the coordinated lanes, aiming to solve the trajectory conflict problem of vehicles on the coordinated lanes;
[0013] Under the high-level decision of the passing sequence of the two-layer reinforcement learning, the priority passing vehicles pass through the intersection at the maximum safe speed, while the conflicting vehicles need to decelerate and stop before the conflict position;
[0014] Through continuous cyclic training and iteration of the two-layer deep reinforcement learning algorithm model, gradually optimize the model parameters, and finally obtain the optimal convergence model; this model can effectively handle various traffic conditions and vehicle behaviors, and improve the passing efficiency and safety of the intersection.
[0015] Furthermore, the constructed road model is a four-way three-lane intersection, with left-turn, straight, and right-turn lanes in each direction. The right-turn lane is uncontrolled. The intersection area is divided into a detection area and a conflict area. The section between the maximum communication range and the stop line is the detection area, and the conflict area is the area where a vehicle may conflict with vehicles from other directions after crossing the stop line. CAVs and HVs coexist on the road. Vehicles that enter the detection area and communicate with the intersection central processor are identified as CAVs, while vehicles that can be detected by the central processor but do not communicate are identified as HVs. All vehicles travel along pre-planned trajectories.
[0016] Furthermore, the driving behaviors of the HVs are mainly divided into two types: following the vehicle in front and driving freely without a vehicle in front. When two free HVs from conflicting directions arrive at the intersection simultaneously, they only need to follow the right-hand rule to pass through the intersection. If there is a vehicle in front of the HV, to avoid collision, the HV needs to maintain a safe distance. Therefore, the target speed of the HV at time should satisfy the following conditions:
[0017]
[0018] where is the maximum speed limit of the road, is the maximum acceleration, is the simulation update step size, is the safe following speed, The formula is:
[0019]
[0020] where is the speed of the vehicle in front at time t, is the distance between vehicles at time t, is the reaction time of the human driver;
[0021] In addition, a random factor is introduced to represent the uncertainty of human driving behavior. The final following speed is:
[0022] .
[0023] Furthermore, for the first-layer deep reinforcement learning algorithm model, the CNN-MPDQN algorithm is adopted. The specific state space, action space, reward function, network architecture, and loss function are designed as follows:
[0024] 4.1 State Space Design
[0025] To visually present the current traffic conditions at the intersection, a grid method is used to discretize the intersection, where the side length of the grid is set to the lane width; the state space is represented by a matrix of the position, speed, and waiting time of vehicles on the lane; specifically, the position matrix is used to indicate whether each grid is occupied by a vehicle, using 1 or 2 to represent the presence of a CAV or HV in the position matrix respectively. When the vehicle occupies more than half of the grid area, the corresponding grid is marked as 1 or 2, otherwise it is marked as 0; the speed matrix records the speed value of the vehicle when there is a vehicle in the grid; the waiting time matrix records the waiting time value of the vehicle in the grid; this method can effectively reflect the traffic flow pattern and vehicle state at the intersection;
[0026] 4.2 Action Space Design
[0027] The parameterized action space consists of a set of discrete actions, , where each discrete action corresponds to a continuous parameter, ; specifically, the discrete action is represented as all candidate lanes. After selecting one lane as the highest-priority lane, the duration of passage of this lane is defined as its corresponding continuous parameter; this action space can be represented as:
[0028]
[0029] where, is the set of continuous parameters for all ;
[0030] 4.3 Reward Function Design
[0031] Design a multi-objective reward function that comprehensively considers efficiency, fairness, and HV , which is defined as:
[0032]
[0033] where, represents the waiting time of vehicles that are penalized for conflicting with the vehicles in the highest-priority lane and causing them to stop and wait;
[0034] is intended to increase the chance for vehicles to become coordinated vehicles so that they can pass through the intersection in coordination, rather than becoming conflicting vehicles and being forced to wait;
[0035] is intended to increase the chance for more HVs to become priority vehicles, rather than becoming conflicting vehicles or coordinated vehicles, which interferes with the safe passage of HVs and the coordination of CAV vehicles in the second layer;
[0036] where, represents the waiting time of the th vehicle, denotes the waiting time threshold, denotes the number of conflicting vehicles, denotes the number of coordinated vehicles, denotes the CAV ratio on the highest-priority lane, , , denotes the weight coefficient;
[0037] 4.4 Network Architecture Design
[0038] The CNN-MPDQN algorithm model structure includes a continuous parameter policy network, a discrete action Q network, and a target network; the traffic high-dimensional space features of the intersection are effectively extracted through CNN; specifically, the continuous parameter policy network takes the current state as input, and after being processed by a convolutional layer, a fully connected layer, and an activation function, outputs the optimal vector of all continuous parameters in the action space; subsequently, this vector is separated so that each discrete action only inputs the parameters related to it, thereby reducing the influence of other continuous parameters on the Q-value estimation; in order to keep the length of the continuous action parameter vector consistent, the irrelevant continuous parameters are set to 0, and then the vectors formed by concatenating the current state vector processed by the convolutional layer and the separated continuous actions are input into the discrete action Q network, and after passing through the fully connected layer and the activation function, output matrix, and the Q-values on the diagonal are extracted from it to form a vector, and finally the optimal Q-value of the discrete action Q network is obtained;
[0039] 4.5 Loss Function Design
[0040] The loss function includes two parts: the discrete action Q network and the continuous parameter policy network; specifically, for the loss function of the discrete action Q network, it is constructed in the form of mean square error, aiming to minimize the target loss:
[0041]
[0042] where, is the n-step target function, expressed as:
[0043]
[0044] where, is the discount factor, and are the discrete action target Q network parameters and the continuous parameter policy target network parameters respectively;
[0045] For the loss function of the continuous parameter policy network, the goal is to maximize the Q-value estimation, defined as:
[0046] .
[0047] Further, the coordinated lanes are a combination of lanes that do not have a trajectory conflict with the highest-priority lane, and the conflicting lanes are a combination of lanes that have a trajectory conflict with the highest-priority lane.
[0048] Further, according to the traffic flow of each coordinated lane, the proportion of the passing duration is allocated respectively. The specific calculation formula for the proportion weight of the th coordinated lane is as follows:
[0049]
[0050] Among them, is the number of coordinated lanes, and are respectively the number of waiting vehicles in the detection area on the th and the th lanes. and are respectively the future predicted traffic flows on the th and the th lanes, which are calculated based on the arrival time of the vehicles. The arrival time follows a Poisson distribution.
[0051] Further, for the second-layer deep reinforcement learning algorithm model, the DQN algorithm is adopted. During the passing duration of the first layer, the reinforcement learning algorithm determines the passing order of the vehicles on the coordinated lanes. The specific state space, action space, and reward function of this model are designed as follows:
[0052] 7.1 State space design
[0053] The grid method is used to discretize the intersection. In the second-layer deep reinforcement learning algorithm model, the passing order of the vehicles on the coordinated lanes is mainly concerned. Therefore, the state space only represents the state information of the vehicles on the coordinated lanes and is described by the matrices of the positions, expected positions, speeds, and waiting times of the vehicles on the coordinated lanes. Specifically, the position matrix is used to represent whether each grid is occupied by a vehicle, the expected position matrix represents the grid that the vehicle will occupy in the future, and is used to reflect the conflict relationship between vehicles. The speed matrix records the speed value of the vehicle when there is a vehicle in the grid. The waiting time matrix records the waiting time value of the vehicle in the grid. This design can accurately represent the vehicle state of the coordinated lanes and provide a basis for optimizing the passing order.
[0054] 7.2 Action space design
[0055] The response action is designed as a combination of the passing orders of the vehicles on the coordinated lanes, which is expressed as:
[0056]
[0057] There are A set of candidate action combinations; based on this, the reinforcement learning algorithm identifies the optimal passing sequence combination;
[0058] 7.3 Reward Function Design
[0059] The reward function of the second layer is defined to minimize the waiting time and passing time of the coordinated vehicles, expressed as:
[0060]
[0061] where, represents the penalty for the waiting time of the coordinated vehicle, represents the penalty for the passing time of the coordinated vehicle, aiming to enable the coordinated vehicle to pass through the intersection as soon as possible; where, represents the th passing time of the coordinated vehicle, represents the passing time threshold, , represents the weight coefficient;
[0062] 7.4 Network Architecture Design
[0063] The DQN algorithm model structure includes a Q-value network and a target network; consistent with the first-layer algorithm model, a CNN is used as the feature extractor; specifically, the Q-value network takes the current state as input, and after being processed by a convolutional layer, a fully connected layer, and an activation function, outputs the optimal Q-value of the action; the target network is used to prevent the Q-value from being updated too quickly, thereby avoiding unstable training. Specifically, in the form of soft update, the parameters of the target network are synchronized with the Q-value network only once every N steps, that is:
[0064]
[0065] where, is the parameter of the target network, is the parameter of the Q-value network;
[0066] 7.5 Loss Function Design
[0067] The loss function is constructed in the form of mean square error, aiming to reduce the loss from the n-step target function:
[0068] .
[0069] Furthermore, under the advanced decision-making of the passing sequence in two-layer reinforcement learning, the priority passing vehicles pass through the intersection at the maximum safe speed, while the conflicting vehicles need to decelerate and stop waiting before the conflict position; specifically, during the driving process of the priority passing CAV vehicles in the coordinated lane, they may encounter free-running HVs in other coordinated lanes. To address this situation, the method adopts an event-triggered re-decision mechanism.
[0070] Furthermore, the method adopts an event-triggered re-decision mechanism. Specifically, when it is detected that a priority CAV vehicle and an HV vehicle may simultaneously occupy the conflict point of the intersection, the system will re-adjust the maximum safe speed of the vehicles passing through the coordinated lane according to the estimated arrival times of the priority CAV and the HV at the conflict point of the intersection. This maximum safe speed can ensure that the priority CAV vehicle and the HV vehicle can pass through the conflict area of the intersection safely without collision. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 It is a schematic flow chart of the collaborative control method of the present invention;
[0072] Figure 2 It is a schematic diagram of the intersection scenario;
[0073] Figure 3 It is a schematic diagram for constructing the state matrix;
[0074] Figure 4 It is a schematic diagram of the CNN-MPDQN algorithm model structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0075] The present invention will be further described in detail below with reference to the accompanying drawings.
[0076] As Figure 1 shown, the collaborative control method based on double-layer parameterized deep reinforcement learning of the present invention is applied to a mixed-flow intersection, and the specific steps are as follows:
[0077] Step 1: Construct a road model for a signal-free manual-autonomous vehicle mixed-flow intersection.
[0078] As Figure 2 shown, the constructed road model is a four-way three-lane intersection, with left-turn, straight-ahead, and right-turn lanes provided in each direction, and the right-turn lane is not controlled; the intersection area is divided into a detection area and a conflict area; the section between the maximum communication range and the stop line is the detection area, and the conflict area is the area where vehicles may conflict with vehicles in other directions after crossing the stop line; CAVs and HVs coexist on the road, and vehicles entering the detection area and communicating with the central processor of the intersection are identified as CAVs, while vehicles that can be detected by the central processor but do not communicate are identified as HVs; all vehicles travel along the pre-planned trajectories. Inside the intersection, the point where two driving trajectories intersect is defined as the conflict point location.
[0079] The collaborative method described in the present invention fully considers the different driving behaviors of HVs and realizes the effective coordination of CAVs on the premise of minimizing interference with the normal driving of HVs. The driving behaviors of HVs considered are mainly divided into two types: following the vehicle in front and free driving without a vehicle in front; when two free HVs from conflicting directions arrive at the intersection simultaneously, they only need to follow the right-hand rule to pass through the intersection; if there is a vehicle in front of the HV, to avoid collision, the HV needs to maintain a safe distance. Therefore, the target speed of the HV at the moment should satisfy the following conditions:
[0080]
[0081] where, is the maximum speed limit of the road, is the maximum acceleration, is the simulation update step size, is the safe following speed, The formula is:
[0082]
[0083] where, is the speed of the vehicle in front at time t, is the distance between vehicles at time t, is the reaction time of the human driver;
[0084] In addition, a random factor is introduced to characterize the uncertainty of human driving behavior. The final following speed is:
[0085] .
[0086] Step 2: Use the CNN-MPDQN algorithm to construct the first-layer deep reinforcement learning algorithm model. Its core task is to select a lane with the highest passing priority from multiple lanes and determine the passing duration of that lane; the specific state space, action space, reward function, network architecture, and loss function are designed as follows:
[0087] Step 2.1: State space design
[0088] To visually present the current traffic conditions at the intersection, the intersection is discretized using a grid method, where the side length of the grid is set to the lane width; as Figure 3 shown, for Figure 1The intersection scene shown is discretized and meshed to construct a state space matrix. It should be noted that the state matrix does not include information about right-turning vehicles or vehicles that have passed through the conflict area. The purpose of this is to focus on the vehicles currently inside the intersection or in front of the conflict area, simplify the state space, and improve the computational efficiency of the algorithm. The state space is represented by matrices of the vehicle's position, speed, and waiting time on the lane. Specifically, the position matrix is used to indicate whether each grid is occupied by a vehicle. A value of 1 or 2 is used to represent the presence of a CAV or HV in the position matrix. When the vehicle occupies more than half of the grid area, the corresponding grid is marked as 1 or 2; otherwise, it is marked as 0. The speed matrix records the speed value of the vehicle when there is a vehicle in the grid. The waiting time matrix records the waiting time value of the vehicle in the grid. This method can effectively reflect the traffic flow state and vehicle state at the intersection.
[0089] Step 2.2: Action Space Design
[0090] The parameterized action space consists of a set of discrete actions, , where each discrete action corresponds to a continuous parameter, ; specifically, the discrete action is represented as all candidate lanes. After selecting one lane as the highest-priority lane, the passing duration of that lane is defined as its corresponding continuous parameter. The action space can be represented as:
[0091]
[0092] Among them, is the set of continuous parameters for all ;
[0093] Step 2.3: Reward Function Design
[0094] Design a multi-objective reward function that comprehensively considers efficiency, fairness, and HV , which is defined as:
[0095]
[0096] Among them, represents punishing the waiting time of the vehicle that stops and waits due to a conflict with the vehicle in the highest-priority lane;
[0097] is aimed at increasing the chance for the vehicle to become a coordinated vehicle so that it can pass through the intersection coordinately instead of being a conflict vehicle and being forced to wait;
[0098] is aimed at increasing the chance for more HVs to become priority vehicles instead of being conflict vehicles or coordinated vehicles, interfering with the safe passage of HVs and the coordination of CAV vehicles in the second layer;
[0099] Among them, represents the waiting time of the th vehicle, represents the waiting time threshold, represents the number of conflicting vehicles, represents the number of coordinated vehicles, represents the CAV ratio on the highest priority lane, , , represents the weight coefficient;
[0100] Step 2.4: Network architecture design
[0101] As Figure 4 shown, the CNN-MPDQN algorithm model structure includes a continuous parameter policy network, a discrete action Q network, and a target network; the traffic high-dimensional spatial features of the intersection are effectively extracted through CNN; specifically, the continuous parameter policy network takes the current state as input, and after being processed by a convolutional layer, a fully connected layer, and an activation function, outputs the optimal vector of all continuous parameters in the action space; subsequently, the vector is separated so that each discrete action only inputs the parameters related to it, thereby reducing the influence of other continuous parameters on the Q-value estimation; in order to keep the length of the continuous action parameter vector consistent, the irrelevant continuous parameters are set to 0, and then the vectors formed by concatenating the current state vector processed by the convolutional layer and the separated continuous actions are input into the discrete action Q network, and after passing through the fully connected layer and the activation function, output matrix, and extract the Q-values on the diagonal to form a vector, and finally obtain the optimal Q-value of the discrete action Q network;
[0102] Step 2.5: Loss function design
[0103] The loss function includes two parts: the discrete action Q network and the continuous parameter policy network; specifically, for the loss function of the discrete action Q network, it is constructed in the form of mean square error, aiming to minimize the target loss:
[0104]
[0105] Among them, is the n-step target function, expressed as:
[0106]
[0107] Among them, is the discount factor, and are the parameters of the discrete action target Q network and the continuous parameter policy target network respectively;
[0108] For the loss function of the continuous parameter policy network, the goal is to maximize the Q-value estimation, which is defined as:
[0109] .
[0110] Step 3: Determine the coordinated lanes and conflicting lanes according to the lane with the highest passing priority selected by the first layer of reinforcement learning; the coordinated lanes are the lane combinations that have no trajectory conflicts with the lane with the highest priority, and the conflicting lanes are the lane combinations that have trajectory conflicts with the lane with the highest priority. As Figure 2 shown, assuming that the S-N (south to north) lane is the lane with the highest priority, then the N-S (north to south), E-N (east to north), and S-E (south to east) lanes are the coordinated lanes, while the other left-turn and straight lanes are the conflicting lanes; during the passing duration, the vehicles on the coordinated lanes can pass through the intersection coordinately, while the vehicles on the conflicting lanes need to stop and wait for the next round of optimization;
[0111] Step 4: Allocate the proportion of the passing duration according to the traffic flow of each coordinated lane respectively, and maximize the utilization of the passing duration to optimize the traffic flow efficiency; among them, the specific calculation formula for the proportion weight of the th coordinated lane is:
[0112]
[0113] Where is the number of coordinated lanes, and are respectively the number of waiting vehicles in the detection area on the th and the th lanes, and are respectively the future predicted traffic flows of the th and the th lanes, which are calculated based on the arrival time of the vehicles, and the arrival time follows a Poisson distribution.
[0114] Step 5: Construct the second-layer deep reinforcement learning algorithm model, adopt the DQN algorithm, and within the passing duration of the first layer, the reinforcement learning determines the passing order of the vehicles on the coordinated lanes, aiming to solve the trajectory conflict problem of the vehicles on the coordinated lanes; the specific state space, action space, reward function, network architecture, and loss function are designed as follows:
[0115] Step 5.1: State space design
[0116] The grid method is used to discretize the intersection. In the second-layer deep reinforcement learning algorithm model, the focus is mainly on the passing order of vehicles on the coordinated lane. Therefore, the state space only represents the state information of vehicles on the coordinated lane and is described by matrices of the positions, desired positions, speeds, and waiting times of vehicles on the coordinated lane. Specifically, the position matrix is used to represent whether each grid is occupied by a vehicle, and the desired position matrix represents the grids that the vehicle will occupy in the future, which is used to reflect the conflict relationship between vehicles. When there is a vehicle in the grid, the speed matrix records the speed value of the vehicle, and the waiting time matrix records the waiting time value of the vehicle in the grid. This design can accurately represent the vehicle state on the coordinated lane and provide a basis for optimizing the passing order.
[0117] Step 5.2: Action space design
[0118] The response action is designed as the combination of passing orders of vehicles on the coordinated lane, which is expressed as:
[0119]
[0120] There are candidate action combinations in total. On this basis, the reinforcement learning algorithm identifies the optimal passing order combination.
[0121] Step 5.3: Reward function design
[0122] The reward function of the second layer is defined to minimize the waiting time and passing time of coordinated vehicles, which is expressed as:
[0123]
[0124] Among them, represents punishing the waiting time of coordinated vehicles, represents punishing the passing time of coordinated vehicles, aiming to make the coordinated vehicles pass through the intersection as soon as possible. Among them, represents the passing time of the th coordinated vehicle, represents the passing time threshold, , represents the weight coefficient.
[0125] Step 5.4: Network architecture design
[0126] The DQN algorithm model structure includes a Q-value network and a target network. Consistent with the first-layer algorithm model, CNN is used as the feature extractor. Specifically, the Q-value network takes the current state as input, and after being processed by the convolutional layer, fully connected layer, and activation function, it outputs the optimal Q-value of the action. The target network is used to prevent the Q-value from being updated too quickly, thereby avoiding unstable training. Specifically, in the way of soft update, the parameters of the target network are synchronized with the Q-value network only every N steps, that is:
[0127]
[0128] Among them, is the target network parameter, is the Q-value network parameter;
[0129] Step 5.5: Loss function design
[0130] The loss function is constructed in the form of mean squared error, aiming to reduce the loss with the n-step objective function:
[0131] .
[0132] Step 6: Under the advanced decision-making of the two-layer reinforcement learning for the passing sequence, the priority passing vehicle passes through the intersection at the maximum safe speed, while the conflicting vehicle needs to decelerate and stop waiting before the conflict position.
[0133] In this step, during the driving process of the priority passing CAV vehicle in the coordinated lane, it may encounter HVs freely driving in other coordinated lanes. To address this situation, the present invention adopts an event-triggered re-decision mechanism: when it is retrieved that the priority CAV vehicle and the HV vehicle may simultaneously occupy the intersection conflict point, the system will re-adjust the maximum safe speed of the passing vehicles in the coordinated lane according to the estimated arrival times of the priority CAV and the HV at the intersection conflict point, and this maximum safe speed can ensure that the priority CAV vehicle and the HV vehicle pass through the intersection conflict area safely without collision.
[0134] Step 7: Through continuous cyclic training and iteration of the two-layer deep reinforcement learning algorithm model, the model parameters are gradually optimized, and finally an optimal convergence model is obtained; this model can effectively cope with various traffic conditions and vehicle behaviors, and improve the passing efficiency and safety of the intersection.
Claims
1. A cooperative control method for mixed traffic flow based on double-layer parameterized deep reinforcement learning, characterized in that, Including the following steps: Construct a road model for a signal-free mixed-flow intersection and divide the intersection area into functional zones; Based on the original multi-channel parameterized deep Q-network MPDQN algorithm, combine it with a convolutional neural network CNN to construct the first-layer deep reinforcement learning algorithm model CNN-MPDQN, and effectively extract the high-dimensional traffic spatial features of the intersection through the CNN; The core task of the CNN-MPDQN algorithm is to select a lane with the highest passing priority from multiple lanes and determine the passing duration of that lane; According to the lane with the highest passing priority selected by the first-layer reinforcement learning, determine the coordinated lane and the conflicting lane; within the passing duration, the vehicles on the coordinated lane can pass through the intersection in coordination, while the vehicles on the conflicting lane need to stop and wait for the next round of optimization; According to the traffic flow of each coordinated lane, allocate the proportion of the passing duration respectively to maximize the utilization of the passing duration and optimize the traffic flow efficiency; Construct the second-layer deep reinforcement learning algorithm model, adopt the deep Q-network DQN algorithm, and within the passing duration of the first layer, use reinforcement learning to determine the passing order of the vehicles on the coordinated lane, aiming to solve the trajectory conflict problem of the vehicles on the coordinated lane; Under the high-level decision of the passing order of the two-layer reinforcement learning, the priority passing vehicles pass through the intersection at the maximum safe speed, while the conflicting vehicles need to decelerate and stop before the conflict position; Through continuous cyclic training and iteration of the two-layer deep reinforcement learning algorithm model, gradually optimize the model parameters, and finally obtain the optimal convergence model; this model can effectively handle various traffic conditions and vehicle behaviors, and improve the passing efficiency and safety of the intersection.
2. The collaborative control method for mixed traffic flow based on double-layer parameterized deep reinforcement learning according to claim 1, characterized in that: The constructed road model is a four-way three-lane intersection, with left-turn, straight-ahead, and right-turn lanes provided in each direction, and the right-turn lane is not controlled; the intersection area is divided into a detection area and a conflict area; the section between the maximum communication range and the stop line is the detection area, and the conflict area is the area where vehicles may conflict with vehicles in other directions after crossing the stop line; both connected autonomous vehicles CAVs and human-driven vehicles HVs coexist on the road. Vehicles entering the detection area and communicating with the central processor of the intersection are identified as CAVs, while vehicles that can be detected by the central processor but do not communicate are identified as HVs; all vehicles travel according to the pre-planned trajectories.
3. The collaborative control method for mixed traffic flow based on double-layer parameterized deep reinforcement learning according to claim 2, characterized in that: The driving behaviors of the HVs are divided into two types: following the vehicle in front and driving freely without a vehicle in front; when two free HVs from conflicting directions arrive at the intersection simultaneously, they only need to follow the right-hand rule to pass through the intersection; if there is a vehicle in front of the HV, to avoid collision, the HV needs to maintain a safe distance. Therefore, the target speed of the HV at time t + Δt should satisfy the following conditions: v des (t + Δt) = min{v max , v(t) + a max Δt, v safe (t + Δt)} Among them, v max is the maximum speed limit of the road, a max is the maximum acceleration, Δt is the simulation update step, v safe is the safe following speed, v safe The formula is: Among them, v pre (t) is the speed of the vehicle ahead at time t, d(t) is the vehicle distance at time t, and τ is the reaction time of a human driver; In addition, a random factor ε ∈ [0, 1] is introduced to represent the uncertainty of human driving behavior, and the final following speed is: v f v(t + Δt) = max{0, v des (t + Δt) - rand(0, εa max )}。 4. The hybrid traffic flow collaborative control method based on double-layer parameterized deep reinforcement learning according to claim 1, characterized in that: For the first-layer deep reinforcement learning algorithm model, the CNN-MPDQN algorithm is adopted. The specific state space, action space, reward function, network architecture and loss function are designed as follows: 4.1 State space design To intuitively present the current traffic conditions at the intersection, the grid method is used to discretize the intersection, where the side length of the grid is set to the lane width; the state space is represented by a matrix of the position, speed and waiting time of vehicles on the lane; specifically, the position matrix is used to represent whether each grid is occupied by a vehicle, and 1 or 2 is used to represent the presence of a CAV or HV in the position matrix. When the vehicle occupies more than half of the grid area, the corresponding grid is marked as 1 or 2, otherwise it is marked as 0; the speed matrix records the speed value of the vehicle when there is a vehicle in the grid; the waiting time matrix records the waiting time value of the vehicle in the grid. This method can effectively reflect the traffic flow state and vehicle state at the intersection. 4.2 Action space design The parameterized action space consists of a set of discrete actions, Α d = [K] = [a1, a2,..., a k , where each discrete action corresponds to a continuous parameter, x k ∈ χ k ; specifically, the discrete actions are represented as all candidate lanes. After selecting one lane as the highest-priority lane, the passing duration of that lane is defined as its corresponding continuous parameter; this action space can be represented as: where χ k is the set of continuous parameters for all a ∈ Α d ; 4.3 Reward function design A multi-objective reward function R1 that comprehensively considers efficiency, fairness and HV is designed, and its definition is: R1 = -(αr w1 + βr k + ηr h ) Among them, represents the waiting time of the vehicle that is penalized for conflicting with the vehicle in the highest-priority lane and thus has to wait and stop. Aim to increase the chance of a vehicle becoming a coordinated vehicle so that it can pass through the intersection in coordination, rather than becoming a conflicting vehicle and being forced to wait; r h = r p The aim is to increase the opportunities for more HVs to become priority vehicles rather than conflict or coordinated vehicles, interfering with the safe passage of HVs and the coordination of CAV vehicles at the second layer; Among them, w i represents the waiting time of the i-th vehicle, T m represents the waiting time threshold, n c represents the number of conflicting vehicles, n k represents the number of coordinated vehicles, r p represents the CAV ratio on the highest-priority lane, and α, β, η represent weight coefficients; 4.4 Network architecture design The CNN-MPDQN algorithm model structure includes a continuous parameter policy network, a discrete action Q network, and a target network; the traffic high-dimensional spatial features of the intersection are effectively extracted by CNN; specifically, the continuous parameter policy network takes the current state as input, and after being processed by a convolutional layer, a fully connected layer, and an activation function, outputs the optimal vector of all continuous parameters in the action space; subsequently, the vector is separated so that each discrete action only inputs the parameters related to it, thereby reducing the influence of other continuous parameters on the Q-value estimation; in order to keep the length of the continuous action parameter vector consistent, the irrelevant continuous parameters are set to 0, and then k vectors formed by concatenating the current state vector processed by the convolutional layer and the separated continuous actions are input into the discrete action Q network. After passing through the fully connected layer and the activation function, it outputs a matrix, and extracts the Q-values on the diagonal to form a vector, and finally obtains the optimal Q-value of the discrete action Q network; 4.5 Loss function design The loss function includes two parts: a discrete action Q network and a continuous parameter policy network; specifically, for the loss function of the discrete action Q network, it is constructed in the form of mean square error, aiming to minimize the target loss; for the loss function of the continuous parameter policy network, the goal is to maximize the Q-value estimation.
5. The hybrid traffic flow collaborative control method based on double-layer parameterized deep reinforcement learning according to claim 1, characterized in that: The coordinated lanes are lane combinations that do not have trajectory conflicts with the highest-priority lane, and the conflicting lanes are lane combinations that have trajectory conflicts with the highest-priority lane.
6. The hybrid traffic flow collaborative control method based on double-layer parameterized deep reinforcement learning according to claim 1, characterized in that: According to the traffic flow of each coordinated lane, the proportion of the passing duration is allocated respectively. The specific calculation formula for the proportion weight of the i-th coordinated lane is: Among them, N k To coordinate the number of lanes, and are the number of waiting vehicles in the detection areas on the i-th and j-th lanes respectively, and are the future predicted traffic flows on the i-th and j-th lanes respectively, calculated based on the arrival times of vehicles, and the arrival times follow a Poisson distribution.
7. The hybrid traffic flow collaborative control method based on double-layer parameterized deep reinforcement learning according to claim 4, characterized in that: For the second-layer deep reinforcement learning algorithm model, the DQN algorithm is adopted. During the passing duration of the first layer, the reinforcement learning algorithm determines the passing order of vehicles on the coordinated lanes; the specific state space, action space and reward function of this model are designed as follows: 7.1 State space design The grid method is used to discretize the intersection. In the second-layer deep reinforcement learning algorithm model, the passing sequence of vehicles on the coordinated lane is mainly concerned. Therefore, the state space only represents the state information of the vehicles on the coordinated lane and is described by matrices of the positions, desired positions, speeds, and waiting times of the vehicles on the coordinated lane. Specifically, the position matrix is used to represent whether each grid is occupied by a vehicle, and the desired position matrix represents the grids that the vehicle will occupy in the future, which is used to reflect the conflict relationship between vehicles. The speed matrix records the speed value of the vehicle when there is a vehicle in the grid, and the waiting time matrix records the waiting time value of the vehicle in the grid. This design can accurately represent the vehicle state on the coordinated lane and provide a basis for optimizing the passing sequence. 7.2 Action Space Design The response action is designed as the combination of the passing sequences of vehicles on the coordinated lane, which is expressed as: There are a total of N k ! candidate action combinations; based on this, the reinforcement learning algorithm identifies the optimal passing order combination; 7.3 Reward Function Design The reward function of the second layer is defined to minimize the waiting time and passing time of the coordinated vehicles, which is expressed as: Among them, represents the penalty for the waiting time of the coordinated vehicle, represents the penalty for the passing time of the coordinated vehicle, aiming to make the coordinated vehicle pass through the intersection as soon as possible; among them, t i represents the passing time of the i-th coordinated vehicle, t m represents the passing time threshold, δ, represents the weight coefficient; 7.4 Network Architecture Design The DQN algorithm model structure includes a Q-value network and a target network. Consistent with the first-layer algorithm model, CNN is used as the feature extractor. Specifically, the Q-value network takes the current state as the input, and after being processed by the convolutional layer, fully connected layer, and activation function, it outputs the optimal Q-value of the action. The target network is used to prevent the Q-value from being updated too quickly, thereby avoiding unstable training. Specifically, in the form of soft update, the parameters of the target network are synchronized with the Q-value network only every N steps, that is: ω' - ←ω' where ω' - is the target network parameter, and ω' is the Q-value network parameter; 7.5 Loss Function Design The loss function is constructed in the form of mean squared error, aiming to reduce the loss with the n-step target function.
8. The hybrid traffic flow cooperative control method based on double-layer parameterized deep reinforcement learning according to claim 1, characterized in that: Under the high-level decision of the passing sequence in the two-layer reinforcement learning, the priority passing vehicle passes through the intersection at the maximum safe speed, while the conflicting vehicle needs to decelerate and stop waiting before the conflict position. During the driving process of the priority passing CAV in the coordinated lane, it may encounter the freely driving HV in other coordinated lanes. To cope with this situation, the method adopts an event-triggered re-decision mechanism.
9. The hybrid traffic flow cooperative control method based on double-layer parameterized deep reinforcement learning according to claim 8, characterized in that: The method adopts an event-triggered re-decision mechanism. When it is retrieved that the priority CAV and the HV may simultaneously occupy the intersection conflict point, the system will re-adjust the maximum safe speed of the passing vehicles in the coordinated lane according to the estimated time for the priority CAV and the HV to reach the intersection conflict point. This maximum safe speed can ensure that the priority CAV and the HV do not collide and safely pass through the intersection conflict area.
Citation Information
Patent Citations
Substation operation and maintenance decision-making method based on multi-agent reinforcement learning
CN118644225A
Vehicle-annunciator cooperative signal control method based on double-layer AMOC
CN118692250A