A vehicle queuing control method based on an improved gap strategy and a depth deterministic strategy gradient algorithm
By improving the vehicle spacing strategy and constructing the MPC framework using the Deep Deterministic Strategy Gradient Algorithm (DDPG), and adjusting parameters in real time, the flexibility problem of vehicle spacing strategy in the CACC system was solved, and efficient and safe vehicle control in dynamic traffic environments was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2026-04-03
AI Technical Summary
In the existing CACC system, the vehicle platoon spacing strategy cannot flexibly respond to different driving conditions, resulting in insufficient platoon stability and safety. Furthermore, the fixed parameter controller has difficulty optimizing vehicle spacing in dynamic traffic environments, affecting traffic efficiency and safety.
An improved inter-vehicle clearance strategy is adopted to construct the MPC framework, and the parameters are adjusted in real time by combining the deep deterministic policy gradient algorithm (DDPG). By considering the speed difference between the preceding vehicle and the following vehicle, the longitudinal control of the vehicle platoon is optimized, and the minimum safe distance between the vehicles is added as a constraint to achieve adaptive control.
It improves the safety and following efficiency of the fleet in dynamic traffic environments, reduces vehicle gap errors, ensures the stability and safety of the vehicle platoon, and optimizes the management of autonomous driving fleets.
Smart Images

Figure CN120544373B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of connected autonomous vehicles and traffic control, specifically to a vehicle queuing control method based on an improved gap strategy and a deep deterministic strategy gradient algorithm. Background Technology
[0002] In recent years, with the rapid development of social civilization and the surge in the number of cars, traffic congestion and traffic safety problems have become increasingly serious. To address these issues, Advanced Driver Assistance Systems (ADAS) have gradually become a research hotspot. Adaptive Cruise Control (ACC) is a typical example, using sensors, radar, and other equipment to perceive road conditions and vehicles ahead, enabling automatic cruise control and following, reducing driver workload and energy consumption. However, with advancements in communication technology, Cooperative Adaptive Cruise Control (CACC), based on ACC, has emerged. CACC utilizes vehicle-to-vehicle communication to obtain more accurate vehicle status information and enable multi-vehicle platooning, thereby improving traffic efficiency and safety.
[0003] In the CACC system, the spacing strategy between platoons has a decisive impact on the stability and efficiency of the entire system, and the size of the vehicle platoon is significantly affected by the spacing strategy. If the platoon spacing is too large, it not only wastes road resources but also makes it easy for other vehicles to enter the platoon, affecting platoon stability; conversely, too small a spacing may increase the risk of rear-end collisions, causing driver stress and affecting safety and comfort. Therefore, maintaining an appropriate spacing is key to ensuring safety, comfort, and energy conservation.
[0004] Furthermore, spacing strategies are not only affected by factors such as vehicle speed and road conditions, but are also closely related to the vehicle's dynamic characteristics. The appropriateness of the controller parameter settings directly determines the vehicle's following behavior pattern. When the vehicle's state changes, fixed controller parameters often cannot cope with diverse driving needs; therefore, selecting suitable controller parameters for different driving states is particularly important. Summary of the Invention
[0005] To address the aforementioned problems in existing technologies, a vehicle queuing control method based on an improved gap strategy and a depth deterministic strategy gradient algorithm is provided. The improved gap strategy is used as the model gap strategy to construct the MPC framework. At the same time, the depth deterministic strategy gradient algorithm is used to adjust the parameters of the model gap strategy and the key control parameters in the MPC framework, enabling the vehicle to dynamically adjust the control strategy according to the real-time vehicle status and road environment.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0007] A vehicle queuing control method based on an improved gap strategy and a depth deterministic strategy gradient algorithm includes the following steps:
[0008] Establish a vehicle gap strategy that considers the speed difference between the vehicle in front and the vehicle itself, construct an MPC framework for the vehicle queue using the aforementioned vehicle gap strategy, and incorporate the minimum safe distance between vehicles into the constraints of the MPC framework.
[0009] The trained deep deterministic policy gradient algorithm model is used to adjust the parameters in the MPC framework in real time to obtain the optimal parameter combination for each state. Then, the optimal parameter combination is substituted into the MPC framework to calculate the optimal control sequence for each time step. Finally, the vehicles in the vehicle queue are longitudinally controlled according to the specified control quantity in the optimal control sequence of each control cycle.
[0010] Furthermore, the calculation formula for the gap strategy considering the speed difference between the preceding vehicle and the following vehicle is as follows:
[0011]
[0012] in, This represents the expected gap between the vehicles in front in the convoy. and Let these represent the speeds of vehicle n and vehicle n-1 in the vehicle queue, respectively. For the desired time interval, Indicates the minimum safe distance. It is the minimum spacing correction term. These are model parameters. This indicates the speed difference between the vehicle in front and the vehicle in front.
[0013] Furthermore, when constructing the MPC framework for the vehicle queue using the aforementioned workshop gap strategy, the specific steps include:
[0014] For each vehicle in the vehicle platoon, a corresponding single-vehicle longitudinal dynamic model is established, taking into account the speed change of the vehicle in front of each vehicle.
[0015] Replace the expected gap of the model corresponding to the MPC framework with the workshop gap strategy;
[0016] Comfort, efficiency, and control input are all taken as optimization control objectives. The objective function and constraints of the vehicle longitudinal dynamics model that takes into account the speed change of the vehicle in front are constructed as the objective function and constraints of the MPC framework.
[0017] Furthermore, the expression for the longitudinal dynamic model of a single vehicle is as follows:
[0018]
[0019]
[0020]
[0021] Where X is the system state variable, U is the system control input, and d is the system disturbance. For the system input, A, B1, and B2 are all coefficient matrices of the state-space equations. The acceleration of the vehicle in front of it. These represent the difference between the actual distance and the expected distance between vehicle n and the vehicle in front in the vehicle queue, the relative speed between vehicle n and the vehicle in front, the acceleration of vehicle n itself, and the velocity of vehicle n itself. Let be the engine time constant for vehicle n. This represents the desired time interval.
[0022] Furthermore, when simultaneously considering comfort, efficiency, and control input as optimization control objectives, and constructing the objective function and constraints for a vehicle longitudinal dynamics model that takes into account changes in the speed of the vehicle ahead, the steps for constructing the objective function include:
[0023] The longitudinal dynamic model of a single vehicle is discretized using the forward Euler method, as shown in the following expression:
[0024]
[0025] In the formula:
[0026]
[0027] Indicates that the system is in The state at any given moment The system sampling time, To control the input, For interference input;
[0028] The objective function is defined as follows:
[0029]
[0030]
[0031] in, represent At this moment The predicted state value at time 10:00. represent At this moment The state reference value at any given time. represent The control input at any given time, i.e., the decision variables of the model. It is a weighted coefficient matrix of spacing error, relative vehicle speed, acceleration, and the derivative of acceleration. These are the weighting coefficients of the control variables. It is the prediction time domain. It controls the time domain.
[0032] Furthermore, when constructing the objective function and constraints of a vehicle longitudinal dynamics model that considers changes in the speed of the vehicle in front, simultaneously taking comfort, efficiency, and control input as optimization control objectives, the steps for constructing the constraints include:
[0033] A relaxation factor is introduced to softly constrain different state variables. The specific constraint relationships are as follows:
[0034]
[0035]
[0036]
[0037]
[0038]
[0039]
[0040]
[0041] in, This represents a 5×5 identity matrix. Represents the slack variables of the model. It refers to a diagonal matrix composed of slack variables. Indicates the lower bound of the constraint. Indicates the upper bound of the constraint, Indicates the lower bound of the slack variable, Indicates the upper bound of the slack variable. , , , , These represent the minimum values of clearance error, relative velocity, acceleration, speed, and control input, respectively. , , , , These represent the maximum values of gap error, relative velocity, acceleration, speed, and control input, respectively. It is a relaxation factor. and It is the relaxation factor for the upper and lower limits of system constraints.
[0042] Furthermore, when incorporating the minimum safe distance between workshops into the constraints of the MPC framework, the maximum parking distance is specifically used as the minimum safe distance between workshops, and the minimum safe distance between workshops is used as the vehicle safety constraint in the constraints. The expression for the maximum parking distance is as follows:
[0043]
[0044] in, It is the delay time of the vehicle braking command. It is the vehicle's maximum deceleration. It is the set minimum following distance. This represents the speed of vehicle n. This represents the speed of vehicle n-1.
[0045] Furthermore, before using the trained deep deterministic policy gradient algorithm model to adjust the parameters in the MPC framework in real time, the steps include training the deep deterministic policy gradient algorithm model, including:
[0046] Real-time acquisition of the clearance error between the current vehicle and the vehicle in front. Current speed difference with the vehicle in front Current acceleration and current speed And update the values of the corresponding parameters in the state space to obtain the current state of the vehicle at the current moment. The current state Input the Actor network to obtain the best action at the current time. The best action at the current moment The expression is as follows:
[0047]
[0048] in, The gap error weighting coefficient of the MPC framework. These are the model parameters for the workshop gap strategy;
[0049] The best action at the current moment After passing through the Critic network, the reward at the current time step is output. The specific representation of the reward at the current time step is as follows:
[0050]
[0051] in, This represents the weighting coefficient of the reward. This is a bonus item for gap error deviation. This is a reward item for relative speed deviation. For the steady-state reward term, the expression is:
[0052]
[0053]
[0054] in, Indicates gap error, Represents the absolute value of the gap error. This represents the absolute value of the relative speed with the vehicle in front. Furthermore, when using the trained deep deterministic policy gradient algorithm model to adjust the parameters in the MPC framework in real time, specifically, the vehicle's own state is acquired at each time step, and this state is input into the Actor network to obtain the action that maximizes the reward function calculated by the Critic network under the vehicle's own state. The gap error weight coefficients of the MPC framework included in this action are then considered. and model parameters for workshop gap strategy As the optimal combination of parameters for the vehicle's own condition.
[0055] Furthermore, the specified control quantity in the control sequence of each control cycle specifically refers to the first control quantity in the control sequence.
[0056] Compared with the prior art, the advantages of the present invention are as follows:
[0057] This invention constructs an MPC framework by considering the gap between the speed of the preceding vehicle and the speed of the vehicle itself, and adds the minimum safe distance between vehicles to the constraints of the MPC framework. This makes the MPC framework predict based on the improved gap strategy, which is more sensitive to changes in the speed of the preceding vehicle and improves the tracking performance. At the same time, it considers the minimum braking distance to prevent rear-end collisions, ensuring vehicle stability and safety.
[0058] This invention employs a Deep Deterministic Policy Gradient (DDPG) algorithm model to adjust the parameters in the MPC framework in real time, ensuring that the optimal controller parameter settings are obtained at every moment, thereby achieving adaptive control. This can effectively reduce the gap error of the fleet and optimize the longitudinal control of the fleet in dynamic traffic environments, thus achieving efficient and safe autonomous driving fleet management. Attached Figure Description
[0059] Figure 1 This is a flowchart of a method according to an embodiment of the present invention.
[0060] Figure 2 This is a control structure diagram of the deep deterministic strategy gradient algorithm model in this embodiment of the invention for real-time adjustment of parameters in the MPC framework.
[0061] Figure 3 This is a schematic diagram of an example scenario of the present invention.
[0062] Figure 4 The gap error curve is for the MPC method using fixed parameters.
[0063] Figure 5 This is the gap error curve of the method in an embodiment of the present invention.
[0064] Figure 6 The curve shows a comparison of vehicle queue length between the method of this invention and the MPC method using fixed parameters.
[0065] Figure 7 This is a curve showing the change in the gap error weighting coefficient of the method in an embodiment of the present invention.
[0066] Figure 8 This is a curve showing the change in model parameters of the method in an embodiment of the present invention. Detailed Implementation
[0067] The present invention will be further described below with reference to the accompanying drawings and specific preferred embodiments, but this does not limit the scope of protection of the present invention.
[0068] Before introducing specific embodiments of the present invention, the relevant concepts will be explained first.
[0069] DDPG (Deep Deterministic Policy Gradient) is a deep reinforcement learning algorithm based on the Actor-Critic framework, used to handle problems with continuous action spaces. DDPG combines policy gradients and a value function, employing a four-network structure: an online policy Actor network, a target policy Actor network, an online policy Critic network, and a target policy Critic network. The online policy Critic network is updated by minimizing the mean squared error, the online policy Actor network is updated by maximizing the Q-value, and finally, the current network parameters are updated using a soft update policy.
[0070] MPC (Model Predictive Control) is an advanced control strategy based on a dynamic model. It generates an optimal control sequence by predicting the dynamic model of the system, thereby achieving closed-loop control of the system. MPC first establishes a dynamic model of the system to predict its future state under different control inputs. Then, it sets an optimization objective function, typically including tracking error and control input, to quantify the degree to which the control objective is achieved. Simultaneously, it considers the system's physical and operational constraints to ensure that the control input and system state are within reasonable ranges. Finally, at each sampling time, based on the current system measurements and the predictive model, it solves a finite-time open-loop optimization problem online to obtain the control sequence, and applies the first element of the control sequence to the controlled object.
[0071] CTH Strategy: Constant Time Headway (CTH) is a spacing control strategy widely used in vehicle queuing control and adaptive cruise control (ACC). CTH maintains a fixed time interval between vehicles, rather than a fixed distance. Specifically, the time interval between the following vehicle and the preceding vehicle remains constant, while the actual distance between vehicles changes with vehicle speed. This strategy maintains relative stability between vehicles at different speeds, effectively ensuring vehicle safety, especially at high speeds. While CTH performs well at high speeds, at low speeds, the relatively large distance between vehicles may lead to reduced road utilization.
[0072] Example
[0073] This embodiment proposes a vehicle longitudinal model predictive control method based on an improved gap policy and a deep deterministic policy gradient (DDPG) algorithm, called DDPG-MPC. This method combines an improved gap policy with the deep deterministic policy gradient (DDPG) algorithm to achieve adaptive optimization control of the vehicle's longitudinal direction in a CACC system. The core lies in using a reinforcement learning-based DDPG model to dynamically adjust the gap error weight parameters and gap policy model parameters of the MPC framework, enabling the vehicle to optimize its following behavior in complex traffic environments and improve safety, comfort, and energy efficiency.
[0074] like Figure 1 As shown, the method in this embodiment specifically includes the following steps:
[0075] S101) MPC Framework Construction Phase: Establish a vehicle gap strategy that considers the speed difference between the preceding vehicle and the following vehicle. Use this gap strategy to construct the MPC framework for the vehicle queue, and incorporate the minimum safe distance between vehicles into the constraints of the MPC framework. Specific steps include:
[0076] S1) Establish a gap strategy that considers the speed difference between the preceding vehicle and the following vehicle as an improved gap strategy, and at the same time establish an equation for the minimum safe distance between vehicles.
[0077] S2) Establish a longitudinal dynamics model of the vehicle and construct an MPC framework in combination with an improved clearance strategy, adding the minimum safe distance between the workshops as a strict constraint condition;
[0078] S102) MPC Framework Application Stage: The trained Deep Deterministic Policy Gradient Algorithm (DDPG) model is used to adjust the parameters in the MPC framework in real time to obtain the optimal parameter combination for each state. Then, the optimal parameter combination is substituted into the MPC framework to calculate the optimal control sequence for each time step. Finally, the vehicles in the vehicle queue are longitudinally controlled according to the specified control quantity in the optimal control sequence of each control cycle. The specific steps include:
[0079] S3) Train the DDPG model and use the trained DDPG model to adjust the gap error weight parameters and gap strategy model parameters in the MPC framework in real time to find the optimal parameter combination for each state.
[0080] S4) In a dynamic traffic environment, the corresponding optimal parameter combination is substituted into the MPC framework, and the optimal control sequence at each time step is obtained by solving the quadratic programming problem.
[0081] S5) Perform longitudinal control of the vehicle according to the specified control quantity in the control sequence of each control cycle.
[0082] The following provides a detailed explanation of each step.
[0083] This embodiment establishes an improved workshop gap strategy through step S1, specifically including the following steps:
[0084] S1.1) Based on the CTH interval strategy, a vehicle spacing strategy that considers the speed difference between the preceding vehicle and the following vehicle is designed. Considering the influence of relative speed on the vehicle spacing, a speed difference-related (VDR) vehicle spacing strategy is proposed to reduce unnecessary changes.
[0085] The CTH interval strategy calculation formula is as follows:
[0086]
[0087] in, This represents the expected gap between the vehicles in front in the convoy. This represents the speed of vehicle n in the vehicle queue. For the desired time interval, This indicates the minimum safe distance.
[0088] The formula for calculating the gap strategy considering the speed difference between the preceding vehicle and the following vehicle is as follows:
[0089]
[0090] in, It is the minimum spacing correction term. These are model parameters. Indicates the speed difference with the vehicle in front. This represents the speed of vehicle n-1 in the vehicle queue.
[0091] Furthermore, considering the minimum braking distance between vehicles and to avoid rear-end collisions, the following equation regarding the minimum safe distance between vehicles should be satisfied:
[0092]
[0093] in It is the delay time of the vehicle braking command. It is the vehicle's maximum deceleration. It is the set minimum following distance.
[0094] S1.2) When the speed of the vehicle ahead is stable, the spacing error between the following vehicle's spacing strategy and the actual spacing must remain convergent; otherwise, the stability of the following vehicle's CACC system cannot be guaranteed. Therefore, a stability analysis of the spacing strategy is required. That is, when... hour, .
[0095] in, , This represents the actual clearance in the workshop. By differentiating both sides, we can obtain:
[0096]
[0097] Analysis of the situation while following another vehicle shows that when At that time, the vehicle speed should be greater than the speed of the vehicle in front, that is Similarly, when At that time, the vehicle speed should be less than the speed of the vehicle in front. Therefore, the relationship between spacing error and relative velocity can be obtained as follows:
[0098]
[0099] in, These are model constant parameters. When the system is in a steady state, By differentiating both sides, we can obtain:
[0100]
[0101] Therefore, we can obtain
[0102]
[0103]
[0104] In conclusion, when hour, ;when hour, Therefore, the spacing error will eventually converge to 0, proving that the designed spacing strategy is stable.
[0105] In step S2 of this embodiment, the improved gap strategy constructed in step S1 is considered to be used to construct the control framework of MPC, and the minimum safe distance between workshops is added as a strict constraint to the constraints of MPC. Specifically, the following steps are included:
[0106] S2.1) For each vehicle in the connected autonomous vehicle fleet, a corresponding longitudinal dynamics model for each vehicle is established, considering the speed change of the vehicle in front of it. The expression for the longitudinal dynamics model of each vehicle is as follows:
[0107]
[0108]
[0109]
[0110] Where X is the system state variable, U is the system control input, and d is the system disturbance. For the system input, A, B1, and B2 are all coefficient matrices of the state-space equations. The acceleration of the vehicle in front of it. These represent the difference between the actual distance and the expected distance between vehicle n and the vehicle in front in the vehicle platoon, respectively; the relative speed between vehicle n and the vehicle in front; the acceleration of vehicle n itself; and the speed of vehicle n itself, where n is the vehicle's sequence number in the connected autonomous vehicle platoon. This represents the engine time constant of the vehicle. This represents the desired time interval.
[0111] S2.2) Use the improved workshop clearance strategy as the model expected clearance corresponding to the MPC framework. Specifically, replace the traditional CTH clearance strategy in the MPC framework with the workshop clearance strategy that considers the speed difference between the preceding vehicle and the following vehicle constructed in step S1.
[0112] S2.3) Comfort, efficiency, and control input are simultaneously used as optimization control objectives. The objective function and constraints of the vehicle longitudinal dynamics model considering the speed change of the preceding vehicle are constructed as the objective function and constraints of the MPC framework.
[0113] From a safety and following performance perspective, theoretically, the vehicle in the middle of the pack should maintain the same speed as the vehicle in front, while keeping the following gap error close to zero. The specific control objective can be expressed as:
[0114]
[0115] For CACC vehicles equipped with both radar and accelerometers, their state space All state variables are measurable; therefore, after discretizing the vehicle dynamics model using the forward Euler method, the expression is as follows:
[0116]
[0117] In the formula
[0118]
[0119] Indicates that the system is in The state at any given moment The system sampling time, To control the input, This is for interference input.
[0120] The objective function is set to the following form:
[0121]
[0122]
[0123] in, represent At this moment The predicted state value at time 10:00. represent At this moment The state reference value at any given time. represent The control input at any given time, i.e., the decision variables of the model. It is a weighted coefficient matrix of spacing error, relative vehicle speed, acceleration, and the derivative of acceleration. These are the weighting coefficients of the control variables. It is the prediction time domain. It controls the time domain.
[0124] During system optimization, constraints must be imposed on each state variable. However, overly strict constraints may prevent the controller from solving. Therefore, this embodiment introduces relaxation factors to softly constrain different state variables. The specific constraint relationships are as follows:
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] in, This represents a 5×5 identity matrix. Represents the slack variables of the model. It refers to a diagonal matrix composed of slack variables. Indicates the lower bound of the constraint. Indicates the upper bound of the constraint, Indicates the lower bound of the slack variable, Indicates the upper bound of the slack variable. , , , , These represent the minimum values of clearance error, relative velocity, acceleration, speed, and control input, respectively. Similarly, , , , , These represent the maximum values of gap error, relative velocity, acceleration, speed, and control input, respectively. It is a relaxation factor. and It is the relaxation factor for the upper and lower limits of system constraints.
[0133] S2.4) Incorporating the minimum safe distance between vehicles into the constraints: To prevent rear-end collisions caused by a small clearance between vehicles in case of emergencies during CAV operation, this embodiment introduces the maximum stopping distance as the minimum safe distance between vehicles, and uses the minimum safe distance between vehicles as a vehicle safety constraint in the constraints. The expression for the maximum stopping distance is as follows:
[0134]
[0135] in, It is the delay time of the vehicle braking command. It is the vehicle's maximum deceleration. It is the set minimum following distance. This represents the speed of vehicle n. This represents the speed of vehicle n-1.
[0136] In step S3 of this embodiment, the DDPG model is first trained, and then the trained DDPG model is used to adjust the key parameters in the MPC framework in real time to find the optimal parameter combination for each state. Specifically, this includes the following steps:
[0137] S3.1) DDPG model training phase:
[0138] The DDPG algorithm performs actions in the environment by exploring or utilizing existing strategies and stores the experience in an experience replay buffer. It then randomly samples a batch of samples from the experience replay buffer. Training was conducted. In this context... Representing CAV in the The state at any given moment, This represents CAV in status. The action to be performed This represents the reward for CAV taking this action. Actions taken on behalf of CAV The new state after training. During training, a soft update strategy is used to gradually adjust the target network parameters to stabilize the learning process:
[0139]
[0140]
[0141] in, These are soft update coefficients used to control the update rate, gradually bringing the parameters of the target network closer to those of the main network. This provides more stable target values, reduces fluctuations during training, and improves the stability and effectiveness of the algorithm.
[0142] Repeat the above steps until the algorithm meets the termination condition.
[0143] Specifically, this embodiment first uses vehicle radar, vehicle-to-vehicle information interaction, and the vehicle's own speed and acceleration detectors to obtain the real-time clearance error between the current vehicle and the vehicle in front. Current speed difference with the vehicle in front Current acceleration and current speed Therefore, the state space is set to include the clearance error between the vehicle and the vehicle in front. Speed difference with the vehicle in front Vehicle acceleration and vehicle speed Specifically, it is expressed as:
[0144]
[0145] Then, the DDPG reinforcement learning method is used to adjust the controller parameters, including: adjusting the vehicle's current state. Input the Actor network to obtain the best action at the current time. The best action at the current moment The expression is as follows:
[0146]
[0147] It is evident that the optimal action Includes two parameters, among which The gap error weighting coefficient of the MPC framework. These are the model parameters for the workshop gap strategy; subsequent real-time parameter adjustments aim to obtain the optimal values of these two parameters. The optimal action at the current moment... After passing through the Critic network, the current time-time reward is output. The current time-time reward consists of the gap error deviation reward item. Relative speed deviation bonus item and steady-state reward items The composition is specifically represented as follows:
[0148]
[0149]
[0150] in, Indicates gap error, Represents the absolute value of the gap error. It represents the absolute value of the relative speed with respect to the vehicle in front.
[0151] The final model reward function is expressed as follows:
[0152]
[0153] in, This represents the weighting coefficient of the reward.
[0154] S3.2) DDPG model application stage:
[0155] like Figure 2 As shown, for a trained DDPG model, when using this model to adjust the parameters in the MPC framework in real time, specifically, the vehicle's own state is obtained at each time step, and the vehicle's own state is input into the Actor network to obtain the action that maximizes the reward function calculated by the Critic network under the vehicle's own state. The gap error weight coefficients of the MPC framework included in the action are then applied. and model parameters for workshop gap strategy As the optimal combination of parameters for the vehicle's own state, using the optimal combination of parameters to set the MPC can obtain the control input at that moment and execute it, thereby obtaining the vehicle's own state at the next moment. By repeating this process, the optimal controller parameter settings can be obtained at each moment.
[0156] In step S4 of this embodiment, in a dynamic traffic environment, the corresponding optimal parameter combination is substituted into the MPC framework, and the optimal control sequence at each time moment is obtained by solving a quadratic programming problem. This process is a common solution to a quadratic programming problem, and the specific solution process is well known in the art and is not the focus of the method in this embodiment, so it will not be described in detail.
[0157] In step S5 of this embodiment, the specified control quantity in the control sequence of each control cycle specifically refers to the first control quantity in the control sequence. Since MPC solves an optimization problem and obtains a control sequence in each control cycle, only the first control quantity of this sequence is executed, and then this process is repeated in the next cycle. Through rolling optimization, MPC can adaptively respond to changes in the system, thereby achieving precise control of the vehicle queue.
[0158] The effectiveness of the method in this embodiment will be verified through specific experiments below.
[0159] This embodiment selects two evaluation indicators: the maximum gap error and the sum of gap errors, as shown below:
[0160]
[0161]
[0162] In the formula This represents the total simulation duration.
[0163] To demonstrate the significant effectiveness of the proposed method (DDPG-MPC) in this embodiment, a simulation experiment using Python is conducted, considering a vehicle convoy of four vehicles, including the lead vehicle. Figure 3 As shown in the figure. A standard MPC control method was used as a comparative experiment, with specific parameter settings as shown below.
[0164] Table 1. Experimental Parameter Values
[0165]
[0166] The acceleration settings for the lead vehicle in the convoy are shown below, and the simulation lasts for 40 seconds.
[0167]
[0168] The acceleration curves of the four following vehicles are shown in the figure. Figure 4 The gap error curve for the fixed-parameter MPC method. Figure 5 This is the gap error curve of the DDPG-MPC method proposed in this embodiment.
[0169] Figure 4 The data shows that the maximum clearance errors of the three following vehicles with fixed parameters MPC are 0.0726, 0.0722, and 0.0716, respectively, with an absolute sum of clearance errors of 8.6932, 8.6931, and 8.6930. In contrast, through... Figure 5 It can be seen that the maximum gap errors of the DDPG-MPC method proposed in this embodiment are 0.0641, 0.06248, and 0.0611, respectively, and the absolute sum of the gap errors is 5.7935, 5.6762, and 5.5545. The results show that the vehicle platoon controlled by the DDPG-MPC method proposed in this embodiment has smaller gap errors, and the following between CAVs is more efficient and safer. Furthermore, when the state of the lead vehicle changes, the gap error of subsequent vehicles is continuously reduced during backpropagation, verifying that the control algorithm meets the requirements of chord stability.
[0170] Figure 6 The graph compares the queue lengths under the two methods. It shows that the queue length of the fleet changes significantly over time under the influence of the CTH strategy in the fixed-parameter MPC method. In the initial stage (0-10s), the queue length rapidly increases from approximately 36m to approximately 52m, indicating that the fleet is accelerating. Between 20s and 30s, the queue length remains stable, indicating that the fleet has entered a steady-state driving phase. Subsequently, after 30s, the queue length gradually returns to the initial level, indicating that the fleet has entered a deceleration phase. Compared to the fixed-parameter MPC method, the DDPG-MPC method proposed in this embodiment can dynamically adjust control parameters according to changes in vehicle state. It reduces the desired gap during acceleration, improving fleet operating efficiency, and increases the desired gap during deceleration, improving fleet safety. This demonstrates the effectiveness of the DDPG-MPC method proposed in this embodiment in autonomous driving fleet management, and helps optimize longitudinal control of the fleet in complex dynamic traffic environments.
[0171] Figure 7 and Figure 8 The gap error weighting coefficients of the DDPG-MPC method proposed in this embodiment are respectively. and workshop gap strategy model parameters The variation curves, combined with the acceleration changes of the leading CAV, show that the gap error weight coefficient Q and the gap strategy model parameter c of the CAV change appropriately with the state of the preceding vehicle, ensuring the efficiency and safety of following. Specifically, as the vehicle's acceleration increases, the controller's gap error weight decreases, and the gap strategy parameter also decreases accordingly. During the acceleration phase, as the vehicle's acceleration decreases, the corresponding gap error weight coefficient Q and gap strategy model parameter c increase. During the deceleration phase, as the vehicle's acceleration decreases, the controller's gap error weight decreases while the gap strategy parameter increases; when the vehicle's acceleration increases, the controller's gap error weight increases while the gap strategy parameter decreases. When the vehicle's acceleration stabilizes, the controller parameters remain at a fixed optimal value. The results show that, compared to the traditional fixed-parameter MPC method, the DDPG-MPC method proposed in this embodiment can dynamically adjust the gap error weight and gap strategy parameter according to the real-time changes in vehicle state and road environment, achieving a smoother control response.
[0172] In summary, this invention proposes a vehicle platoon control algorithm (DDPG-MPC) based on an improved gap strategy and a deep deterministic policy gradient algorithm (DDPG), specifically for the cooperative control of vehicle-to-everything (CAV) fleets, aiming to improve fleet safety and following efficiency in dynamic traffic environments. This invention improves the fixed headway (CTH) strategy, comprehensively considering the influence of safety distance and relative speed, and proposes an adaptive distance control strategy. Furthermore, this invention combines a vehicle longitudinal dynamics model and uses an improved gap strategy to construct a vehicle longitudinal model predictive control (MPC) framework. Further, this invention uses the DDPG algorithm to adjust the weight matrix coefficients and gap strategy parameters in the MPC framework in real time, enabling the fleet to dynamically optimize the control strategy according to vehicle state and road environment. Finally, simulation experiments verify the effectiveness of the proposed strategy. Experimental results show that compared with the traditional fixed-parameter MPC method, the proposed DDPG-MPC method can significantly reduce fleet gap error, improve the stability and safety of longitudinal fleet control, and achieve efficient and safe autonomous driving fleet management.
[0173] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A vehicle queuing control method based on an improved gap strategy and a depth deterministic strategy gradient algorithm, characterized in that, Includes the following steps: Establish a vehicle gap strategy that considers the speed difference between the vehicle in front and the vehicle itself, construct an MPC framework for the vehicle queue using the aforementioned vehicle gap strategy, and incorporate the minimum safe distance between vehicles into the constraints of the MPC framework. Training a deep deterministic policy gradient algorithm model includes: Real-time acquisition of the clearance error between the current vehicle and the vehicle in front. Current speed difference with the vehicle in front Current acceleration and current speed And update the values of the corresponding parameters in the state space to obtain the current state of the vehicle at the current moment. The current state Input the Actor network to obtain the best action at the current time. The best action at the current moment The expression is as follows: in, The gap error weighting coefficient of the MPC framework. These are the model parameters for the workshop gap strategy; The best action at the current moment After passing through the Critic network, the reward at the current time step is output. The specific representation of the reward at the current time step is as follows: in, This represents the weighting coefficient of the reward. This is a bonus item for gap error deviation. This is a reward item for relative speed deviation. For the steady-state reward term, the expression is: in, Indicates gap error, Represents the absolute value of the gap error. It represents the absolute value of the relative speed with respect to the vehicle in front; The trained deep deterministic policy gradient algorithm model is used to adjust the parameters in the MPC framework in real time to obtain the optimal parameter combination for each state. Then, the optimal parameter combination is substituted into the MPC framework to calculate the optimal control sequence for each time step. Finally, the vehicles in the vehicle queue are longitudinally controlled according to the specified control quantity in the optimal control sequence of each control cycle.
2. The vehicle queuing control method based on the improved gap strategy and depth deterministic strategy gradient algorithm according to claim 1, characterized in that, The formula for calculating the gap strategy considering the speed difference between the vehicle in front and the vehicle itself is as follows: in, This represents the expected gap between the vehicles in front in the convoy. and Let these represent the speeds of vehicle n and vehicle n-1 in the vehicle queue, respectively. For the desired time interval, Indicates the minimum safe distance. It is the minimum spacing correction term. These are model parameters. This indicates the speed difference between the vehicle in front and the vehicle in front.
3. The vehicle queuing control method based on the improved gap strategy and depth deterministic strategy gradient algorithm according to claim 1, characterized in that, When constructing the MPC framework for vehicle queuing using the aforementioned workshop gap strategy, the specific steps include: For each vehicle in the vehicle platoon, a corresponding single-vehicle longitudinal dynamic model is established, taking into account the speed change of the vehicle in front of each vehicle. Replace the expected gap of the model corresponding to the MPC framework with the workshop gap strategy; Comfort, efficiency, and control input are all taken as optimization control objectives. The objective function and constraints of the vehicle longitudinal dynamics model that takes into account the speed change of the vehicle in front are constructed as the objective function and constraints of the MPC framework.
4. The vehicle queuing control method based on the improved gap strategy and depth deterministic strategy gradient algorithm according to claim 3, characterized in that, The expression for the longitudinal dynamics model of a single vehicle is as follows: Where X is the system state variable, U is the system control input, and d is the system disturbance. For the system input, A, B1, and B2 are all coefficient matrices of the state-space equations. The acceleration of the vehicle in front of it. These represent the difference between the actual distance and the expected distance between vehicle n and the vehicle in front in the vehicle queue, the relative speed between vehicle n and the vehicle in front, the acceleration of vehicle n itself, and the velocity of vehicle n itself. Let be the engine time constant for vehicle n. This represents the desired time interval.
5. The vehicle queuing control method based on the improved gap strategy and depth deterministic strategy gradient algorithm according to claim 3, characterized in that, When simultaneously considering comfort, efficiency, and control input as optimization control objectives, and constructing the objective function and constraints for a vehicle longitudinal dynamics model that takes into account changes in the speed of the vehicle ahead, the steps involved in constructing the objective function include: The longitudinal dynamic model of a single vehicle is discretized using the forward Euler method, as shown in the following expression: In the formula: Indicates that the system is in The state at any given moment The system sampling time, To control the input, To interfere with the input, These are model parameters; The objective function is defined as follows: in, represent At this moment The predicted state value at time 10:
00. represent At this moment The state reference value at any given time. represent Time-based control input, It is a weighted coefficient matrix of spacing error, relative vehicle speed, acceleration, and the derivative of acceleration. These are the weighting coefficients of the control variables. It is the prediction time domain. It controls the time domain.
6. The vehicle queuing control method based on the improved gap strategy and depth deterministic strategy gradient algorithm according to claim 5, characterized in that, When simultaneously considering comfort, efficiency, and control input as optimization control objectives, and constructing the objective function and constraints for a vehicle longitudinal dynamics model that takes into account changes in the speed of the vehicle ahead, the steps for constructing the constraints include: A relaxation factor is introduced to softly constrain different state variables. The specific constraint relationships are as follows: in, This represents a 5×5 identity matrix. Represents the model slack variables. It refers to a diagonal matrix composed of slack variables. Indicates the lower bound of the constraint. Indicates the upper bound of the constraint, Indicates the lower bound of the slack variable, Indicates the upper bound of the slack variable. , , , , These represent the minimum values of clearance error, relative velocity, acceleration, speed, and control input, respectively. , , , , These represent the maximum values of gap error, relative velocity, acceleration, speed, and control input, respectively. It is a relaxation factor. and It is the relaxation factor for the upper and lower limits of system constraints.
7. The vehicle queuing control method based on the improved gap strategy and depth deterministic strategy gradient algorithm according to claim 1, characterized in that, When the minimum safe distance between workshops is added to the constraints of the MPC framework, the maximum parking distance is specifically used as the minimum safe distance between workshops, and the minimum safe distance between workshops is used as the vehicle safety constraint in the constraints. The expression for the maximum parking distance is as follows: in, It is the delay time of the vehicle braking command. It is the vehicle's maximum deceleration. It is the set minimum following distance. This represents the speed of vehicle n. This represents the speed of vehicle n-1.
8. The vehicle queuing control method based on the improved gap strategy and depth deterministic strategy gradient algorithm according to claim 1, characterized in that, When using a trained deep deterministic policy gradient algorithm model to adjust the parameters in the MPC framework in real time, specifically, the vehicle's own state is acquired at each time step, and the vehicle's own state is input into the Actor network to obtain the action that maximizes the reward function calculated by the Critic network under the vehicle's own state. The gap error weight coefficients of the MPC framework included in the action are then applied. and model parameters for workshop gap strategy As the optimal combination of parameters for the vehicle's own condition.
9. The vehicle queuing control method based on the improved gap strategy and depth deterministic strategy gradient algorithm according to claim 1, characterized in that, The specified control quantity in the control sequence of each control cycle specifically refers to the first control quantity in the control sequence.
Citation Information
Patent Citations
Reinforcement learning automatic driving fleet control method based on model predictive control guidance
CN116088530A