A braking control method based on the comfort of autonomous vehicles
Through the TD3 deep reinforcement learning algorithm, the Q matrix of the model's predictive control is trained in stages, the weight coefficient is dynamically adjusted, and the braking force distribution is optimized, which solves the problem of large changes in pitch angles of autonomous vehicles at different braking stages, and improves comfort and real-time adaptability.
Patent Information
- Application Number
- CN202411636033.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-15
AI Technical Summary
In the braking control of autonomous driving vehicles, it is difficult to effectively reduce the change in the vehicle pitch angle during braking under a dynamic environment, resulting in poor passenger comfort, especially the pitch angles at different braking stages vary greatly.
The TD3 deep reinforcement learning algorithm is used to train the Q matrix of the model predictive control in stages, dynamically adjust the Q weight coefficient, combine the model prediction controller, optimize the braking force distribution, and generate the most comfortable control sequence.
It improves the comfort of autonomous vehicles at different braking stages, reduces jitter in pitch angle changes, and improves the comfort of passenger experience and the real-time adaptability of the system.
Smart Images

Figure CN119568090B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and discloses a braking control method based on the comfort of autonomous vehicles. Background Art
[0002] In recent years, with the development of autonomous driving technology, comfort has become one of the key factors in enhancing the passenger experience. Among them, braking control plays a crucial role in affecting the riding experience. Braking control can be divided into emergency braking and normal braking. Currently, comfortable braking for normal conditions is mainly divided into three categories. One is to control the vehicle suspension, and different total braking forces and front-rear axle distribution ratios are allocated according to different vehicle speeds to ensure that the vehicle pitch angle always remains within a comfortable range during braking. The second is to optimize the braking deceleration with the help of empirical formulas under normal conditions, thereby improving the comfort of the vehicle. The third is to jointly control the vehicle suspension and braking system, but the cost is relatively high. Although many scholars have conducted a lot of research, there are still deficiencies:
[0003] 1. Many scholars use braking force distribution for comfortable braking and derive the relationship between deceleration and braking force. For example, observing the change of the minimum pitch angle and using feedforward control for braking force distribution, but this method can only ensure the control effect under quasi-static conditions.
[0004] 2. There are also many scholars who focus on suspension control and use two model predictive controllers to control the distribution of the total braking force and the front-rear axle braking force. However, it is impossible to give full play to the advantage of decoupling the braking force of the control system, and the complexity of the two model predictive controllers is relatively high. Therefore, most scholars set the weight coefficient of the model predictive controller as a fixed value, which greatly reduces the braking effect and affects the passenger experience.
[0005] In summary, at present, many comfortable braking methods use model predictive controllers to control vehicles. However, during braking, the pitch angles in the early, middle, and late stages of braking are quite different, and the Q weight coefficient of the model predictive controller is set as a fixed value, and the surrounding environment and vehicle speed change at any time. Therefore, it is difficult to achieve the comfortable braking effect. Summary of the Invention
[0006] The present invention discloses a braking control method based on the comfort of autonomous vehicles, and the braking control method based on the comfort of autonomous vehicles includes the following steps:
[0007] Step 1, collect the speed and distance information of the obstacle ahead, analyze the collected road information and vehicle information, and then comprehensively judge whether it is an emergency condition in combination with the current speed of the vehicle itself, the speed of the vehicle ahead, and the distance information from the obstacle.
[0008] Step 2: Learn the Q matrix in model predictive control through the TD3 deep reinforcement learning algorithm;
[0009] Step 3: The model predictive control model estimates the current braking stage based on the information input data, and calls the Q matrix with the best comfort for the current stage from the policy library for optimal solution;
[0010] Step 4: Call the trained Q matrix for optimal solution to generate the most comfortable control sequence;
[0011] Step 5: Transmit the control sequence to the vehicle control module for signal conversion to generate braking torque, so that the vehicle pitch angle is within a comfortable range throughout the process.
[0012] Furthermore, in Step 2, the following steps are also included:
[0013] Step 21: Build a TD3 training module;
[0014] Step 22: Continuously adjust the hyperparameters during the training process, and observe the control effects of acceleration, pitch angle, and pitch angular velocity to achieve the most matching reward weight coefficients;
[0015] Step 23: Use the TD3 deep reinforcement learning algorithm to train the Q weight coefficients of acceleration, pitch angle, and pitch angular velocity in the middle and late stages of braking, evaluate the reward evaluation effects of different weight coefficients, and put the trained Q weight coefficients back into the policy library.
[0016] Furthermore, in Step 21, in order to ensure that the vehicle can brake smoothly, the reward function of the TD3 training module is set as a function related to vehicle comfort:
[0017] r a =-γ1·|a - a T |
[0018] r θ =-γ2·|θ - θ T |
[0019]
[0020] They are the rewards for acceleration a, pitch angle θ, and pitch angular velocity respectively. γ1 is the reward weight coefficient for acceleration, γ2 is the reward weight coefficient for pitch angle, γ3 is the reward weight coefficient for pitch angular velocity, a T represents the target acceleration, θ T represents the target pitch angle, represents the target pitch angular velocity;
[0021] r总 The expression is the sum of three terms:
[0022]
[0023] where w1, w2, and w3 are the weight coefficients for the acceleration reward stage, pitch angle reward stage, and pitch rate stage, respectively;
[0024] Training is carried out in stages. The braking stage is mainly divided into the early - mid stage and the late stage. In the early - mid stage of vehicle braking, q in MPC a has a relatively large weight in this stage, followed by the pitch angle, and finally the pitch rate. The reward expression for the early - mid stage is:
[0025]
[0026] w 11 represents the acceleration weight coefficient in the early - mid stage of braking, w 21 represents the pitch angle weight coefficient in the early - mid stage of braking, w 31 represents the pitch rate weight coefficient in the early - mid stage of braking.
[0027] In the late stage of vehicle braking, should have a relatively large weight in this stage. So in the late stage of braking, the pitch rate is the main target, followed by the pitch angle, and finally the acceleration. The reward expression for the late stage is:
[0028]
[0029] w 21 represents the acceleration weight coefficient in the late stage of braking, w 22 represents the pitch angle weight coefficient in the late stage of braking, w 32 represents the pitch rate weight coefficient in the late stage of braking.
[0030] Furthermore, in step 21, the target network is updated in a soft - update manner. The soft - update expression is:
[0031] θ target ←τ·θ+(1 - τ)·θ target
[0032] τ is the soft - update coefficient, θ represents the current network parameters, and θ target represents the target network parameters;
[0033] The expression of the objective function is as follows:
[0034]
[0035] For the objective function J with respect to the policy network parameters θμ Gradient of is the gradient of the Q function with respect to the action, is the gradient of the policy network with respect to the action, i is the sample number, and N is the total number of samples;
[0036] The parameters θ of the Q-value function are optimized by minimizing the mean square error between the predicted value and the target value Q , so the expression of the loss function is:
[0037]
[0038] Q(s i , a i |θ Q ) represents the predicted value, and y i represents the target value.
[0039] Furthermore, in step 3, the model predictive controller continuously obtains the Q-matrix parameters from the TD3 training module, matches different weight coefficients in different braking phases, generates corresponding actions and states, and puts them back into the experience pool of the TD3 training module. This process repeats continuously to update the Q-matrix library.
[0040] Furthermore, in step 4, according to the model predictive control principle, the comfort optimization function can be obtained as:
[0041]
[0042] Where J represents the comfort optimization function, m represents the sampling step, m = 1:1:N P , N P represents the prediction step, y(k + m|t) represents the predicted value of the control output, y ref (k + m|t) represents the reference value of the control output, N C represents the control step, Δu(k + m|t) represents the control input increment, Q is the system output quantity weight system matrix, R is the control quantity weight coefficient matrix, ρ is the weight coefficient, and ε is the relaxation factor; q a is the acceleration weight coefficient, q θ is the vehicle pitch angle weight coefficient, is the vehicle pitch angular velocity weight coefficient;
[0043] The minimum value min J corresponding to the comfort objective function solver is found, and to prevent sudden changes in the control quantity, the following constraint conditions are added:
[0044] u min (k + m) ≤ u(k + m|k) ≤ u max (k + m)
[0045] Δu min(k + m) ≤ Δu(k + m|k) ≤ Δu max (k + m)
[0046] u min represents the minimum value of the control input, u max represents the maximum value of the control input, Δu max represents the maximum value of the control increment, Δu mim represents the minimum value of the control increment;
[0047] Referring to the desired vehicle speed, based on the output parameter Y of the acceleration prediction model, a series of optimal acceleration increments ΔU(t) under the system constraint conditions can be solved. Taking the first acceleration increment Δu(k|t) in this series and adding the acceleration control amount at the previous moment, the current acceleration control amount u can be obtained. The beneficial effects achieved by the present invention are:
[0048] Improve braking comfort: Use the TD3 deep reinforcement learning algorithm to train the Q matrix of the model predictive control in stages. The adaptive MPC adaptively adjusts the Q weight coefficient to adapt to different braking stages for solution, reducing the jitter caused by the change of the pitch angle, thereby improving comfort.
[0049] Real-time adaptability: The dynamically adjusted MPC can better adapt to the continuously changing driving environment, thereby improving the flexibility and reaction speed of the autonomous driving system. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a schematic flow chart of a braking control method based on the comfort of an autonomous vehicle proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0051] The present invention will be further described below in conjunction with specific embodiments, and the advantages and features of the present invention will become clearer as the description progresses. However, these embodiments are exemplary only and do not constitute any limitation to the scope of the present invention. Those skilled in the art should understand that the details and forms of the technical solutions of the present invention can be modified or replaced without departing from the spirit and scope of the present invention, but these modifications and replacements all fall within the protection scope of the present invention.
[0052] As Figure 1 shown, a braking control method based on the comfort of an autonomous vehicle according to the present invention includes the following steps:
[0053] Step 1, first use sensors such as on-vehicle lidar and binocular cameras to collect the speed and distance information of the obstacles ahead, analyze the collected road information and vehicle information, and then comprehensively judge whether it is in an emergency working condition or a non-emergency working condition in combination with the current speed of the vehicle itself, the speed of the vehicle ahead, and the distance information from the obstacles.
[0054] Specifically, it further includes the following steps:
[0055] Step 11, scene recognition;
[0056] Relevant measurement data are obtained through an in-vehicle lidar, a binocular camera, and other sensors, and the data are input into the scene recognition module. The scene recognition module can estimate based on these measurement data to obtain the speed of the autonomous vehicle, the speed of the vehicle ahead, the distance from the vehicle ahead to the autonomous vehicle, the pitch angle of the vehicle, and so on.
[0057] Step 12, make a judgment based on the road information and vehicle information obtained by the data acquisition module, so as to perform working condition switching;
[0058] The main working process is as follows: If the vehicle speed of the vehicle ahead obtained from the scene recognition module decreases sharply or other sudden scenarios occur, combined with the excessive displacement distance of the brake pedal, the vehicle switches to the emergency braking working condition. At this time, safety is greater than comfort, and emergency braking should be performed. On the contrary, if the data acquisition module obtains that the vehicle ahead has stopped or is in front of a red light at an intersection, combined with the fact that the displacement distance of the brake pedal is not very large, the vehicle is in the normal working condition braking, and then subsequent comfortable braking is performed.
[0059] Step 2, learn the Q matrix in model predictive control through the TD3 deep reinforcement learning algorithm.
[0060] The TD3 training module is for the TD3 deep reinforcement learning algorithm to learn the Q matrix parameters in model predictive control, and generate corresponding predictive model outputs according to the road information and speed information for subsequent optimal control calculations.
[0061] The TD3 training module sets a suitable reward function with the goals of minimizing acceleration, minimizing pitch angle, and minimizing pitch angular velocity to evaluate the performance of different weight coefficients of the Q matrix in the model predictive control system. During the training process, the TD3 algorithm interacts with the environment to collect experience and optimize. First, the policy network is used to select an action. Taking the current state as the input, the currently selected action, that is, the corresponding Q matrix, is applied to the model predictive controller to obtain the next state and reward. At the same time, the experience is stored in the experience replay buffer, and then the policy network and the value function network are updated by minimizing the loss function defined by the TD3 algorithm.
[0062] Specifically, it further includes the following steps:
[0063] Step 21, construct the TD3 training module
[0064] The observation state of the reinforcement learning model for outputting the predictive model is the longitudinal speed, longitudinal acceleration, and vehicle pitch angle, and the predictive model output can be represented by a sequence:
[0065] [v(t), a(t)]......[v(t + P - 1), a(t + P - 1)]
[0066] v(t) represents velocity, a(t) represents acceleration, and P represents the prediction time step.
[0067] TD3 is an actor - critic - based algorithm that includes a policy network (Actor) and two value networks (Critic). The input of the policy network (Actor) is the state observed by the agent, and the output is the action u(s|θ μ )
[0068] where θ μ is the parameter of the policy network. The purpose of the policy network is to maximize the Q - value output by the Critic network. The value network (Critic) evaluates the Q - value Q(s, a|θ Q ) of the state - action pair (s, a), where θ Q is the parameter of the Critic network. Two Critic networks are used to calculate the Q - value of the action, and the minimum value of them is taken as the target value to reduce over - estimation. The target Q - value satisfies the following expression:
[0069]
[0070] where r is the immediate reward, μtarget is the target policy network, γ is the discount factor, and the value of the discount factor is between 0 and 1. The larger the value, the higher the importance of future rewards. In the actual braking control process, to ensure that the vehicle can brake smoothly, the reward function of the deep reinforcement learning model is set as a function related to vehicle comfort:
[0071] r a = -γ1·|a - a T |
[0072] r θ = -γ2·|θ - θ T |
[0073]
[0074] are the rewards for acceleration a, pitch angle θ, and pitch angular velocity respectively. γ1, γ2, γ3 are their respective reward weight coefficients. The smaller the gap between acceleration, pitch angle, and pitch angular velocity and the target value, the higher the reward value. The expression of r 总 is the sum of the three:
[0075]
[0076] Where w1, w2, and w3 are the weight coefficients for the acceleration reward stage, pitch angle reward stage, and pitch angular velocity stage respectively. Training is carried out in stages here. The braking stage is mainly divided into the middle and early stages and the late stage. In the middle and early stages of vehicle braking, the main focus is on the safety distance of the vehicle, so acceleration is the main target. In MPC, q a should account for a relatively large weight in this stage, followed by the pitch angle, and finally the pitch angular velocity. The reward expression for the middle and early stages is:
[0077]
[0078] In the late stage of vehicle braking, the vehicle speed is relatively slow. When the vehicle speed drops to 0, the vehicle will experience "jitter". In MPC should account for a relatively large weight in this stage. Therefore, in the late stage of braking, the pitch angular velocity is the main target, followed by the pitch angle, and finally the acceleration. The reward expression for the late stage is
[0079]
[0080] Training in this phased manner can better match the required weight coefficients.
[0081] The target network update method is soft update, and the soft update expression is:
[0082] θ target ←τ·θ+(1 - τ)·θ target
[0083] τ is the soft update coefficient. The expression of the target function is as follows:
[0084]
[0085] is the gradient of the target function J with respect to the policy network parameter θ μ . For each sample i, the gradient of the Q function with respect to the action and the gradient of the policy network with respect to the action are summed, and then averaged to obtain the final gradient estimate. Then, the parameters θ Q of the Q value function are optimized by minimizing the mean square error between the predicted value and the target value. Therefore, the expression of the loss function is:
[0086]
[0087] Step 22: Continuously adjust the hyperparameters during training, and observe the control effects of acceleration, pitch angle, and pitch angular velocity to achieve the most matching γ1, γ2, and γ3;
[0088] Step 23, in addition, the acceleration, pitch angle, and pitch angular velocity Q weight coefficients are different for the middle and late stages of braking. Therefore, the deep reinforcement learning algorithm TD3 is used to train the three weight coefficients, evaluate the reward evaluation effects of different weight coefficients, and put the trained Q weight coefficients back into the policy library.
[0089] Step 3, the model predictive control model estimates the current braking stage (middle and late stages) based on the information input data, and calls the Q matrix with the best comfort for the current stage from the policy library for optimization and solution.
[0090] The MPC actual control module is based on the vehicle dynamics model. By inputting the vehicle state and environmental information in real time, it uses model predictive control to calculate the optimal control strategy. And the MPC actual control module can call the Q weight coefficient that is currently most suitable for MPC from the TD3 policy library to facilitate the solution of the MPC optimal value, and finally generate a control signal with strong comfort.
[0091] The real-time adjustment of the weight coefficient is that the TD3 training module dynamically changes the Q weight parameter in the comfort objective function during the MPC control process, so that the optimal control signal sequence output by the MPC is further optimized. It effectively guarantees the comfort of the vehicle. The following are several key aspects of the real-time adjustment of the weight coefficient:
[0092] For a complete braking control, the braking process is divided into the early stage, middle stage, and late stage. The jitter of the vehicle in the middle and early stages of braking is mainly caused by the change in acceleration. For the late stage of braking, the dizziness of passengers mainly comes from the vehicle jitter caused by the sudden change in the pitch angle of the vehicle, which leads to a decline in the passenger riding experience.
[0093] Therefore, in different braking stages, the focus of the model predictive control should be different. This braking method combines the TD3 algorithm with the model predictive control for joint training. In the middle and early stages of braking, set a larger acceleration weight coefficient q a and pitch angle weight coefficient q θ , because in the middle and early stages of braking, the braking distance accounts for a relatively large proportion of the entire braking process and it is necessary to strictly follow the expectations. Therefore, the acceleration and pitch angle are the main optimization objectives, while the pitch angular velocity is the secondary optimization objective. In the late stage of braking, the vehicle speed is slow, and when the vehicle slides to a stop, there will be a sudden change in the pitch angle and the body will shake. So in the late stage of braking, the pitch angular velocity is the main optimization objective, the secondary optimization objective is the pitch angle, and the secondary secondary optimization objective is the acceleration. The TD3 training module should adjust and set a larger pitch angular velocity weight coefficient
[0094] The model predictive controller continuously obtains the Q matrix parameters from the TD3 training module, matches different weight coefficients in different braking stages, generates corresponding actions and states, and then puts them back into the experience pool of the TD3 training module. This process repeats continuously to update the Q matrix library. Eventually, the model predictive controller can obtain the Q matrix parameters after training in different braking stages, and then generate a smooth braking signal.
[0095] Step 4: Call the trained Q matrix for optimization and solution to generate the most comfortable control sequence;
[0096] The optimal value solution is obtained by optimizing and calculating the vehicle state and its change amount to ensure that the optimal control sequence can be output in different braking stages, thereby improving the comfort of passengers. The following is the specific principle of the optimal value solution:
[0097] Construction of the comfort objective function:
[0098] In the model predictive control model, the vehicle acceleration and its own pitch angle are mainly predicted. The state variables include acceleration, pitch angle, and pitch angular velocity.
[0099] The optimizer can solve for the optimal acceleration under the constraint conditions according to the desired vehicle speed and the comfort optimization function, and transmit the control signal to the controller.
[0100] According to the model predictive control principle, the comfort optimization function can be obtained as:
[0101]
[0102] In the formula, Q is the system output quantity weight system matrix, R is the control quantity weight coefficient matrix, ρ is the weight coefficient, and ε is the relaxation factor. q a 、q θ 、 respectively correspond to the acceleration weight coefficient, the vehicle pitch angle weight coefficient, and the vehicle pitch angular velocity weight coefficient.
[0103] To obtain the optimal control sequence of the MPC controller, the minimum value min J of the comfort objective function solver is calculated, and to prevent the sudden change of the control quantity, the following constraint conditions are added:
[0104] u min (k + m) ≤ u(k + m|k) ≤ u max (k + m)
[0105] Δu min (k + m) ≤ Δu(k + m|k) ≤ Δu max (k + m)
[0106] Referring to the expected vehicle speed, a series of optimal acceleration increments ΔU(t) under the system constraint conditions can be solved according to the output parameter Y of the acceleration prediction model. Taking the first acceleration increment Δu(k|t) in this series and adding the acceleration control amount at the previous moment, the current acceleration control amount u can be obtained.
[0107] Step 5: Transmit the control sequence to the vehicle control module for signal conversion to generate braking torque and perform operations, so that the vehicle pitch angle is within a comfortable range throughout the process.
[0108] After receiving the control signal, the vehicle control module performs corresponding braking actions. If the data acquisition module determines that the vehicle is in an emergency, the vehicle control module will perform emergency braking. If the vehicle control module receives the braking instruction from the MPC optimization solution module, it will generate a corresponding braking control amount for comfortable braking. Then, the vehicle and surrounding data are obtained in real time through the vehicle-mounted sensor kit and transmitted to the scene recognition module, and this process repeats continuously to achieve comfortable braking.
[0109] The present invention is not limited to the above specific embodiments. Those of ordinary skill in the art can implement the present invention in various other specific embodiments according to the embodiments and the disclosed content of the drawings. Therefore, any design that adopts the design structure and concept of the present invention and makes some simple transformations or modifications falls within the protection scope of the present invention.
Claims
1. A braking control method based on the comfort of an autonomous vehicle, characterized in that, The braking control method based on the comfort of autonomous vehicles includes the following steps: Step 1: Collect the speed and distance information of the obstacle ahead, analyze the collected road information and vehicle information, and comprehensively judge whether it is an emergency condition by combining the current speed of the vehicle itself, the speed of the vehicle ahead, and the distance information from the obstacle. Step 2: Use the TD3 deep reinforcement learning algorithm to learn the Q matrix in the model predictive control. Step 3: The model predictive control model estimates the current braking stage based on the information input data, and calls the Q matrix with the best comfort for the current stage from the policy library for optimization and solution. Step 4: Call the trained Q matrix for optimization and solution to generate the most comfortable control sequence. Step 5: Transmit the control sequence to the vehicle control module for signal conversion to generate a braking torque, so that the vehicle pitch angle is within a comfortable range throughout the process. In Step 2, the following steps are also included: Step 21: Construct a TD3 training module. Step 22: Continuously adjust the hyperparameters during the training process, and observe the control effects of acceleration, pitch angle, and pitch angular velocity to achieve the most matching reward weight coefficients. Step 23: Use the TD3 deep reinforcement learning algorithm to train the Q weight coefficients of acceleration, pitch angle, and pitch angular velocity in the middle and late stages of braking, evaluate the reward evaluation effects of different weight coefficients, and put the trained Q weight coefficients back into the policy library. In Step 21, in order to ensure that the vehicle can brake smoothly, the reward function of the TD3 training module is set as a function related to vehicle comfort: r a = -γ1·|a - a T | r θ = -γ2·|θ - θ T | are the acceleration a, the pitch angle θ, and the pitch angular velocity rewards, where γ1 is the reward weight coefficient for acceleration, γ2 is the reward weight coefficient for pitch angle, γ3 is the reward weight coefficient for pitch angular velocity, a T represents the target acceleration, θ T represents the target pitch angle, represents the target pitch angular velocity; r 总 The expression of is the sum of three terms: where w1, w2, and w3 are the weight coefficients of the acceleration reward stage, the pitch angle reward stage, and the pitch angular velocity stage, respectively. Training is carried out in stages. The braking stage is mainly divided into the first half and the second half. In the first half of vehicle braking, q in MPC a has a relatively large weight in this stage, followed by the pitch angle, and finally the pitch rate. The reward expression in the first half is as follows: w 11 represents the acceleration weight coefficient in the middle stage before braking, w 21 represents the pitch angle weight coefficient in the middle stage before braking, w 31 represents the pitch angular velocity weight coefficient in the middle stage before braking; In the later stage of vehicle braking, should account for a relatively large weight in this stage. Therefore, in the later stage of braking, the pitch angular velocity is the main target, the pitch angle is the second, and finally the acceleration. The reward expression in the later stage is: w 21 represents the acceleration weight coefficient in the late braking stage, w 22 represents the pitch angle weight coefficient in the late braking stage, w 32 represents the pitch angular velocity weight coefficient in the late braking stage.
2. The braking control method based on the comfort of an autonomous vehicle according to claim 1, wherein In Step 21, the target network update method is soft update, and the soft update expression is: θ target ← τ·θ+(1 - τ)·θ target τ is the soft update coefficient, θ represents the current network parameters, and θ target represents the target network parameters; The expression of the objective function is as follows: For the gradient of the objective function J with respect to the parameters θ of the policy network μ , is the gradient of the Q function with respect to the action is the gradient of the policy network with respect to the action, i is the sample number, and N is the total number of samples; Optimize the parameter θ of the Q-value function by minimizing the mean squared error between the predicted value and the target value Q , so the expression of the loss function is: Q(s i ,a i |θ Q ) represents the predicted value, and y i represents the target value.
3. The braking control method based on the comfort of an autonomous vehicle according to claim 1, wherein In Step 3, the model predictive controller continuously obtains the Q matrix parameters from the TD3 training module, matches different weight coefficients in different braking stages, generates corresponding actions and states, and then puts them back into the experience pool of the TD3 training module. This process repeats continuously to update the Q matrix library.
4. The braking control method based on the comfort of an autonomous vehicle according to claim 1, wherein In Step 4, according to the model predictive control principle, the comfort optimization function can be obtained as: Where J represents the comfort optimization function, m represents the sampling step, m = 1:1:N P , N P represents the prediction step, y(k+m|t) represents the predicted value of the control output, y ref (k+m|t) represents the reference value of the control output, N C represents the control step, Δu(k+m|t) represents the control input increment, Q is the system output quantity weight system matrix, R is the control quantity weight coefficient matrix, ρ is the weight coefficient, ε is the relaxation factor; q a is the acceleration weight coefficient, q θ is the vehicle pitch angle weight coefficient, is the vehicle pitch angular velocity weight coefficient; Solve the minimum value min J of the comfort objective function solver, and prevent the situation of sudden change of the control quantity. Therefore, the following constraint conditions are added: u min (k + m) ≤ u(k + m|k) ≤ u max (k + m) Δu min (k + m) ≤ Δu(k + m|k) ≤ Δu max (k + m) u min represents the minimum value of the control input, u max represents the maximum value of the control input, Δu max represents the maximum value of the control increment, Δu min represents the minimum value of the control increment; Referring to the expected vehicle speed, according to the output parameter Y of the acceleration prediction model, a series of optimal acceleration increments ΔU(t) under the system constraint conditions can be solved. Take the first acceleration increment Δu(k|t) of this series, and add the acceleration control quantity at the previous moment to obtain the current acceleration control quantity u.
Citation Information
Patent Citations
Electric vehicle economical self-adaptive cruise control method and system based on reinforcement learning
CN114771520A
Train emergency parking method based on small sample data enhanced ensemble learning
CN114852129A