Model predictive controller parameter self-calibration method based on deep reinforcement learning
Through deep reinforcement learning and first-order inertial link modeling steering delay characteristics, the complexity and adaptability of MPC controller parameter calibration problems are solved, and high-precision path tracking control of autonomous vehicles under complex operating conditions is realized.
Patent Information
- Application Number
- CN202510782510.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In autonomous vehicles, existing MPC controllers have problems such as high parameter calibration dimensions, complex nonlinear coupling relationships, long calibration periods, poor working conditions adaptability and insufficient online self-adjustment capabilities, resulting in a decrease in control accuracy and robustness.
Using a method based on deep reinforcement learning, the vehicle status data is collected in real time through the sensor network, a multi-dimensional reward function and DQN algorithm are designed, combined with the modeled steering delay characteristics of the first-order inertial link, the self-calibration of the model predicted controller parameters is realized, and the controller parameters are optimized to adapt to complex working conditions.
It improves the control accuracy and stability of autonomous driving vehicles in complex scenarios, shortens the calibration cycle, enhances the adaptability and robustness of the controller, and realizes high-precision path tracking control.
Smart Images

Figure CN120295146A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle intelligent control, and particularly relates to a method for self-calibrating model predictive controller parameters based on deep reinforcement learning. Background Art
[0002] At present, with the rapid development of intelligent driving technology, the path tracking control system, as a core component of the intelligent driving system, its core function is to ensure the lateral motion control accuracy of the autonomous driving vehicle along the preset trajectory during dynamic driving, and at the same time, to take into account good handling stability. Model predictive control (MPC) has been widely used in the field of vehicle lateral control due to its rolling optimization and constraint handling capabilities.
[0003] However, the existing MPC technology faces many key bottlenecks in practical engineering applications. Currently, MPC controllers generally use simplified two-degree-of-freedom vehicle dynamics models or quasi-static tire models for rolling prediction. However, these models do not fully consider a key problem existing in the actual vehicle steering system, that is, there is a phase lag between the controller output steering angle command and the actual actuator response. This defect will directly lead to the deterioration of the lateral motion control performance of the autonomous driving vehicle, and further affect the control accuracy. In addition, the performance of the MPC controller highly depends on the reasonable configuration of parameters such as the cost function weight matrix, and the existing technology generally uses manual calibration methods, which have obvious defects: 1) Calibration dimension problem: The MPC parameter space has a high dimension, and there is a non-linear coupling relationship between parameters. Traditional trial-and-error methods or grid search methods need to traverse more than ten thousand experiments, and the calibration period is long, which seriously restricts the development efficiency of the controller; 2) Poor working condition adaptability: Fixed parameter sets are difficult to adapt to dynamic driving scenarios; 3) Lack of online self-adjustment ability: Existing methods lack a parameter adaptive adjustment and optimization mechanism based on real-time vehicle states, resulting in a significant decrease in the robustness of the controller under time-varying working conditions and being unable to meet the requirements of diverse driving scenarios.
[0004] In recent years, artificial intelligence technology has developed explosively and been applied to the field of intelligent driving vehicles. Breakthroughs in artificial intelligence technologies such as deep reinforcement learning have brought new opportunities for the calibration of model predictive controller parameters. Based on this, the present invention proposes a method for self-calibrating model predictive controller parameters based on deep reinforcement learning. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for self-calibrating model predictive controller parameters based on deep reinforcement learning, aiming to solve the problems proposed in the above background art.
[0006] The purpose of the present invention is achieved through the following technical solutions: A method for self-calibrating model predictive controller parameters based on deep reinforcement learning includes the following steps: Step 1: Vehicle information collection and reference path information acquisition; Arrange a sensor network on the vehicle to collect real - vehicle state data in real - time and simultaneously receive the reference path information transmitted by the upper - layer system; Step 2: Data information pre - processing and state variable calculation; Pre - process the data from the sensors and calculate the state variables required by the model predictive control controller according to the error tracking control framework: lateral deviation, lateral deviation rate of change, heading angle deviation, and heading angle deviation rate of change; Step 3: Design of MPC lateral controller considering steering delay characteristics; Model the delay characteristics of the front - axle and rear - axle steering systems of the vehicle as a first - order inertial link and augment and fuse them as state variables into the error tracking control architecture of the full - by - wire four - wheel steering vehicle model to design the MPC lateral controller; Step 4: Design of parameter self - calibration strategy based on deep reinforcement learning; Take the pre - processed vehicle state data and the front - axle and rear - axle steering angle data as the state space, and take the weight matrix coefficients 、 Output of the model predictive controller as the action space; Use the reinforcement learning algorithm based on the DQN algorithm to design a multi - dimensional reward function, comprehensively evaluate the lateral tracking error, heading tracking error, and control smoothness, and obtain the parameter self - calibration strategy of the model predictive controller based on deep reinforcement learning; Step 5: Send control instructions to the full - by - wire four - wheel steering vehicle and update the vehicle state; Use the output parameters of the trained calibration strategy and apply them to the MPC lateral controller considering steering delay characteristics. The MPC lateral controller calculates the front - axle steering angle and rear - axle steering angle commands and sends them to the actuators of the full - by - wire four - wheel steering vehicle, and updates the state variables based on the vehicle - feedback state data to achieve closed - loop control.
[0007] Furthermore, in the above - mentioned Step 1, the vehicle state data collected by the sensors in real - time includes the heading angle of the current point of the vehicle 、the longitudinal vehicle speed at the vehicle's center of mass 、the lateral vehicle speed at the vehicle's center of mass 、the position information of the vehicle itself in the global coordinate system 、the yaw - rate information of the current point of the vehicle 、the longitudinal acceleration information at the vehicle's center of mass 、the lateral acceleration information at the vehicle's center of mass ; The received reference path information includes the position of the target point at the current moment 、the reference heading angle 、the curvature of the target point .
[0008] Furthermore, in the step 2, the specific steps for calculating the state variables required by the model predictive controller are as follows: Step 2.1: Calculate the lateral deviation and the heading angle deviation of the vehicle based on the position information of the target point and the position information of the vehicle itself. Step 2.2: Derive the rate of change of the heading angle deviation from the obtained heading angle deviation. Step 2.3: The rate of change of the lateral deviation can be obtained by using the longitudinal vehicle speed information, the lateral speed information, and the rate of change of the heading angle deviation.
[0009] Furthermore, the specific steps of the step 3 are as follows: Step 3.1: Select the state vector and the control variable , where is the lateral deviation, is the rate of change of the lateral deviation, is the heading deviation, is the rate of change of the heading deviation, is the transpose operation, is the front axle steering angle of the vehicle, is the rear axle steering angle of the vehicle; Step 3.2: Use a first-order inertia link to characterize the relationship between the actual vehicle steering angle and the target steering angle. The modeling process of the steering delay characteristic is as follows: ; where is the actual vehicle steering angle, is the target vehicle steering angle at the previous moment, is the inertia time constant, is the Laplace variable; Augment the actual front axle steering angle and the actual rear axle steering angle as the delayed states into the state vector of the model predictive controller to obtain the state space equation of the model predictive controller: ; where , are the actual front axle and rear axle steering angles of the vehicle, , are the target front axle and rear axle steering angles of the vehicle at the previous moment, is the state variable of the system at the -th moment, is the state variable of the system at the -th moment, is the sampling time, is the inertial time constant, , are the cornering stiffnesses of the front and rear axles of the vehicle, is the total vehicle mass; is the longitudinal vehicle speed at the vehicle's center of mass; and are the distances from the center of mass to the front and rear axles of the vehicle; is the vehicle's moment of inertia; Step 3.3: Define the objective function in model predictive control as: ; In the formula is the objective function, and are the prediction horizon and the control horizon respectively, is the time step, is the control variable of the system, and represent the predicted output variable and the reference output variable respectively, represents a positive definite matrix of the output weight, represents a positive definite matrix of the control variable weight, is the relaxation factor weight, is the relaxation factor; where, , , in the formula are the weight factors of the lateral error, the change rate of the lateral error, the heading error, and the change rate of the heading error respectively, are the weight factors of the front axle steering angle and the rear axle steering angle of the vehicle respectively; The constraint condition is , in the formula and are the minimum and maximum values of the front axle steering angle and the rear axle steering angle of the vehicle respectively, is the control variable of the system at the th moment.
[0010] Furthermore, the specific steps of Step 4 are as follows: Definition of state space and action space: Action space ; State space includes the lateral deviation , the change rate of the lateral deviation , the heading deviation , the change rate of the heading deviation , the target steering angle of the front axle at the previous moment 、 the target steering angle of the rear axle at the previous moment , the curvature of the reference path , vehicle yaw rate , longitudinal acceleration at the center of mass , lateral acceleration information ; Compound reward function design: Construct a compound reward function architecture centered around vehicle tracking accuracy and control smoothness. Define the immediate reward generated during system state transition as: , where are the lateral error, heading error, the action input at time is the control quantity at time the lateral error control accuracy reward, the heading error control accuracy reward; among them , , where and are the thresholds of lateral deviation and heading deviation respectively; Based on deep reinforcement learning, realize the self-calibration of model predictive controller parameters, specifically including: Step 4.1: Initialize the experience pool M to store the experience data of the full-line-by-wire four-wheel steering vehicle , with a capacity of N; Step 4.2: Randomly initialize the parameters of the main network and the target network ; Step 4.3: Set the number of training rounds of the reinforcement learning algorithm based on the DQN algorithm , the maximum number of training steps in each round T , and conduct training; in the current state, randomly explore the action space with a probability of , and start the exploitation mode with a probability of to select , use the value of as the weight value input of the model predictive controller, the system generates the reward feedback at time , the action decided by the agent acts on the environment, and the output state at time is obtained; Step 4.4: Store the experience data of the system in the experience pool M, randomly sample N historical data from the experience pool M, and calculate for each data using the target network: , where is the value of the target network, is the discount factor, The action value for the target network; Step 4.5: Use the stochastic gradient descent method for optimization to minimize the target loss , and update the main network parameters; Step 4.6: Repeat the training to update the main network parameters, and synchronize the target network with the main network every C steps; Step 4.7: If the termination condition of the maximum number of iterations is reached, the training ends; otherwise, go back to Step 4.3 to continue the training.
[0011] Furthermore, the main network and the target network adopt the same deep neural network structure, with a three-layer network, 100 neurons in each layer, and the linear ReLU as the activation function.
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The present invention characterizes the relationship between the actual vehicle steering angle and the target steering angle through a first-order inertia link, effectively solving the problems of the extended adjustment time and reduced robustness of the control system caused by the response hysteresis of the actuator of the fully-wireless four-wheel steering vehicle, and improving the chassis motion control ability in complex scenarios.
[0013] 2. The present invention combines the deep reinforcement learning technology and innovatively designs a multi-objective reward function, solving the multi-objective coordinated optimization problem of lateral error, heading error, and control smoothness, and realizing the high-precision and stable lateral control of the fully-wireless four-wheel steering vehicle under complex working conditions.
[0014] 3. The present invention realizes the autonomous calibration and optimization of multiple parameters of the model predictive controller through real-time interactive training and closed-loop state update. The generalization performance of reinforcement learning enables the vehicle to adapt to the vast majority of complex working conditions after sufficient training, solving the technical problems of the long calibration period, poor adaptability, and insufficient robustness of the traditional manual calibration method, and improving the path tracking accuracy. Description of the Drawings
[0015] Figure 1 It is the method flow chart of the present invention.
[0016] Figure 2 It is the trajectory tracking effect, vehicle speed, and lateral deviation of the traditional model predictive controller under the 69 km / h double lane change working condition.
[0017] Figure 3 It is the trajectory tracking effect, vehicle speed, and lateral deviation of the model predictive controller based on deep reinforcement learning under the 80 km / h double lane change working condition. Detailed Embodiments
[0018] For a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, the following detailed description of the technical solution of the present invention is provided, but it should not be construed as a limitation on the scope of implementation of the present invention.
[0019] The present invention provides a method for self-calibrating the parameters of a model predictive controller based on deep reinforcement learning. The flowchart is as Figure 1 shown, and the method includes the following steps: Step 1: Vehicle information collection and reference path information acquisition; Arrange a sensor network (accelerometer, displacement sensor, IMU module) on the vehicle to collect vehicle state data in real time, and at the same time receive the reference path information transmitted by the upper-level system; The vehicle state data collected by the sensor in real time includes the heading angle of the current vehicle point , the longitudinal vehicle speed at the vehicle center of mass , the lateral vehicle speed at the vehicle center of mass , the position information of the vehicle itself in the global coordinate system , the yaw rate information of the current vehicle point , the longitudinal acceleration information at the vehicle center of mass , the lateral acceleration information at the vehicle center of mass ; The received reference path information includes the target point position at the current moment , the reference heading angle , the curvature of the target point .
[0020] Step 2: Data information preprocessing and state quantity calculation; Use a preset method to preprocess the data from the sensor (detect possible incorrect data, such as data points outside the reasonable range, and screen, correct, or eliminate outliers) to ensure the accuracy and applicability of the data. Calculate the state quantities required for the model predictive controller according to the error tracking control framework: lateral deviation, lateral deviation change rate, heading angle deviation, heading angle deviation change rate, and form a state vector , where is the lateral deviation, is the lateral deviation change rate, is the heading deviation, is the heading deviation change rate, is the transpose operation; The specific steps for calculating the state quantities required for the model predictive controller are as follows: Step 2.1: According to the position information of the target point and the position information of the vehicle itself , calculate the lateral deviation of the vehicle and the heading angle deviation ; ; Step 2.2: According to the obtained heading angle deviation Derive the change rate of the heading angle deviation ; Step 2.3: Utilize the longitudinal vehicle speed information , lateral speed information and the change rate of the heading angle deviation to obtain the change rate of the lateral deviation .
[0021] ; Since the heading angle deviation between the vehicle and the path reference point is small during high-speed tracking, the above equation can be linearly approximated using small-angle trigonometric functions, resulting in: ; Step 3: Design of the MPC lateral controller considering the steering delay characteristic; Model the delay characteristics of the front and rear axle steering systems of the vehicle as a first-order inertia link and augment and fuse them as state variables into the model error tracking control architecture of the fully-by-wire four-wheel steering vehicle. Based on this, design the MPC lateral controller.
[0022] The specific process of this step is as follows: Step 3.1: Based on the state vector , control variable , where is the front axle steering angle of the vehicle, is the rear axle steering angle of the vehicle.
[0023] Step 3.2: Use a first-order inertia link to characterize the relationship between the actual vehicle steering angle and the target steering angle. The steering delay characteristic modeling process is as follows: ; where is the actual vehicle steering angle, is the target vehicle steering angle at the previous moment, is the inertia time constant, is the Laplace variable.
[0024] Augment the actual front axle steering angle and the actual rear axle steering angle as delay states into the state vector of the model predictive controller to obtain the state space equation of the model predictive controller: ; wherein and are the actual steering angles of the front axle and rear axle of the vehicle, and are the target steering angles of the front axle and rear axle of the vehicle at the previous moment, is the state variable of the system at the th moment, is the state variable of the system at the th moment, is the sampling time, is the inertia time constant, and are the cornering stiffnesses of the front axle and rear axle of the vehicle, is the total vehicle mass; is the longitudinal vehicle speed at the vehicle center of mass; and are the distances from the center of mass to the front axle and rear axle of the vehicle; is the moment of inertia of the vehicle.
[0025] Step 3.3: Define the objective function in the model predictive control as: ; wherein is the objective function, and are the prediction horizon and control horizon respectively, is the time step, is the control variable of the system, and represent the predicted output variable and the reference output variable respectively, represents a positive definite matrix of the output weight, represents a positive definite matrix of the control variable weight, is the slack factor weight, is the slack factor; wherein, , , wherein are the weight factors of the lateral error, the change rate of the lateral error, the heading error, and the change rate of the heading error respectively, are the weight factors of the steering angle of the front axle and the steering angle of the rear axle of the vehicle respectively.
[0026] Convert the MPC problem into a quadratic programming problem, and the constraint condition is , wherein and are the minimum and maximum values of the steering angle of the front axle and the rear axle of the vehicle respectively, is the control variable of the system at the th moment.
[0027] Step 4: Design of parameter self-calibration strategy based on deep reinforcement learning; The preprocessed vehicle state data and the front and rear axle steering angle data are used as the state space, and the positive definite matrix of the output weight of the model predictive controller and the positive definite matrix of the control quantity weight , The output is used as the action space; a reinforcement learning algorithm based on the DQN algorithm is used to design a multi-dimensional reward function to comprehensively evaluate the lateral tracking error, heading tracking error, and control smoothness, and a parameter self-calibration strategy for the model predictive controller based on deep reinforcement learning is obtained.
[0028] The specific process of this step is as follows: Definition of state space and action space: Considering the influence of actual weight coefficients on the controller effect, the action space is selected to reduce the amount of action . The state space includes the information required for the vehicle to complete the tracking path control task: lateral deviation , lateral deviation change rate , heading deviation , heading deviation change rate , the target steering angle of the front axle at the previous moment 、 The target steering angle of the rear axle at the previous moment , reference path curvature , vehicle yaw rate , longitudinal acceleration at the center of mass , lateral acceleration information .
[0029] Design of composite reward function: Considering that the parameter self-calibration method of the model predictive controller based on deep reinforcement learning focuses on vehicle tracking accuracy and control smoothness, a composite reward function architecture is constructed, and the immediate reward generated when the system state transitions is defined as: , where in the formula are the weight coefficients of the lateral error, heading error, the action amount input at time, the lateral error control accuracy reward, and the heading error control accuracy reward respectively, is the control quantity at time, is the lateral error control accuracy reward, is the heading error control accuracy reward; among them , , where in the formula and are the thresholds of the lateral deviation and the heading deviation respectively.
[0030] Implement parameter self-calibration of the model predictive controller based on deep reinforcement learning (deep Q network, i.e., DQN), specifically including: Step 4.1: Initialize the experience pool M to store the experience data of the full-line controlled four-wheel steering vehicle , with a capacity of N; where is the output state at time is the action, is the reward, is the output state at time
[0031] Step 4.2: Randomly initialize the parameters of the main network and the target network parameters. The main network and the target network adopt the same deep neural network structure. Considering the requirements of system complexity and real-time performance, a three-layer network is used. Each layer of network neurons is connected pairwise, and the number of neurons in each layer is 100. The linear ReLU is used as the activation function
[0032] Step 4.3: Set the number of training rounds of the reinforcement learning algorithm based on the DQN algorithm , the maximum number of training steps in each round , and conduct training. At the current state, randomly explore the parameters of the cost function in the model predictive controller with probability , and select with probability to start the exploitation mode , where is the state observation of the system at time , is the action amount of the system at time; Input the value of as the weight value of the model predictive controller, and the system generates the reward feedback at time , and the action decided by the agent acts on the environment to obtain the output state at time .
[0033] Step 4.4: Store the experience data of the system in the experience pool M, randomly sample N historical data from the experience pool M, and calculate for each data using the target network: , where is the value of the target network, is the discount factor, is the action value of the target network
[0034] Step 4.5: Use the stochastic gradient descent method for optimization to minimize the target loss (where is the action value of the main network), and update the parameters of the main network accordingly
[0035] Step 4.6: Repeatedly train to update the main network parameters, and synchronize the target network with the main network every C steps to make the target network more stable than the training network.
[0036] Step 4.7: If the termination condition of the maximum number of iterations is reached, the training ends; otherwise, go back to Step 4.3 to continue training.
[0037] Step 5: Send the control instruction to the full-wire-controlled four-wheel steering vehicle and update the vehicle state.
[0038] Output the parameters using the trained calibration strategy and apply them to the MPC lateral controller considering the steering delay characteristics. The MPC lateral controller calculates the front axle steering angle and rear axle steering angle commands and sends them to the actuators of the full-wire-controlled four-wheel steering vehicle, and updates the state variables based on the vehicle feedback state data to achieve closed-loop control, ensuring the dynamic accuracy of path tracking.
[0039] The following describes the specific implementation of the present invention in detail with specific embodiments.
[0040] Embodiment 1: High-adhesion double lane change test path tracking control condition; In this embodiment, a full-wire-controlled four-wheel steering vehicle platform is used to conduct in-vehicle test comparison and verification on the model predictive controller, and the same trajectory is tracked within the error range of lateral deviation of ±0.2 m. This autonomous vehicle is equipped with front and rear axle wire-controlled steering systems, a high-precision vehicle-mounted integrated navigation and positioning system, and a pair of GNSS measurement antennas, which can provide the vehicle position and vehicle state in real time. During the test, the prototype controller runs the control algorithm in real time, and collects the positioning information and the chassis information of the full-wire-controlled four-wheel steering vehicle through the CAN bus. Figure 2 Shows the trajectory tracking effect, vehicle speed, and lateral deviation of the traditional model predictive controller under the 69 km / h double lane change condition; Figure 3 Shows the trajectory tracking effect, vehicle speed, and lateral deviation of the model predictive controller based on deep reinforcement learning in the present invention under the 80 km / h double lane change condition.
[0041] From Figure 2 and Figure 3It can be seen that both the traditional fixed-parameter model predictive control method (using a traditional model predictive controller) and the model predictive control method based on deep reinforcement learning (using a model predictive controller based on deep reinforcement learning) can complete the path tracking control task under the double lane change condition. The traditional fixed-parameter model predictive control method cannot meet the error range of lateral deviation of ±0.2 m at a vehicle speed of 69 km / h, and the maximum error reaches 0.3 m. Moreover, after traveling 200 m, oscillation occurs near the expected trajectory. However, the model predictive control method based on deep reinforcement learning adopted in the present invention can stably and accurately track the reference path, and still keep the lateral error within the range of 0.2 m at a vehicle speed of 80 km / h. Compared with the traditional fixed-parameter model predictive control method, the maximum vehicle speed of the method of the present invention is increased by 15.94%, and the change of the control quantity is relatively gentle.
[0042] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent.
Claims
1. A method for self-calibrating the parameters of a model predictive controller based on deep reinforcement learning, characterized in that It includes the following steps: Step 1: Vehicle information collection and reference path information acquisition; A sensor network is arranged on the vehicle to collect real vehicle state data in real time, and at the same time, receive the reference path information transmitted by the upper-layer system; Step 2: Data information preprocessing and state quantity calculation; Preprocess the data from the sensors, and calculate the state quantities required by the model predictive control controller according to the error tracking control framework: lateral deviation, lateral deviation change rate, heading angle deviation, and heading angle deviation change rate; Step 3: Design of MPC lateral controller considering steering delay characteristics; Model the delay characteristics of the front axle and rear axle steering systems of the vehicle as a first-order inertial link, and augment and fuse it as a state variable into the error tracking control architecture of the full-wireless four-wheel steering vehicle model to design the MPC lateral controller; Step 4: Design of parameter self-calibration strategy based on deep reinforcement learning; Taking the preprocessed vehicle state data and the front and rear axle steering angle data as the state space, and using the weight matrix coefficients of the model predictive controller , as the output as the action space; using a reinforcement learning algorithm based on the DQN algorithm, designing a multi-dimensional reward function, comprehensively evaluating the lateral tracking error, heading tracking error, and control smoothness, to obtain a self-calibration strategy for the parameters of the model predictive controller based on deep reinforcement learning; Step 5: Send the control command to the full-wireless four-wheel steering vehicle and update the vehicle state; Use the output parameters of the trained calibration strategy and apply them to the MPC lateral controller considering steering delay characteristics. The MPC lateral controller calculates the front axle steering angle and rear axle steering angle commands and sends them to the actuators of the full-wireless four-wheel steering vehicle, and updates the state quantities based on the state data fed back by the vehicle to achieve closed-loop control.
2. The method for self-calibrating the parameters of a model predictive controller based on deep reinforcement learning according to claim 1, characterized in that, In the said step 1, the vehicle state data collected in real time by the sensor includes the heading angle of the current point of the vehicle , the longitudinal vehicle speed at the vehicle's center of mass , the lateral vehicle speed at the vehicle's center of mass , the position information of the vehicle itself in the global coordinate system , the yaw rate information of the current point of the vehicle , the longitudinal acceleration information at the vehicle's center of mass , the lateral acceleration information at the vehicle's center of mass ; The received reference path information includes the position of the target point at the current moment , the reference course angle , and the curvature of the target point .
3. The method for self-calibrating the parameters of a model predictive controller based on deep reinforcement learning according to claim 1, wherein In step 2, the specific steps for calculating the state quantities required by the model predictive controller are as follows: Step 2.1: Calculate the lateral deviation and heading angle deviation of the vehicle according to the position information of the target point and the position information of the vehicle itself; Step 2.2: Derive the change rate of the heading angle deviation from the obtained heading angle deviation; Step 2.3: The change rate of the lateral deviation can be obtained by using the vehicle longitudinal speed information, lateral speed information, and the change rate of the heading angle deviation.
4. The method for self-calibrating the parameters of a model predictive controller based on deep reinforcement learning according to claim 1, wherein The specific steps of step 3 are as follows: Step 3.1: Select the state vector , the control quantity , where is the lateral deviation, is the change rate of lateral deviation, is the heading deviation, is the change rate of heading deviation, is the transpose operation, is the front axle steering angle of the vehicle, is the rear axle steering angle of the vehicle; Step 3.2: Use a first-order inertial link to characterize the relationship between the actual vehicle steering angle and the target steering angle. The steering delay characteristic modeling process is as follows: ; where is the actual steering angle of the vehicle, is the target vehicle steering angle at the previous moment, is the inertia time constant, is the Laplace variable; The actual front axle rotation angle and the actual rear axle rotation angle are augmented as a delay state into the state vector of the model predictive controller to obtain the state space equation of the model predictive controller: ; where , are the actual steering angles of the front and rear axles of the vehicle, , are the target steering angles of the front and rear axles of the vehicle at the previous moment, is the state variable of the system at the th moment, is the state variable of the system at the th moment, is the sampling time, is the inertia time constant, , are the cornering stiffnesses of the front and rear axles of the vehicle, is the total vehicle mass; is the longitudinal vehicle speed at the vehicle's center of mass; and are the distances from the center of mass to the front and rear axles of the vehicle; is the moment of inertia of the vehicle; Step 3.3: Define the objective function in model predictive control as: ; where is the objective function, and are the prediction horizon and control horizon respectively, is the time step, is the control input of the system, and represent the predicted output and reference output respectively, is a positive definite matrix representing the output weight, is a positive definite matrix representing the control input weight, is the slack factor weight, is the slack factor; where, , , where are the weight factors of the lateral error, the change rate of the lateral error, the heading error, and the change rate of the heading error respectively, are the weight factors of the front axle angle and the rear axle angle of the vehicle respectively; The constraint conditions are , where and are the minimum and maximum values of the front axle angle and the rear axle angle of the vehicle respectively, is the control quantity of the system at the th moment.
5. The method for self-calibrating model predictive controller parameters based on deep reinforcement learning according to claim 1, wherein The specific steps of step 4 are as follows: Definition of State Space and Action Space: Action Space ; State Space includes lateral deviation , rate of change of lateral deviation , heading deviation , rate of change of heading deviation , target angle of the front axle at the previous moment 、 target angle of the rear axle at the previous moment , reference path curvature , vehicle yaw rate , longitudinal acceleration at the center of mass , lateral acceleration information ; Composite Reward Function Design: Construct a composite reward function architecture centered around vehicle tracking accuracy and control smoothness. Define the immediate reward generated during system state transition as: , where are the lateral error, heading error, the weight coefficients of the action quantity input at time , the lateral error control accuracy reward, and the heading error control accuracy reward respectively. is the control quantity at time is the lateral error control accuracy reward, is the heading error control accuracy reward; among which , , where and are the thresholds of the lateral deviation and the heading deviation respectively; Realize the parameter self-calibration of the model predictive controller based on deep reinforcement learning, specifically including: Step 4.1: Initialize the experience pool M to store the experience data of the full X-by-wire four-wheel steering vehicle , with a capacity of N; Step 4.2: Randomly initialize the parameters of the main network and the target network ; Step 4.3: Set the number of training rounds for the reinforcement learning algorithm based on the DQN algorithm , the maximum number of training steps in each round T , for training; in the current state, The probability of randomly exploring the action space ,by The probability of opening the exploit mode selection ,Will The value of is used as the weight value input of the model predictive controller, and the system generates Reward feedback at all times , the action decided by the agent acts on the environment, and we get Output status at the moment ; Step 4.4: Store the empirical data of the system in the empirical pool M, and randomly sample N historical data from the empirical pool M , and calculate for each data using the target network: , where is the value of the target network, is the discount factor, is the action value of the target network; Step 4.5: Use the stochastic gradient descent method for optimization to minimize the objective loss , and update the main network parameters; Step 4.6: Repeat training to update the main network parameters, and synchronize the target network with the main network every C steps; Step 4.7: If the termination condition of the maximum number of iterations is reached, the training ends; otherwise, go back to step 4.3 to continue training.
6. The method for self-calibrating model predictive controller parameters based on deep reinforcement learning according to claim 5, wherein The main network and the target network adopt the same deep neural network structure, with a three-layer network, 100 neurons in each layer, and using linear ReLU as the activation function.
Citation Information
Patent Citations
Model predictive control trajectory tracking control system and method based on reinforcement learning
CN114967676A
Man-vehicle cooperative steering control method based on reinforcement learning corner weight distribution
CN115062539A
Automatic driving vehicle transverse control method based on DRL-MPC
CN117360544A
Vehicle trajectory tracking and control method based on reinforcement learning MPC
CN118605160A
Four-wheel independent steering and driving vehicle path tracking control method based on compound control framework
CN119535962A
Cited By
Intelligent automobile longitudinal control method based on dynamic feature learning optimization
CN120595610A
A longitudinal control method for intelligent vehicles based on dynamic feature learning optimization
CN120595610B
Trajectory tracking control method and device
CN121165713A
Hydrological forecasting model multi-flood-session multi-target parameter calibration method based on deep reinforcement learning
CN121598785A