Dynamic drift method and device based on deep reinforcement learning mode, and medium

By employing a dynamic drift method based on deep reinforcement learning, combined with various vehicle models and controllers, the drift control problem of autonomous driving systems under extreme conditions was solved, achieving efficient and stable improvement in vehicle cornering performance.

CN121411166APending Publication Date: 2026-01-27FOSHAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511810136.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Autonomous driving systems struggle to achieve effective drift control under extreme conditions. Traditional methods suffer from poor adaptability, high computational costs, and difficulty in guaranteeing stability.

Method used

A dynamic drift method based on deep reinforcement learning is adopted, which combines a three-degree-of-freedom vehicle single-track dynamics model, a tire model, a control guidance trajectory tracking dynamics model, and a deep reinforcement learning agent to construct an MPC drift controller, a PID longitudinal controller, and an LQR lateral controller to achieve targeted processing of drift and tracking.

Benefits of technology

It improves the vehicle's cornering performance under extreme conditions, enhances computational efficiency and stability, and is highly adaptable, enabling it to switch drift strategies at the optimal time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121411166A_ABST
    Figure CN121411166A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic drift method and device based on a deep reinforcement learning mode and a medium, and relates to the technical field of data processing, and the method comprises the steps: constructing a three-degree-of-freedom vehicle monorail dynamics model, a tire model and a control guide trajectory tracking dynamics model; constructing a front wheel tire inverse model according to the tire model; a rear wheel transverse force prediction model based on a Koopman operator is constructed; calculating a drift balance point in real time according to the three-degree-of-freedom vehicle monorail dynamics model; constructing an MPC drift controller according to the control guide trajectory tracking dynamics model, a front wheel tire inverse model, a rear wheel transverse force prediction model and a drift balance point; constructing an LQR transverse controller according to the control guide trajectory tracking dynamics model; a PID longitudinal controller is constructed; and constructing a DRL intelligent agent and carrying out mode switching in combination with the MPC drift controller, the PID longitudinal controller and the LQR transverse controller. By the adoption of the method, the comprehensive performance of the vehicle under the limiting working condition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a dynamic drift method, device and medium based on deep reinforcement learning. Background Technology

[0002] Typically, autonomous driving systems are designed for standard operating conditions, keeping the vehicle within a stable linear range to ensure high safety. However, in extreme conditions such as emergency cornering on slippery surfaces, vehicle dynamics exhibit high nonlinearity. Traditional control strategies, due to their conservative performance, struggle to fully utilize the vehicle's physical limits, potentially leading to accidents caused by insufficient avoidance.

[0003] Drifting, as an advanced driving technique that utilizes tire force saturation to achieve rapid cornering, can significantly expand the handling limits of a vehicle. Applying drifting technology to autonomous driving would greatly improve the active safety and maneuverability of vehicles under extreme conditions. However, current autonomous driving systems still have the following shortcomings in dynamic drift control: (1) Traditional methods generally rely on heuristic rules with preset fixed thresholds for switching between regular turning and drift control. This logic, which relies on expert experience, is poorly adaptable to complex and ever-changing working conditions, affecting cornering efficiency and safety. (2) In order to accurately describe the strong nonlinear dynamics of drift, although the nonlinear model predictive control (NMPC) is effective, it will form a non-convex optimization problem with extremely high computational cost, and its solution time is difficult to meet the real-time requirements of drift. (3) Directly using end-to-end deep learning to control vehicle drift makes it difficult to theoretically guarantee the stability and safety of the controller due to its black box characteristics.

[0004] Therefore, this invention proposes a dynamic drifting method to solve the above problems and improve the overall performance of vehicles under extreme conditions. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a dynamic drift method, device and medium based on deep reinforcement learning, which can improve the overall performance of vehicles under extreme conditions.

[0006] To address the aforementioned technical problems, this invention provides a dynamic drift method based on deep reinforcement learning, comprising: constructing a three-degree-of-freedom vehicle single-track dynamics model, a tire model, and a control-guided trajectory tracking dynamics model, wherein the control-guided trajectory tracking dynamics model uses the vehicle's center of gravity sideslip angle, yaw rate, longitudinal velocity, lateral error, and heading angle error as state variables, and the front wheel lateral force and rear wheel longitudinal force as control variables; constructing a front wheel tire inverse model based on the tire model, wherein the front wheel tire inverse model is used to characterize the three-dimensional mapping relationship between the front wheel lateral force, road adhesion coefficient, and front wheel sideslip angle; and constructing a rear wheel lateral force prediction model based on the Koopman operator, wherein the rear wheel lateral force prediction model takes the rear wheel sideslip angle, rear wheel longitudinal force, longitudinal velocity, and yaw rate as inputs, and... The rear wheel lateral force is the output; the drift equilibrium point is calculated in real time based on the three-degree-of-freedom vehicle single-track dynamics model; an MPC drift controller is constructed based on the control-guided trajectory tracking dynamics model, the front wheel tire inverse model, the rear wheel lateral force prediction model, and the drift equilibrium point; an LQR lateral controller is constructed based on the control-guided trajectory tracking dynamics model, wherein the LQR lateral controller uses lateral error, lateral error change rate, heading angle error, and heading error change rate as state variables, and the front wheel steering angle as the control variable; a PID longitudinal controller is constructed, which is used to control the error between the desired longitudinal speed and the actual longitudinal speed and output the rear wheel longitudinal force; a DRL agent is constructed and combined with the MPC drift controller, PID longitudinal controller, and LQR lateral controller for mode switching.

[0007] As an improvement to the above scheme, the steps of constructing a three-degree-of-freedom vehicle single-track dynamics model, a tire model, and a control-guided trajectory tracking dynamics model include: constructing a three-degree-of-freedom vehicle single-track dynamics model based on the vehicle's lateral motion, yaw motion, and longitudinal motion; the mechanical parameters of the three-degree-of-freedom vehicle single-track dynamics model include the center of gravity sideslip angle, yaw rate, longitudinal velocity, front wheel lateral force, and rear wheel longitudinal force; characterizing the vehicle's tire lateral force using the Magic Tire formula to generate a tire model; constructing an error model between the vehicle and the reference path; the error parameters of the error model include lateral error and heading angle error; and constructing a control-guided trajectory tracking dynamics model based on the Frenet coordinate system based on the three-degree-of-freedom vehicle single-track dynamics model, the tire model, and the error model.

[0008] As an improvement to the above scheme, the step of constructing the front tire inverse model based on the tire model includes: constructing all the reference combinations between the front tire lateral force, road surface adhesion coefficient and front tire vertical force; calculating the front tire slip angle corresponding to each reference combination according to the Magic Tire Formula; and constructing the three-dimensional mapping relationship between the front tire lateral force, road surface adhesion coefficient and front tire slip angle to form the front tire inverse model.

[0009] As an improvement to the above scheme, the step of constructing a rear wheel lateral force prediction model based on the Koopman operator includes: constructing a nonlinear function, which is used to characterize the transformation of the original input data to a high-dimensional observable vector, wherein the original input data includes the rear wheel slip angle, the rear wheel longitudinal force, the vehicle longitudinal velocity, and the yaw rate; constructing a linear model, which is used to characterize the linear relationship between the high-dimensional observable vector and the rear wheel lateral force; and constructing a fitting model, which is used to fit the original input data using regularized least squares.

[0010] As an improvement to the above scheme, the nonlinear function includes polynomial terms, trigonometric terms, coupling terms, other nonlinear terms, and constant terms; the polynomial terms are used to capture the nonlinear saturation characteristics of the rear wheel sideslip angle; the trigonometric terms are used to describe the periodic or sinusoidal characteristics of the rear wheel sideslip angle; the coupling terms are used to characterize the frictional circular coupling effect of the tire; the other nonlinear terms are used to handle the symmetrical action of the force; and the constant terms are used to transform the input sample into a high-dimensional observable space.

[0011] As an improvement to the above scheme, the step of calculating the drift equilibrium point in real time based on the three-degree-of-freedom vehicle single-track dynamics model includes: constructing the mapping relationship between the yaw rate, longitudinal velocity, reference turning radius, and reference centroid sideslip angle when the vehicle is drifting; constructing all reference combinations between the reference road surface adhesion coefficient, reference turning radius, and reference centroid sideslip angle; and calculating the drift equilibrium point corresponding to each reference combination based on the three-degree-of-freedom vehicle single-track dynamics model, tire model, control guidance trajectory tracking dynamics model, and mapping relationship.

[0012] As an improvement to the above scheme, the step of constructing an MPC drift controller based on the control-guided trajectory tracking dynamics model, the front wheel tire inverse model, the rear wheel lateral force prediction model, and the drift equilibrium point includes: optimizing the tire model based on the front wheel tire inverse model and the rear wheel lateral force prediction model; constructing a linear error model based on the control-guided trajectory tracking dynamics model and the drift equilibrium point; discretizing the linear error model to generate a discrete linear error model; optimizing the discrete linear error model into an augmented model with control increment as input; constructing a linear prediction model based on the augmented model; constructing an objective function, which is to minimize the weighted sum of the state error and the control increment; and integrating the linear prediction model, the objective function, and preset control constraints to construct a quadratic programming model.

[0013] As an improvement to the above scheme, the optimization solution steps of the LQR lateral controller include: calculating the optimal gain matrix; calculating the state feedback rate based on the optimal gain matrix; minimizing the quadratic cost function based on the state feedback rate; minimizing the lateral error and heading error using the minimized quadratic cost function; differentiating the minimized quadratic cost function to generate the state feedback matrix; and calculating the optimal control variables based on the state feedback matrix.

[0014] As an improvement to the above scheme, the state space of the DRL agent includes a vector space of the vehicle's real-time dynamics, path tracking error, and geometric features of the path ahead; the action space of the DRL agent includes a discrete set of maintaining the current mode and requesting a mode switch.

[0015] As an improvement to the above scheme, the reward function of the DRL agent is: ;in, Basic reward items, For mode-related rewards, This is a reward for switching modes.

[0016] Accordingly, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described dynamic drift method based on deep reinforcement learning patterns.

[0017] Accordingly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described dynamic drift method based on deep reinforcement learning patterns.

[0018] This invention presents a dynamic drift method based on deep reinforcement learning to comprehensively improve the cornering performance of vehicles under extreme conditions. Implementing this invention has the following beneficial effects: In this invention, the underlying layer of the DRL agent uses an MPC drift controller for drift control and a PID longitudinal controller and an LQR lateral controller for conventional path tracking, thus achieving targeted processing of drift and tracking. This invention constructs a rear wheel lateral force prediction model based on the Koopman operator, which simplifies the highly nonlinear tire force problem and improves the computational efficiency of the MPC drift controller.

[0019] Furthermore, this invention introduces a reward function for learning the optimal policy switching timing during drift, which can effectively guide the DRL agent to learn to perform switching actions during the optimal drift timing window, rather than simply reacting to the state.

[0020] More preferably, in this invention, the top layer of the DRL agent is trained using the DQN algorithm, which is responsible for the policy switching decision during dynamic drift, and has strong adaptability. Attached Figure Description

[0021] Figure 1 This is a flowchart of an embodiment of the dynamic drift method based on deep reinforcement learning mode of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It is hereby declared that the directional terms such as up, down, left, right, front, back, inside, and outside used in this text are based solely on the accompanying drawings and are not intended to specifically limit the invention.

[0023] See Figure 1 , Figure 1 The flowchart illustrating an embodiment of the dynamic drift method based on deep reinforcement learning patterns of the present invention is shown, which includes: S101, construct a three-degree-of-freedom vehicle single-track dynamics model, tire model and control guidance trajectory tracking dynamics model; The dynamic model for controlling the guiding trajectory tracking uses the vehicle's center of gravity sideslip angle, yaw rate, longitudinal velocity, lateral error, and heading angle error as state variables, and the lateral force of the front wheels and the longitudinal force of the rear wheels as control variables. Accordingly, the steps for constructing a control-guided trajectory tracking dynamics model for trajectory tracking control include: (1) Construct a three-degree-of-freedom vehicle monorail dynamics model based on the vehicle's lateral motion, yaw motion and longitudinal motion; This invention addresses drift control, and therefore employs a three-degree-of-freedom vehicle single-track dynamics model that considers the vehicle's lateral, yaw, and longitudinal motions. (1) (2) (3) in: For vehicle quality; Longitudinal velocity; This is the distance from the front axle to the vehicle's center of gravity. This is the distance from the rear axle to the vehicle's center of gravity. Let be the moment of inertia of the vehicle about its vertical axis; This is the lateral force on the front wheel; This is the lateral force on the rear wheel; This is the longitudinal force of the rear wheel; The steering angle of the front wheels; This refers to the yaw rate; It is the centroid sideslip angle.

[0024] (2) The lateral force of the vehicle's tires is characterized by the magic tire formula to generate a tire model; The lateral force of the tire is calculated using the Magic Tire formula, as shown below: (4) in: The parameters for the Magic Tire formula are calibrated using tire lateral force data from CarSim. The parameters for the Magic Tire formula are calibrated using tire lateral force data from CarSim. The road surface adhesion coefficient; In this embodiment, the coefficient of friction is measured for the tire. The value of is 1; This refers to the lateral force of the tire, and its subscript is... These represent the front wheels and the rear wheels, respectively. The vertical force of the tire, its subscript These represent the front wheels and the rear wheels, respectively. The tire slip angle is indicated by its subscript. These represent the front wheel and the rear wheel, respectively.

[0025] Furthermore, front wheel slip angle The calculation formula is: (5) Correspondingly, rear wheel slip angle The calculation formula is: (6) During typical cornering, the total force on the rear tires is not saturated, but when the vehicle is drifting, the total force on the rear tires reaches saturation. The coupling relationship between the lateral force and the longitudinal force of the rear wheels can be obtained, as shown below: (7) in, For the lateral force of the rear wheel, For the vertical force of the rear wheel, This is the longitudinal force of the rear wheel.

[0026] (3) Construct an error model between the vehicle and the reference path; It should be noted that the error parameters of the error model include lateral error and heading angle error.

[0027] The lateral error is calculated by determining the positional error of the vehicle's current position relative to the nearest point of the trajectory projection in the geodetic coordinate system. , Then the lateral error for: (8) in, For lateral error, This represents the vehicle's positional deviation along the X-axis. The vehicle's X-axis coordinate. The X-axis coordinate of the point closest to the vehicle on the reference trajectory. This represents the vehicle's positional deviation along the Y-axis. For the vehicle's Y-axis coordinate, The Y-axis coordinate of the point closest to the vehicle on the reference trajectory. The reference heading angle is the reference trajectory at the nearest point.

[0028] The heading angle error is: (9) in, For heading angle error, This is the vehicle's heading angle.

[0029] Furthermore, the lateral error, the distance along the reference path, and the heading angle error are used to represent the other three states of vehicle motion, and their kinematic equations can be expressed as follows: (10) (11) in: The vehicle's heading angle; This refers to the yaw rate, also known as the heading angular velocity. is the heading angle error, which is the angle error between the heading angle and the reference heading angle; The curvature of the reference trajectory; This is the velocity of the vehicle's center of gravity.

[0030] (4) Based on the three-degree-of-freedom vehicle single-track dynamics model, tire model and error model, construct a control guidance trajectory tracking dynamics model based on the Frenet coordinate system.

[0031] To achieve trajectory tracking, a control-guided trajectory tracking dynamics model considering tracking errors was designed using the Frenet coordinate system, which describes the error between the vehicle and the reference path.

[0032] Combining formulas (1) to (11), the control-guided trajectory tracking dynamics model used for trajectory tracking control can be summarized as follows: (12) Among them, the state variables of the control-guided trajectory tracking dynamics model are The control quantity (i.e., control input) of the control-guided trajectory tracking dynamics model is: .

[0033] Therefore, this invention is based on a three-degree-of-freedom vehicle single-track dynamics model, uses the magic tire formula to characterize the lateral force of the vehicle's tires, and introduces the Frenet coordinate system to establish a control and guidance trajectory tracking dynamics model to describe the trajectory tracking error.

[0034] S102, Construct the inverse model of the front tires based on the tire model; The inverse model of the front tire is used to characterize the three-dimensional mapping relationship between the front tire lateral force, the road adhesion coefficient, and the front tire slip angle. Accordingly, the steps for constructing the inverse model of the front tire based on the tire model include: (1) Construct all the reference combinations between the lateral force of the front wheel, the road adhesion coefficient and the vertical force of the front wheel; (2) Calculate the front wheel slip angle corresponding to each reference combination according to the magic tire formula; (3) Construct a three-dimensional mapping relationship between the front wheel lateral force, road surface adhesion coefficient and front wheel side slip angle to form a front wheel tire inverse model.

[0035] By inputting the lateral force of the front wheel, the corresponding front wheel steering angle can be mapped using the tire inverse model, effectively handling the nonlinear dynamic characteristics of the front wheel. The specific three-dimensional mapping relationship is shown below: (13) in, For the front wheel steering angle, The obtained front wheel slip angle, For the lateral force of the front wheel, For the vertical force of the front wheel, For longitudinal vehicle speed, This refers to the lateral speed.

[0036] As can be seen from the above, as long as the front wheel is in the unsaturated region, given a certain front wheel lateral force, front wheel vertical force and road adhesion coefficient, there exists a unique front wheel slip angle. By traversing all possible front wheel slip angle candidates, the three-dimensional relationship can be stored as a three-dimensional lookup table. Therefore, after obtaining the optimal front wheel lateral force, the front wheel steering angle required to achieve this force can be quickly solved through this three-dimensional lookup table, which greatly simplifies the calculation.

[0037] S103, Construct a rear wheel lateral force prediction model based on the Koopman operator; The rear wheel lateral force prediction model takes the rear wheel sideslip angle, rear wheel longitudinal force, longitudinal velocity and yaw rate as inputs and the rear wheel lateral force as output. It should be noted that the core of the theory of constructing a rear wheel lateral force prediction model using the Koopman operator is to elevate the state of the nonlinear dynamic system to a high-dimensional observation function space. Although the state evolution of the original system is nonlinear, the observables in this space are linear. Therefore, by finding a finite-dimensional approximation matrix of the operator, a linear model describing the dynamics of the system can be constructed.

[0038] Accordingly, the steps for constructing a rear wheel lateral force prediction model based on the Koopman operator include: (1) Construct a nonlinear function; Nonlinear functions are used to characterize the transformation of raw input data into a high-dimensional observable vector. The raw input data includes rear wheel slip angle, rear wheel longitudinal force, vehicle longitudinal velocity, and yaw rate. In this invention, the training samples are in the radius The input data required for the rear wheel lateral force prediction model is obtained from the steady-state drift condition in a fixed circle. Output data Lateral force of the rear wheel Then, the input and output data for training are standardized by subtracting their mean and dividing by their standard deviation; input data and output data There are complex nonlinear relationships between them. By constructing a nonlinear mapping, the input data can be elevated to a high-dimensional observable vector. The corresponding nonlinear mapping (i.e., the nonlinear function) is as follows: (14) in: It is a nonlinear mapping; For input data, and ; For polynomial terms; These are trigonometric function terms; For coupling terms; For other nonlinear terms; This is a constant term.

[0039] As shown above, nonlinear functions include polynomial terms, trigonometric terms, coupling terms, other nonlinear terms, and constant terms. The following sections will explain polynomial terms, trigonometric terms, coupling terms, other nonlinear terms, and constant terms in detail: I. Polynomial Terms polynomial terms The nonlinear saturation characteristic used to capture the rear wheel sideslip angle is given by the following formula: (15) II. Trigonometric Function Terms Trigonometric function terms The formula used to describe the periodic or sinusoidal characteristics of the rear wheel slip angle is as follows: (16) III. Coupling Terms Coupling terms The formula used to characterize the frictional circular coupling effect of tires is as follows: (17) IV. Other Nonlinear Terms Other nonlinear terms The formula is used to handle symmetrical forces as follows: (18) V. Constant Terms constant term Used to transform input samples into a high-dimensional observable space.

[0040] Under normal circumstances, The constant term, as a bias in a linear system, can be used to represent each input sample. It is transformed into a high-dimensional observable space.

[0041] (2) Construct a linear model, which is used to characterize the linear relationship between the original state variables at the next time step, the high-dimensional observable vector, and the lateral force of the rear wheel;

[0042] in, The state matrix of the linear model. Let be the input matrix of the linear system.

[0043] Output data of the rear wheel lateral force prediction model With observable vector There exists a simple linear relationship This linear relationship (i.e., the linear model) can be expressed as: (19) in, This is the output value of the k-th training sample. Let be the high-dimensional input vector corresponding to the output value of the k-th training sample. Let be the constant row vector (i.e., the linear relationship) to be solved.

[0044] (3) Construct a fitting model, which is used to fit the original input data using regularized least squares.

[0045] To characterize the dynamic evolution of the high-dimensional state and the mapping relationship of the lateral force of the rear wheel, it is necessary to simultaneously solve the system state matrix. Control matrix and output mapping vector Considering that direct solution may lead to overfitting or unstable results due to data noise, this invention uniformly adopts the regularized least squares method for solution.

[0046] definition The current state. For the state at the next moment, To control the input, To represent the actual lateral force of the rear wheel, a joint input matrix is ​​constructed. .so , and The solution expression is as follows: (20) in: and The cross-correlation matrix between the output data and the boosted input was calculated separately. and The Gram matrix of the boosted input was calculated separately. For regularization terms; Through the above data standardization process, an observable measurement space is constructed, and linear systems are identified and models are solved. Linear relationships can then be obtained through training. To predict the lateral force of the rear wheels Thus, a model for predicting the lateral force of the rear wheels is constructed.

[0047] Therefore, this invention designs a set of nonlinear functions consisting of polynomial terms, trigonometric function terms, coupling terms, other nonlinear terms, and constant terms to define the transformation from an original input state containing rear wheel slip angle, rear wheel longitudinal force, longitudinal velocity, and rear wheel vertical force to a high-dimensional observable space. Subsequently, by performing regularized least squares fitting on the original input data, the linear dynamic matrix in this high-dimensional space can be identified, thereby quickly predicting the rear wheel lateral force in a linear operation manner.

[0048] S104, calculates the drift equilibrium point in real time based on the three-degree-of-freedom vehicle monorail dynamics model; The drift equilibrium point refers to the critical point at which a driver can maintain a stable drift state by precisely controlling the steering, accelerator, or brakes during a drift.

[0049] Accordingly, the steps for real-time calculation of the drift equilibrium point based on the three-degree-of-freedom vehicle single-track dynamics model include: (1) Construct the mapping relationship between the yaw rate, longitudinal velocity, reference turning radius and reference centroid sideslip angle when the vehicle is drifting; During steady-state drift, the vehicle state and curvature satisfy the following relationship: (twenty one) (twenty two) Substituting formula (22) into formula (21), we can see that for each point on the reference trajectory, when the reference turning radius... And reference centroid side slip angle It can be seen that the longitudinal velocity can be obtained. With yaw rate Mapping relationship:

[0050] (2) Construct all reference combinations between the reference road surface adhesion coefficient, the reference turning radius, and the reference centroid sideslip angle; (3) Based on the three-degree-of-freedom vehicle single-track dynamics model, tire model, control guidance trajectory tracking dynamics model and mapping relationship, calculate the drift equilibrium point corresponding to each reference combination.

[0051] For a three-degree-of-freedom vehicle single-track dynamics model, let the state derivative... The following algebraic equation can be obtained: (twenty three) (twenty four) (25) in: The longitudinal velocity under drift equilibrium conditions; The lateral force on the front wheel during drift equilibrium; The lateral force on the rear wheel during drift equilibrium; The longitudinal force on the rear wheel during drift equilibrium; The front wheel steering angle under drift equilibrium conditions; The yaw rate under drift equilibrium conditions; Reference centroid sideslip angle.

[0052] By combining equations (4) to (6) and equations (21) to (25), for this system of five variables and three equations, given a certain reference turning radius... And reference centroid side slip angle That is, the reference drift equilibrium state combined with path information can be obtained by solving three equations, and according to the state variables in formula (12), the drift equilibrium point needs to be calculated as the yaw rate in the drift equilibrium state. and longitudinal velocity in drift equilibrium state Therefore, it is possible to iterate through different road surface adhesion coefficients. Reference turning radius And reference centroid side slip angle Generate the yaw rate in the drift equilibrium state. longitudinal velocity in drift equilibrium state The drift equilibrium point lookup table provides a real-time reference equilibrium point that depends on the current state.

[0053] Therefore, based on a three-degree-of-freedom vehicle single-track dynamics model, this invention establishes a mapping relationship between state variables such as yaw rate and longitudinal velocity during vehicle drift and given reference turning radius and reference centroid sideslip angle. By traversing different reference road surface adhesion coefficients, reference turning radii, and reference centroid sideslip angles, a series of drift equilibrium points are generated and stored as a three-dimensional lookup table. This allows for the rapid retrieval of the corresponding drift equilibrium point under the current operating condition based on real-time path information and control objectives.

[0054] S105. Based on the control guidance trajectory tracking dynamics model, the front wheel tire inverse model, the rear wheel lateral force prediction model, and the drift equilibrium point, an MPC (model predictive control) drift controller is constructed. Accordingly, based on the control guidance trajectory tracking dynamics model, the front wheel tire inverse model, the rear wheel lateral force prediction model, and the drift equilibrium point, the steps for constructing the MPC drift controller include: (1) Optimize the tire model based on the inverse model of the front tire and the prediction model of the lateral force of the rear tire; For the control-guided trajectory tracking dynamics model, when and When the time derivatives are all 0, the drift equilibrium point can be obtained. , ); Accordingly, state error can be defined. for: (26) In addition, control input error for: (27) in: It is a state variable, and ; To control the input, and .

[0055] This invention decouples and linearizes the nonlinear tire force, and transforms the rear wheel lateral force in formula (4) into a linear equation. The rear wheel lateral force prediction model constructed using formulas (14) to (20) is used instead, and the front wheel lateral force is predicted using the front wheel tire inverse model of formula (13). Converted to front wheel steering angle Command to avoid introducing lateral forces to the front wheels With front wheel steering angle The strong coupling nonlinear relationship between them.

[0056] (2) Construct a linear error model based on the dynamic model of the control-guided trajectory tracking and the drift equilibrium point; It should be noted that step (1) above only linearizes the main nonlinear dynamics in the control problem; the lateral force of the front wheels still exists in the control-guided trajectory tracking dynamics model. With front wheel steering angle The nonlinear term. Therefore, it is necessary to modify the dynamic model of the control-guided trajectory tracking at any drift equilibrium point ( , A first-order Taylor expansion near the 0° ... (28)

[0057]

[0058] in, and These are the Jacobian matrices calculated at the equilibrium point, respectively.

[0059] (3) Discretize the linear error model to generate a discrete linear error model; Discretizing the linear error model of formula (28) yields the discrete linear error model: (29)

[0060]

[0061] in, Sampling time, It is an identity matrix.

[0062] (4) Optimize the discrete linear error model into an augmented model with the control increment as input; Set the optimization variable as the control increment. The discrete linear error model can be rewritten as an augmented form with the control increment as input: (30) (31) in:

[0063]

[0064]

[0065]

[0066] (5) Construct a linear prediction model based on the augmented model; Define the prediction time domain as The control time domain is defined as , The linear prediction model for system variables and system output can be expressed as follows: (32) in:

[0067]

[0068]

[0069]

[0070] (6) Construct the objective function; To ensure the performance of the MPC drift controller, an objective function needs to be designed to guarantee the ability to track the drift equilibrium state, maintain control stability, and ensure the feasibility of the solution at each time step. The corresponding objective function is as follows: (33) in, Here is the weight matrix for the state error. To control the increment matrix, As a relaxation factor, This is the weight matrix for the relaxation factor.

[0071] As shown above, the objective function is to minimize the weighted sum of the state error and the control increment.

[0072] (7) Integrate the linear prediction model, objective function and preset control constraints to construct a quadratic programming model.

[0073] To facilitate the calculation of the optimization problem and ensure real-time performance, the optimization problem in formula (33) is transformed into a quadratic programming model through matrix transformation: (34)

[0074] in:

[0075]

[0076] To control the lower limit of input increment; To control the upper limit of input increment; To control the lower limit of the input; To control the upper limit of input.

[0077] It should be noted that the lateral force of the front wheels The upper and lower limits are determined by the vehicle's vertical force and the maximum lateral grip that the current road surface adhesion coefficient can provide; the longitudinal force of the rear wheels The upper and lower limits are constrained by the theory of rear wheel friction circle, and are usually taken as the conservative extreme value as a fixed constraint; the control input increment is usually used as the controller tuning parameter and is determined through simulation experiments.

[0078] Therefore, this invention constructs an MPC drift controller by using a dynamic model for controlling the directional trajectory tracking, an inverse model for the front tires, a prediction model for the lateral force of the rear wheels, and a drift equilibrium point. By integrating the prediction equation, objective function, and control constraints, a standard quadratic programming problem is formed, which is then solved to obtain the optimal lateral force of the front wheels and the longitudinal force of the rear wheels.

[0079] S106, Construct an LQR (linear-quadratic regulator) lateral controller based on the control-guided trajectory tracking dynamics model; It should be noted that the LQR lateral controller adopts a control-guided trajectory tracking dynamic model, with lateral error, lateral error rate of change, heading angle error, and heading error rate of change as state variables, and the front wheel steering angle as the control variable. , .

[0080] Accordingly, the optimization solution steps for the LQR lateral controller include: (1) Calculate the optimal gain matrix; (2) Calculate the state feedback rate based on the optimal gain matrix; (3) Minimize the quadratic cost function based on the state feedback rate; In the optimization process, the optimal gain matrix is ​​calculated. , to make the state feedback rate Minimize the quadratic cost function: (35) in, Let cost function be Let be the state error vector. The state weight matrix is... To control the weight matrix.

[0081] (4) Minimize the lateral error and heading error by minimizing the quadratic cost function; By minimizing the lateral and heading errors in the control-guided trajectory tracking dynamics model using a quadratic cost function, stable tracking of the desired path can be achieved.

[0082] (5) Take the derivative of the minimized quadratic cost function to generate the state feedback matrix; Taking the derivative of equation (35), we can obtain the state feedback matrix. The expression: (36) in, To control the input matrix, Solve the matrix for the Riccati equation. The system state matrix, This is the cross-weight matrix.

[0083] (6) Calculate the optimal control variables based on the state feedback matrix.

[0084] Based on the state feedback matrix The optimal control variables can be calculated. To achieve optimal lateral control: (37) Therefore, in terms of lateral control, the present invention uses an LQR lateral controller; the LQR lateral controller aims to minimize the lateral error, the rate of change of the lateral error, the heading angle error, and the rate of change of the heading angle error, and calculates the optimal front wheel steering angle by solving the Riccati equation.

[0085] S107, Construct a PID (proportional integral derivative) longitudinal controller; It should be noted that the PID longitudinal controller is used to control the error between the desired longitudinal speed and the actual longitudinal speed and output the longitudinal force of the rear wheels.

[0086] The expression for a PID longitudinal controller is as follows: (38)

[0087]

[0088] in, For proportional gain, For integral gain, This is the differential gain.

[0089] Therefore, in terms of longitudinal control, this invention designs a PID longitudinal controller based on the deviation between the desired longitudinal speed and the actual speed, with the desired longitudinal speed as the target, and outputs longitudinal driving force.

[0090] S108, Construct a DRL (deep reinforcement learning) agent and combine it with the MPC drift controller, PID longitudinal controller and LQR lateral controller to perform mode switching; The state space and action space of the DRL agent are explained in detail below: I. State Space The state space of the DRL agent includes a vector space of the vehicle's real-time dynamics, path tracking error, and geometric features of the path ahead. In this invention, the state space of the DRL agent is designed as an 11-dimensional vector space that includes the vehicle's real-time dynamics, path tracking error, and geometric features of the path ahead. , The state vector At each time step The state components are constructed and input into the DRL agent; after all state components are combined into the final 11-dimensional vector, they are linearly normalized to [the specified values ​​based on preset upper and lower bounds]. The interval [1,1] is used to accommodate the input requirements of the neural network.

[0091] II. Movement Space The action space of a DRL agent consists of a discrete set of actions for maintaining the current mode and for requesting a mode switch.

[0092] In this invention, the action space of the DRL agent is designed as a discrete set. The continuous values ​​output by the DRL agent are mapped to two discrete actions. To achieve integrated control of conventional path tracking and dynamic drift, this invention decomposes the dynamic cornering task into three logical modes to construct a clearly structured decision-making environment. These three modes are: (1) Cruise mode, the controller is a PID longitudinal controller / LQR lateral controller, used for stable path tracking under normal road conditions, and the DRL agent does not participate in decision-making in this mode; (2) Drift preparation mode Ready, the controller is a PID longitudinal controller / LQR lateral controller, as a transition stage before entering a sharp turn, allowing the DRL agent to decide the best time to drift in this mode; (3) Drift control mode: The controller is the MPC drift controller, which performs drifting action in sharp turns, and the DRL agent determines the best exit time.

[0093] Therefore, when action 0 is performed, the DRL agent requests to remain in Ready or Drift; when action 1 is performed, the DRL agent requests to switch the control strategy, that is, to switch from Ready to Drift or from Drift to Cruise.

[0094] Furthermore, in order to ensure the decision-making safety of the DRL agent in dynamic drift tasks, this invention designs a curve entry judgment architecture based on safety boundary constraints, which integrates deterministic safety boundary rules with the agent's learning strategy.

[0095] First, the vehicle enters Cruise mode on a straight road, at which point safety boundary determination is placed at the highest level of decision-making. The safety boundary relies on path prediction and is forcibly triggered by a deterministic rule based on the geometry of the road ahead, with the following logic: (1) Determine whether the curvature of the reference point of the vehicle's current path is very small; (2) Whether the maximum curvature within the preview range is large enough; (3) The trend of the curvature in front must be positive.

[0096] Therefore, switching from Cruise to Ready mode is only allowed after the switching conditions are met. The decision-making power is only delegated to the DRL agent for the most timing-critical decisions—entering and exiting drift—allowing it to utilize its learned timing judgment capabilities. This design, through hierarchical decision-making, ensures that the DRL agent is authorized to execute the optimal switching decision only when the safety boundary conditions are fully met.

[0097] Secondly, since the maintenance strategy and the switching strategy constitute a discrete action space, this invention selects the DQN algorithm. Specifically, the DQN algorithm is based on a five-layer neural network architecture, including one input layer, three hidden layers, and one output layer; each hidden layer has 256 neurons and uses the ReLU activation function; the target update frequency... The learning rate is 1. The gradient is 0.001, and the Adam optimizer is used for gradient update with a gradient threshold of 1; the exploration rate is... The initial value is set to 1, and it decays at a rate of 0.001 until it reaches the minimum exploration rate of 0.01, after which it remains constant; discount factor The value is 0.995, for the experience replay pool. Batch size: 100,000 The maximum number of training rounds is 256. The value is 1000; the training track is a single U-turn track with a 40m straightaway. and a transition curve of 30-50m And the 18-26m circular arc section Random selection is used to generate the data. The algorithm flow is shown in Table 1 below: Table 1

[0098] Furthermore, to guide the reinforcement learning agent in learning complex drift driving strategy switching, this invention designs a hierarchical, event- and state-based reward function, aiming to balance vehicle stability, path tracking accuracy, drift posture, and the accuracy of decision-making timing. Specifically, the reward function of the DRL agent... Represented as: (39) in, Basic reward items, For mode-related rewards, This is a reward for switching modes.

[0099] The following sections provide detailed explanations of the basic rewards, mode-related rewards, and mode-switching rewards: I. Basic Rewards Basic Rewards Includes a constant survival reward. To incentivize the agent to continue operating, and a penalty term based on lateral error. To prevent vehicles from deviating from the preset path. That is, the basic reward consists of two parts: (40) in: Survival Rewards Mainly each time step Give the agent a small positive reward to encourage it to continue operating; Off-track penalty When the vehicle's lateral error The absolute value exceeds the preset threshold At that time, a huge negative penalty is imposed.

[0100] II. Rewards Related to the Model Mode-related rewards Different reward signals are provided based on the vehicle's current control mode. This design guides the DRL agent to perform specific objectives at different stages of the task, providing timely rewards in the Ready mode. Optimize cornering timing; achieve a comprehensive total drift bonus in Drift mode. To help the agent determine the appropriate drift duration; in straight-line mode, Cruise uses instability penalties. This helps the agent determine its recovery status after exiting a bend. Specifically, it is represented as follows: (41) in: Timing Rewards The reward or penalty is based on how close the magnitudes of the sideslip angle and yaw rate are to the expected drift value when the vehicle switches from Ready to Drift strategy. Total Drift Rewards Including side slip angle bonus Yaw rate bonus Efficiency rewards and improper drift penalty Among them, the side slip angle bonus Used to encourage vehicles under the Drift strategy Close Yaw rate bonus Used to encourage vehicles under the Drift strategy Close Efficiency Rewards Used to encourage maintaining a relatively stable speed during drifting; penalty for improper drifting. Used for road sections that are close to a straight line (i.e., the current curvature). Less than the maximum linear curvature Drifting will be penalized.

[0101] Instability penalty This is when the vehicle is in straight-line mode (Cruise). Exceeding the stability threshold Punishment will be imposed.

[0102] III. Rewards for Mode Switching Rewards for mode switching Triggered only when a specific state transition occurs, and rewarded by the timing of exiting the corner. and stable bonus after exiting the corner This is a component used to directly evaluate the merits of the switching behavior itself. That is, the reward for mode switching. It consists of two parts: (42) in: Corner exit timing bonus When a vehicle switches from the Drift strategy to the Cruise strategy, the timing of exiting a corner is penalized based on geometric characteristics, using a corner exit weight coefficient. Multiply by the mass of the exit curve The out-of-bend mass is , =0.5 / 1.0, =0.5 / deg2rad(15), where 1 and deg2rad(15) are defined unit penalty errors. Error normalization is performed by dividing the actual error by the unit penalty value. 0.5 is the overall sensitivity coefficient, which determines the rate at which the exit quality decreases with respect to this total penalty term.

[0103] Stable bonus after exiting the corner It rewards a reduction in the sideslip angle and penalizes an increase in it shortly after exiting a corner.

[0104] Therefore, by using a comprehensive, hierarchical reward function, the DRL agent can be effectively guided to perform switching actions at the most appropriate time window and learn the optimal strategy.

[0105] Accordingly, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described dynamic drift method based on deep reinforcement learning patterns.

[0106] Meanwhile, the present invention also discloses a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described dynamic drift method based on deep reinforcement learning patterns.

[0107] In summary, the dynamic drift method based on deep reinforcement learning in this invention comprehensively improves the cornering performance of vehicles under extreme conditions. Implementing this invention has the following beneficial effects: (1) This invention proposes a hierarchical dynamic drift control architecture; wherein, the top layer uses a DRL agent trained with a deep Q-network (DQN) to be responsible for the policy switching decision of dynamic drift, the bottom layer uses an MPC drift controller to be responsible for drift control, and a PID longitudinal controller and an LQR lateral controller to be responsible for conventional path tracking. (2) This invention designs a reward function for learning the optimal policy switching timing during drift; it models the drift policy switching as a Markov decision process (MDP) and effectively guides the DRL agent to learn to perform the switching action at the optimal drift timing window, rather than simply reacting to the state; (3) This invention proposes a linear prediction model for rear wheel lateral force based on Koopman operator theory, which simplifies the highly nonlinear tire force problem and improves the computational efficiency of MPC drift controller.

[0108] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A dynamic drift method based on deep reinforcement learning patterns, characterized in that, include: The three-degree-of-freedom vehicle single-track dynamics model, tire model and control-guided trajectory tracking dynamics model, wherein the control-guided trajectory tracking dynamics model uses the vehicle's center of gravity sideslip angle, yaw rate, longitudinal velocity, lateral error and heading angle error as state variables, and the front wheel lateral force and rear wheel longitudinal force as control variables. Based on the tire model, a front tire inverse model is constructed. The front tire inverse model is used to characterize the three-dimensional mapping relationship between the front tire lateral force, road adhesion coefficient and front tire slip angle. A rear wheel lateral force prediction model based on the Koopman operator is constructed. The rear wheel lateral force prediction model takes the rear wheel sideslip angle, rear wheel longitudinal force, longitudinal velocity and yaw rate as inputs and the rear wheel lateral force as output. The drift equilibrium point is calculated in real time based on the three-degree-of-freedom vehicle monorail dynamics model. Based on the control guidance trajectory tracking dynamics model, the front tire inverse model, the rear wheel lateral force prediction model, and the drift equilibrium point, an MPC drift controller is constructed. An LQR lateral controller is constructed based on the control guidance trajectory tracking dynamic model. The LQR lateral controller uses lateral error, lateral error change rate, heading angle error, and heading error change rate as state variables, and front wheel angle as control variable. A PID longitudinal controller is constructed, which is used to control the error between the desired longitudinal speed and the actual longitudinal speed and output the longitudinal force of the rear wheel; A DRL agent is constructed and combined with the MPC drift controller, PID longitudinal controller and LQR lateral controller to perform mode switching.

2. The dynamic drift method based on deep reinforcement learning patterns as described in claim 1, characterized in that, The steps for constructing the three-degree-of-freedom vehicle single-track dynamics model, tire model, and control guidance trajectory tracking dynamics model include: A three-degree-of-freedom vehicle single-track dynamics model is constructed based on the vehicle's lateral motion, yaw motion, and longitudinal motion. The mechanical parameters of the three-degree-of-freedom vehicle single-track dynamics model include the center of gravity sideslip angle, yaw rate, longitudinal velocity, front wheel lateral force, and rear wheel longitudinal force. The lateral forces of a vehicle's tires are characterized using the Magic Tire Formula to generate a tire model. Construct an error model between the vehicle and the reference path, wherein the error parameters of the error model include lateral error and heading angle error; Based on the three-degree-of-freedom vehicle single-track dynamics model, tire model, and error model, a control and guidance trajectory tracking dynamics model based on the Frenet coordinate system is constructed.

3. The dynamic drift method based on deep reinforcement learning patterns as described in claim 1, characterized in that, The step of constructing the inverse model of the front tire based on the tire model includes: Construct all the reference combinations between the front wheel lateral force, road surface adhesion coefficient, and front wheel vertical force; Based on the Magic Tire Formula, calculate the front wheel slip angle corresponding to each benchmark combination; A three-dimensional mapping relationship is constructed between the front wheel lateral force, the road surface adhesion coefficient, and the front wheel sideslip angle to form a front wheel tire inverse model.

4. The dynamic drift method based on deep reinforcement learning patterns as described in claim 1, characterized in that, The steps for constructing a rear wheel lateral force prediction model based on the Koopman operator include: A nonlinear function is constructed to characterize the transformation of the original input data into a high-dimensional observable vector. The original input data includes the rear wheel slip angle, the rear wheel longitudinal force, the vehicle longitudinal velocity, and the yaw rate. A linear model is constructed to characterize the linear relationship between the high-dimensional observable vector and the lateral force of the rear wheel; A fitting model is constructed, which is used to fit the original input data using regularized least squares.

5. The dynamic drift method based on deep reinforcement learning patterns as described in claim 4, characterized in that, The nonlinear function includes polynomial terms, trigonometric function terms, coupling terms, other nonlinear terms, and constant terms; The polynomial terms are used to capture the nonlinear saturation characteristics of the rear wheel sideslip angle; The trigonometric function terms are used to describe the periodic or sinusoidal characteristics of the rear wheel slip angle; The coupling term is used to characterize the frictional circular coupling effect of the tire; The other nonlinear terms are used to handle the symmetrical action of forces; The constant term is used to transform the input sample into a high-dimensional observable space.

6. The dynamic drift method based on deep reinforcement learning patterns as described in claim 2, characterized in that, The step of calculating the drift equilibrium point in real time based on the three-degree-of-freedom vehicle single-track dynamics model includes: Construct the mapping relationship between the yaw rate, longitudinal velocity, reference turning radius, and reference centroid sideslip angle during vehicle drift; Construct all reference combinations between the reference road surface adhesion coefficient, the reference turning radius, and the reference centroid sideslip angle; Based on the three-degree-of-freedom vehicle single-track dynamics model, tire model, control guidance trajectory tracking dynamics model and mapping relationship, the drift equilibrium point corresponding to each reference combination is calculated respectively.

7. The dynamic drift method based on deep reinforcement learning patterns as described in claim 2, characterized in that, The steps for constructing the MPC drift controller based on the control guidance trajectory tracking dynamics model, the front wheel tire inverse model, the rear wheel lateral force prediction model, and the drift equilibrium point include: The tire model is optimized based on the inverse model of the front tire and the lateral force prediction model of the rear tire. A linear error model is constructed based on the control-guided trajectory tracking dynamics model and the drift equilibrium point; The linear error model is discretized to generate a discrete linear error model; The discrete linear error model is optimized into an augmented model with the control increment as input; Construct a linear prediction model based on the augmented model; Construct an objective function that minimizes the weighted sum of state error and control increment; The linear prediction model, objective function, and preset control constraints are integrated to construct a quadratic programming model.

8. The dynamic drift method based on deep reinforcement learning patterns as described in claim 1, characterized in that, The optimization solution steps for the LQR lateral controller include: Calculate the optimal gain matrix; Calculate the state feedback rate based on the optimal gain matrix; Minimize the quadratic cost function based on the state feedback rate; The lateral error and heading error are minimized by minimizing the quadratic cost function. Differentiate the minimized quadratic cost function to generate the state feedback matrix; Calculate the optimal control variables based on the state feedback matrix.

9. The dynamic drift method based on deep reinforcement learning patterns as described in claim 1, characterized in that, The state space of the DRL agent includes a vector space of the vehicle's real-time dynamics, path tracking error, and geometric features of the path ahead. The action space of the DRL agent includes a discrete set of modes for maintaining the current mode and for requesting a mode switch.

10. The dynamic drift method based on deep reinforcement learning patterns as described in claim 1, characterized in that, The reward function of the DRL agent is: in, Basic reward items, For mode-related rewards, This is a reward for switching modes.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the dynamic drift method based on deep reinforcement learning patterns as described in any one of claims 1 to 10.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the dynamic drift method based on deep reinforcement learning patterns as described in any one of claims 1 to 10.

Citation Information

Cited By

  • Hybrid-order Koopman tensor decomposition and CBF fused vehicle limit control method

    CN121947518A