Tire force online estimation method based on deep reinforcement learning
Through the deep reinforcement learning method, an MPC controller and magic tire model are designed, and the actor-critic neural network is used to perform online tire force estimation, which solves the data collection problems under high cost and extreme working conditions, realizes accurate estimation of tire force in the nonlinear region, and improves vehicle stability and safety.
Patent Information
- Application Number
- CN202510738761.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-05
AI Technical Summary
Existing tire force estimation technologies have the disadvantages of high computational complexity, high data acquisition cost, and high risk under extreme working conditions. In addition, traditional methods have large estimation errors in the nonlinear region of the tire, making it difficult to achieve accurate online tire force estimation.
An MPC controller and magic tire model are designed based on deep reinforcement learning. An intelligent agent is constructed through an actor-critic neural network and trained with the Carsim/Simulink model to achieve online tire force estimation.
It reduces data acquisition costs and risks, improves the accuracy and real-time performance of tire force estimation, adapts to complex driving scenarios, and enhances vehicle stability and safety.
Smart Images

Figure CN120597712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle tire force estimation, and in particular to a method for online tire force estimation based on deep reinforcement learning. Background Art
[0002] Tire force is a key factor affecting vehicle stability and safety. Existing tire force estimation techniques primarily rely on two approaches: those based on Kalman filtering and those based on tire model numerical calculations. The former suffers from exponentially increasing algorithmic complexity due to difficulties linearizing nonlinear systems, while the latter is limited by the structural characteristics of the model. Tire models can be categorized into three main types: theoretical, semi-empirical, and empirical. Theoretical models suffer from insufficient analytical accuracy and computational dimensionality explosion. Semi-empirical models, limited by their simplified assumptions, still require extensive calibration. Parameter determination for empirical models, such as the magic tire model, requires high experimental costs and suffers from poor generalization.
[0003] Currently, the main approaches to Magic Tire parameter estimation include neural networks and genetic algorithms. Genetic algorithms rely on initial values, and when these values are poorly chosen, they can easily fall into local optima, resulting in poor Magic Tire parameter identification. Neural network-based tire force estimation requires large datasets, but acquiring real-world vehicle data is expensive and highly risky under extreme operating conditions. Furthermore, the Magic Tire parameters identified by these methods are all constant values, but tires are highly nonlinear. Using constant values to calculate tire forces increases estimation errors when the tire enters the nonlinear region. Summary of the Invention
[0004] The purpose of this invention is to propose a method for online estimation of tire forces based on deep reinforcement learning, which better realizes the calculation of tire forces and the nonlinear description of tire forces, thereby improving the safety and stability of the vehicle.
[0005] To achieve the above objectives, the present invention proposes a method for online tire force estimation based on deep reinforcement learning, which includes the following steps:
[0006] Step S1, designing an MPC controller, establishing a two-degree-of-freedom dynamic model of the vehicle and a two-degree-of-freedom linear time-varying discrete prediction model of the vehicle;
[0007] Step S2: designing a tire model based on the magic tire calculation formula;
[0008] Step S3: Standardize the observation space and action space, where the observation space includes the road adhesion coefficient, tire slip angle, vertical load, and lateral acceleration, and the action space is the parameters of the magic tire calculation formula;
[0009] Step S4, defining the reinforcement learning stage reward function, error termination function and vehicle initialization function;
[0010] Step S5: constructing an Actor-Critic neural network based on the state space and the action space;
[0011] Step S6: Establishing a deep reinforcement learning agent update and iteration mechanism;
[0012] Step S7: Build and train a high-fidelity Carsim / Simulink joint deep reinforcement learning model to verify the accuracy of tire force estimation.
[0013] Preferably, in step S1, the specific steps of designing the MPC controller are as follows:
[0014] Step S11: Establish a two-degree-of-freedom dynamic model of the vehicle according to Newton's second law. The formula is as follows:
[0015]
[0016] Where m is the vehicle mass, v2 and v1 are the lateral and longitudinal velocities of the vehicle, K1 and K2 are the cornering stiffness of the front and rear axles, a and b are the distances from the center of mass of the vehicle to the front and rear axles, σ is the front wheel turning angle of the vehicle, and I zz is the moment of inertia, ω is the vehicle yaw rate, is the vehicle yaw angle angular acceleration;
[0017] Step S12: Establish a vehicle lateral tracking error e d , lateral tracking error change rate Vehicle heading angle error and the rate of change of vehicle heading angle error The vehicle tracking error dynamics model for state information;
[0018] Step S13: According to the vehicle tracking error dynamics model, select is the state variable, u(t)=σ is the control variable, satisfying the state space equation:
[0019]
[0020] in, is the state vector, f[·] is the nonlinear vector function;
[0021] Step S14: According to step S13, the state space equation is obtained, and the working point (ξ(t), u(t)) of the vehicle tracking error dynamics model is linearized at time t to obtain a linear time-varying equation, which is as follows:
[0022]
[0023] in, is the current state vector, A d (t) is the state Jacobian matrix, B d (t) is the input Jacobian matrix, x(t) is the original state vector;
[0024] Step S15: discretize the linear time-varying equation obtained in step S14 using the forward Euler method. The discretized equation at time k is:
[0025]
[0026] Among them, A d (k)=I+T′A(t), A d (k) is the discrete-time state matrix, A(t) is the continuous-time state matrix, and B d (k) = T′B(t), B d (k) is the discrete-time input matrix, B(t) is the continuous-time input matrix, is the control increment, T′ is the control step size, and I is the identity matrix;
[0027] Step S16: establishing a two-degree-of-freedom linear time-varying discrete prediction model for the vehicle according to the discretization equation of step S15;
[0028] Step S17: Design a cost function according to step S16, taking into account the error size and the size of the control amount in the path tracking process, achieving a good path tracking effect with the minimum control amount, and introducing weighted coefficients and relaxation factors to design the cost function J, and perform quadratic programming solution.
[0029] Preferably, in step S12, the calculation formula of the vehicle tracking error dynamics model is as follows:
[0030]
[0031] in, is the state space equation of the vehicle tracking error dynamics model; u = delta_f, is the front wheel angle; is the system state matrix; is the input matrix; is the interference matrix, is the reference heading angular velocity.
[0032] Preferably, in step S16, based on ζ(k+1|t)=[ξ(k|t)u(k-1|t)] T , the state space equation of the prediction model is obtained as:
[0033]
[0034] in, ζ(k|t) is the extended state vector, To expand the discrete state matrix, is the expanded discrete input matrix, Δu(k|t) is the control increment, η(k|t) is the predicted input vector, is the expanded output matrix, A k,t is the state matrix of the original linearized system, B k,t is the input matrix, I m is the m-dimensional unit matrix, n is the state dimension, and m is the control dimension;
[0035] Assumptions The prediction interval is N p , the control interval is N c , and get the prediction model equation:
[0036] Y(t)=ψζ(t|t)+ΘΔU(t);
[0037] in, Predict output values for the system at multiple future moments, To reflect the impact of state variables on future output, is the future control increment, ΔU(t) is the dynamic response to the future output, η(t+f|t) is the system predicted output based on time t in the f-th step, for f-th power, used to recursively predict the state of the next f steps, f = 1, 2, ..., N p , is the input matrix treated as a constant.
[0038] Preferably, in step S17, a cost function J is designed and a quadratic programming solution is performed. The specific formula is as follows:
[0039] minJ=min[ΔU(t) T ,ε]H t [ΔU(t) T ,ε] T +G t [ΔU(t) T ,ε];
[0040] Where ε is the relaxation factor, T is the matrix transpose, is the Hessian matrix, Q is the weighted matrix of the state variables, θ is the rolling prediction matrix, ρ is the proportional coefficient, G t =[2E(t) T Qθ t ] is the coefficient vector of the first-order term in the cost function, E(t)=ψ(t)ξ(t)-Y ref(t) is the initial error between the predicted output and the reference trajectory, ψ(·) is the state input mapping matrix, ξ(·) is the error state vector, and Y ref (·) is the reference trajectory output sequence, θ t To control the gain;
[0041] The optimal sequence obtained in the control time domain is: ΔU t =[Δu(t),Δu(t+1),...,Δu(t+N c -1)] T , where Δu(t) is the current control increment.
[0042] Preferably, in step S2, a tire model is designed based on the magic tire calculation formula, which is as follows:
[0043] F y =-D*sin(C*atan(B*xE*atan(B*x-atan(B*x))))+Sv;
[0044] Among them, F y is the tire lateral force, D is the peak factor, B is the stiffness factor, E is the curvature factor, C is the curve shape factor, Sv is the vertical drift of the curve, and x is the tire side slip angle.
[0045] Preferably, in step S3, the observation space includes the road adhesion coefficient μ, the tire side slip angle α, the vertical load F z , vehicle lateral acceleration a y ,The action space is the five secondary parameters of the magic tire formula, which are C, B, D, E, and Sv.
[0046] Preferably, in step S4, the calculation formula of the stage-by-stage reward function is as follows:
[0047]
[0048] Among them, reward is the reward value, force is the current tire force, and err is the absolute value of the error between the tire forces;
[0049] The calculation formula of the error termination function is as follows:
[0050]
[0051] Among them, 1 means stop training, 0 means continue training, and err_rate represents the relative error;
[0052] The calculation formula of the vehicle initialization function is as follows:
[0053] initialSpeed=30+rand*100;
[0054] Among them, initialspeed is the initial speed, and rand is a random number in [0,1].
[0055] Preferably, in step S5, the Critic network structure includes an input layer, a hidden layer and an output layer, the input layer includes an obsInLyr layer and an actInlyr layer, both layers have a dimension of 5, the hidden layer includes an add layer and multiple fully connected layers and a ReLu layer, each fully connected layer contains 80 neurons, and QValLyr is the output layer; the Actor network structure includes an input layer, a hidden layer and an output layer, the number of fully connected layer neurons in the hidden layer is 100, and the output layer includes a fully connected layer, a tanhLayer activation layer and a scalingLayer scaling layer.
[0056] Preferably, in step S6, a deep reinforcement learning agent update and iteration mechanism is established, and the specific steps are as follows:
[0057] Step S61: Through the vehicle initialization function, each training vehicle is put into a different state;
[0058] Step S62: Update the agent. The steps are as follows:
[0059] Step S621: Initialize the critic network Q(S,A;φ) and initialize the target critic network parameter φ t =φ, where S is the state space, A is the action space, φ is the Critic network parameter, t is the target critic network parameter;
[0060] Step S622: Initialize the Actor network π(s;θ) and initialize the target Actor network parameters θ t =θ, where s is the current observation state;
[0061] Step S623: Initialize the experience buffer R;
[0062] Step S624, initializing random detection noise N;
[0063] Step S625: Select action A based on the current state and the detected noise N t =π(S t ;θ)+N t , where A t is the action at time t, S t is the observed state at time t, N t is the detection noise at time t;
[0064] Step S626: Execute action A t , observed reward value rt and the next observation state S t+1 ;
[0065] Step S627: Store transition (S t ,A t ,r t ,S t+1 ) to the experience buffer R;
[0066] Step S628: Randomly extract K experiences (S i ,A i ,r i ,S i+1 ), randomly sampled by MiniBatchSize, where S i is the state in the i-th experience, A i is the action in the i-th experience, r i is the reward in the i-th experience, S i+1 is the next state in the i-th experience, i is a positive integer;
[0067] Step S629: Set the value function y i =r i +γQ t (S i+1 ,π t (S i+1 θ t );φ t ), where y i is the TD target value, γ is the discount factor, π t (·) is the current Actor strategy network, Q t (·) is the current Critic valuation network;
[0068] Step S6210: Update the critic network using the minimum loss function. The loss function is as follows:
[0069]
[0070] Among them, K is the MiniBatcSize size, and M is the number of samples involved in gradient calculation;
[0071] Step S6211: Update the Actor parameters using the strategy to maximize the expected discounted reward:
[0072]
[0073] Where A=π(S i ;θ), is the gradient operator with respect to the parameter θ, is the gradient operator for action A;
[0074] Step S6212: Soft-update the target network parameters using the following smoothing factor τ:
[0075]
[0076] Therefore, the present invention proposes a method for online tire force estimation based on deep reinforcement learning, which has the following beneficial effects:
[0077] (1) Reduce costs and data acquisition risks: The present invention uses deep reinforcement learning technology, eliminating the need to rely on a large number of real-vehicle tests to acquire data, thus avoiding the high cost of data acquisition and the risks of data acquisition under extreme working conditions, thereby improving the safety and economy of data acquisition.
[0078] (2) Adaptive Tire Parameter Adjustment: Utilizing a deep reinforcement learning agent, the Magic Tire parameters can be adaptively adjusted based on vehicle body state information. Under different driving conditions, the tire model parameters can dynamically change. Compared to traditional tire models with fixed parameters, this model can better adapt to complex and changing real-world driving scenarios, improving the real-time and accuracy of tire force estimation.
[0079] (3) Improving tire force estimation accuracy: The deep reinforcement learning model is trained to learn the complex patterns of tire force variations. When the tire enters the nonlinear region, the model adaptively adjusts parameters to overcome the problem of increased estimation error caused by using fixed values to calculate tire force. This significantly improves the accuracy of tire force estimation and provides more reliable data support for vehicle stability and safety control.
[0080] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 This is a flow chart of a method for online tire force estimation based on deep reinforcement learning according to the present invention;
[0082] Figure 2 This is the Critic network structure diagram in the present invention;
[0083] Figure 3 This is the Actor network structure diagram in the present invention;
[0084] Figure 4 A flowchart for updating the reinforcement learning agent in the present invention;
[0085] Figure 5 This is a comparison diagram of the front tire forces under a 72km double lane change condition in an embodiment of the present invention;
[0086] Figure 6This is a lateral acceleration diagram of a vehicle under a double lane-changing condition at 72 km / h in an embodiment of the present invention;
[0087] Figure 7 This is a comparison diagram of the front tire forces under a 108km continuous lane change condition in an embodiment of the present invention;
[0088] Figure 8 This is a diagram of the vehicle's lateral acceleration during a 108km / h continuous lane change condition in an embodiment of the present invention. DETAILED DESCRIPTION
[0089] To make the technical solutions, advantages, and objectives of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below. The described embodiments are part of the embodiments of the present invention, not all of them. Based on the described embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0090] Unless otherwise defined, technical or scientific terms used in the present invention shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs.
[0091] like Figure 1 FIG. 1 is a flowchart of a method for online tire force estimation based on deep reinforcement learning according to the present invention. The specific steps are as follows:
[0092] S1. Design an MPC controller and establish a two-degree-of-freedom vehicle dynamics model and a two-degree-of-freedom linear time-varying discrete prediction model for the vehicle. The specific steps for designing the MPC controller are as follows:
[0093] S11. According to Newton's second law, a two-degree-of-freedom dynamic model of the vehicle is established. The formula is as follows:
[0094]
[0095] Where m is the vehicle mass, v2 and v1 are the lateral and longitudinal velocities of the vehicle, K1 and K2 are the cornering stiffness of the front and rear axles, a and b are the distances from the center of mass of the vehicle to the front and rear axles, σ is the front wheel turning angle of the vehicle, and I zz is the moment of inertia, ω is the vehicle yaw rate, is the vehicle yaw angle angular acceleration;
[0096] S12. In order to realize MPC control tracking of the target path, a dynamic tracking error model is established based on the vehicle's two-degree-of-freedom dynamic model, taking into account the target path tracking error. The lateral tracking error e d , lateral tracking error change rate Vehicle heading angle error and the rate of change of vehicle heading angle error Satisfies the following formula:
[0097]
[0098] in, is the lateral acceleration error, is the lateral acceleration of the vehicle, is the vehicle yaw angular velocity, is the reference heading angular velocity, is the vehicle yaw angular acceleration, is the yaw angular acceleration error;
[0099] The vehicle lateral tracking error e d , lateral tracking error change rate Vehicle heading angle error and the rate of change of vehicle heading angle error The vehicle tracking error dynamics model is the state information, and the calculation formula is as follows:
[0100]
[0101] in, is the state space equation of the vehicle tracking error dynamics model; u = delta_f, is the front wheel angle; is the system state matrix; is the input matrix; is the interference matrix.
[0102] S13, according to the vehicle tracking error dynamics model, select is the state quantity, u(t)= is the control quantity, satisfying the state space equation:
[0103]
[0104] in, is the state vector, f[·] is the nonlinear vector function;
[0105] S14. According to step S13, the state space equation is obtained, and the working point (ξ(t), u(t)) of the vehicle tracking error dynamics model is linearized at time t to obtain a linear time-varying equation, which is as follows:
[0106]
[0107] in, is the current state vector, A d (t) is the state Jacobian matrix, B d (t) is the input Jacobian matrix, x(t) is the original state vector;
[0108] S15. Discretize the linear time-varying equation obtained in step S14 using the forward Euler method. The discretized equation at time k is:
[0109]
[0110] Among them, A d (k)=I+T′A(t), A d (k) is the discrete-time state matrix, A(t) is the continuous-time state matrix, and B d (k) = T′B(t), B d (k) is the discrete-time input matrix, B(t) is the continuous-time input matrix, is the control increment, T′ is the control step size, and I is the identity matrix;
[0111] S16: Establish a two-degree-of-freedom linear time-varying discrete prediction model for the vehicle based on the discretization equation of step S15, based on ζ(k+1|t)=[ξ(k|t)u(k-1|t)] T , the state space equation of the prediction model is obtained as:
[0112]
[0113] in, ζ(k|t) is the extended state vector, To expand the discrete state matrix, is the expanded discrete input matrix, Δu(k|t) is the control increment, η(k|t) is the predicted input vector, is the expanded output matrix, A k,t is the state matrix of the original linearized system, B k,t is the input matrix, I m is the m-dimensional unit matrix, n is the state dimension, and m is the control dimension;
[0114] Assumptions The prediction interval is N p , the control interval is N c , and get the prediction model equation:
[0115] Y(t)=ψζ(t|t)+ΘΔU(t);
[0116] in, Predict output values for the system at multiple future moments, To reflect the impact of state variables on future output, is the future control increment, ΔU(t) is the dynamic response to the future output, η(t+f|t) is the system predicted output based on time t in the f-th step, for f-th power, used to recursively predict the state of the next f steps, f = 1, 2, ..., N p , is the input matrix treated as a constant.
[0117] S17. Design a cost function according to step S16, taking into account the error size and the size of the control amount in the path tracking process, and achieve a good path tracking effect with the minimum control amount. In addition, introduce the weight coefficient and relaxation factor to design the cost function J, and perform quadratic programming to solve the cost function. The specific formula is as follows:
[0118] minJ=min[ΔU(t) T ,ε]H t [ΔU(t) T ,ε] T +G t [ΔU(t) T ,ε];
[0119] Where ε is the relaxation factor, T is the matrix transpose, is the Hessian matrix, Q is the weighted matrix of the state variables, θ is the rolling prediction matrix, ρ is the proportional coefficient, G t =[2E(t) T Qθ t ] is the coefficient vector of the first-order term in the cost function, E(t)=ψ(t)ξ(t)-Y ref (t) is the initial error between the predicted output and the reference trajectory, ψ(·) is the state input mapping matrix, ξ(·) is the error state vector, and Y ref (·) is the reference trajectory output sequence, θ t To control the gain;
[0120] The optimal sequence obtained in the control time domain is: ΔU t =[Δu(t),Δu(t+1),...,Δu(t+N c -1)] T , where Δu(t) is the current control increment.
[0121] S2. Design a tire model based on the magic tire calculation formula. The formula is as follows:
[0122] F y =-D*sin(C*atan(B*xE*atan(B*x-atan(B*x))))+Sv;
[0123] Among them, F y is the tire lateral force, D is the peak factor, B is the stiffness factor, E is the curvature factor, C is the curve shape factor, Sv is the vertical drift of the curve, and x is the tire side slip angle.
[0124] S3, standardize the observation space and action space. The observation space includes the road adhesion coefficient μ, the tire side slip angle α, and the vertical load F on the tire. z , vehicle lateral acceleration a y The upper limit of the observation range is [inf; inf; inf; inf; inf], and the lower limit of the observation range is [-inf; -inf; -inf; -inf; -inf]; the action space is the five secondary parameters of the magic tire formula, namely C, B, D, E, and Sv, the upper limit of the action output range is [2; 0.9; 13000; -0.21; 1], and the lower limit of the action output range is [1; 0.19; 1500; -0.27; -19].
[0125] S4. Define a reinforcement learning reward function for each stage. By calculating the absolute value of the difference between the actual and reference tire forces, the agent is guided to continuously learn and optimize. Four thresholds are set to ensure that the difference between the tire force calculated by the magic tire formula and the reference tire force is within 10% at time t and at any longitudinal speed. The specific reward function logic is as follows:
[0126]
[0127] Among them, reward is the reward value, force is the current tire force, and err is the absolute value of the error between the tire forces;
[0128] An error termination function is designed to prevent the agent from exceeding the environment boundary due to excessive error during training, thereby improving the efficiency of reinforcement learning. The reinforcement learning termination conditions are designed as follows:
[0129]
[0130] Among them, 1 means stop training, 0 means continue training, and err_rate represents the relative error;
[0131] Design a vehicle initialization function to ensure that the agent is in a different initial state in each round of training. Use the initial speed of the initialized vehicle, calculated as follows:
[0132] initialSpeed=30+rand*100;
[0133] Among them, initialspeed is the initial speed, and rand is a random number in [0,1].
[0134] S5. Construct Actor-Critic neural network based on state space and action space; Figure 2As shown in the figure, the Critic network structure includes an input layer, a hidden layer, and an output layer. The input layer includes an obsInLyr layer and an actInlyr layer, which are used to process observation information and action information respectively. The dimensions of both layers are 5. The hidden layer includes an add layer and multiple fully connected layers and ReLu layers. The add layer adds the outputs of the observation path and the action path, so that the network can consider the influence of observation information and action information at the same time. Each fully connected layer contains 80 neurons, each of which is connected to all neurons in the previous layer, and the weights of the connections between neurons are continuously adjusted during the training process. QValLyr is the output layer, which represents the expected future reward of performing an action in a given state. The learning rate of the Critic network is 10 -3 .
[0135] like Figure 3 As shown in the figure, the Actor network structure includes an input layer, a hidden layer, and an output layer. The number of neurons in the fully connected layer in the hidden layer is 100. The output layer includes a fully connected layer, a tanhLayer activation layer, and a scalingLayer scaling layer. In order to match the network output dimension with the action space dimension, the number of neurons is the action output dimension. In order to enable the agent to fully explore the action output range, the tanhLayer layer is used to limit the output between [-1, 1], and the output is scaled and offset by the scaling layer. Among them, the Actor network learning rate is 10 -4 .
[0136] S6, such as Figure 4 As shown in Figure 2, we establish a deep reinforcement learning agent update and iteration mechanism. The specific steps are as follows:
[0137] S61. Using the vehicle initialization function, the training vehicle is placed in a different state each time;
[0138] S62: Update the agent. The steps are as follows:
[0139] S621, initialize the critic network Q (S, A; φ), and initialize the target critic network parameter φ t =φ, where S is the state space, A is the action space, φ is the Critic network parameter, t is the target critic network parameter;
[0140] S622, initialize the Actor network π(s;θ) and initialize the target Actor network parameters θ t =θ, where s is the current observation state;
[0141] S623, initializing the experience buffer R;
[0142] S624, initializing random detection noise N;
[0143] S625: Select action A based on the current state and the detected noise N t =π(S t ;θ)+N t , where A t is the action at time t, S t is the observed state at time t, N t is the detection noise at time t;
[0144] S626: Execute action A t , observed reward value r t and the next observation state S t+1 ;
[0145] S627: Storage transition (S t ,A t ,r t ,S t+1 ) to the experience buffer R;
[0146] S628: Randomly extract K experiences (S i ,A i ,r i ,S i+1 ), randomly sampled by MiniBatchSize, where S i is the state in the i-th experience, A i is the action in the i-th experience, r i is the reward in the i-th experience, S i+1 is the next state in the i-th experience, i is a positive integer;
[0147] S629: Set value function y i =r i +γQ t (S i+1 ,π t (S i+1 θ t );φ t ), where y i is the TD target value, γ is the discount factor, π t (·) is the current Actor strategy network, Q t (·) is the current Critic valuation network;
[0148] S6210: Update the critic network using the minimum loss function. The loss function is as follows:
[0149]
[0150] Among them, K is the MiniBatcSize size, and M is the number of samples involved in gradient calculation;
[0151] S6211: Update Actor parameters using a strategy to maximize expected discounted rewards:
[0152]
[0153] Where A=π(S i ;θ), is the gradient operator with respect to the parameter θ, is the gradient operator for action A;
[0154] S6212: Soft-update the target network parameters using the following smoothing factor τ:
[0155]
[0156] S7. Build and train a high-fidelity Carsim / Simulink combined deep reinforcement learning model to verify tire force estimation accuracy. First, initialize the vehicle speed in step S61, ensuring each training session is in a different vehicle state. Next, use the MPC controller in step S1 to track the target path and obtain the front wheel angle, which serves as input to Carsim. Carsim then outputs the vehicle data required for deep reinforcement learning. Finally, the deep reinforcement learning agent outputs action parameters to the Magic Tire model in step S2 to calculate tire forces. This cycle continues to complete deep reinforcement learning training.
[0157] To verify the online tire force estimation method based on deep reinforcement learning designed in the present invention, tire force estimation under linear and nonlinear conditions was selected for accuracy verification. This example uses the Carsim-Simulink joint simulation platform to verify the accuracy of tire estimation. The verification is performed by comparing the reference tire force output by Carsim with the estimated tire force. The experiment uses medium-speed continuous lane change conditions and high-speed double lane change conditions to test the effect of deep reinforcement learning.
[0158] The vehicle model parameters and MPC controller parameters are shown in Table 1 and Table 2:
[0159] Table 1 Vehicle parameters
[0160] Vehicle parameters Numerical unit Mass m 1413 kg Distance from center of mass to front axle a 1.015 m Distance from center of mass to rear axle b 1.895 m <![CDATA[Moment of inertia I zz > 1536.7 <![CDATA[kg·m 2 ]]> Gravity coefficient g 9.8 <![CDATA[kg·m 2 ]]>
[0161] Table 2MPC controller parameters
[0162] Parameter name Numerical Discrete time T 0.01s <![CDATA[Prediction horizon N c > 40 <![CDATA[Control time domain N p > 50
[0163] like Figure 5-8As shown, under the conditions of continuous lane change at 72 km / h and double lane change at 108 km / h, the tire force estimated by the inventive method is highly consistent with the reference tire force, and the vehicle lateral acceleration is stable, indicating that it can accurately estimate the tire force and effectively improve the vehicle driving stability.
[0164] It is worth noting that the contents not elaborated in detail in the present invention are all prior art and are well known to those skilled in the art.
[0165] Therefore, the present invention provides a method for online tire force estimation based on deep reinforcement learning. Through innovative technical means, it successfully overcomes the difficult problems of high cost of acquiring real vehicle data and high risk of data collection under extreme working conditions, while breaking through the limitations of traditional tire models. It not only solves the problem of narrow applicable working conditions for fixed tire parameters, but also greatly improves the calculation accuracy in the nonlinear region of the tire, bringing a new breakthrough in the field of vehicle tire force estimation.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the same. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solutions of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for online tire force estimation based on deep reinforcement learning, characterized in that: The following steps are involved: Step S1, designing an MPC controller, establishing a two-degree-of-freedom dynamic model of the vehicle and a two-degree-of-freedom linear time-varying discrete prediction model of the vehicle; Step S2: designing a tire model based on the magic tire calculation formula; Step S3: Standardize the observation space and action space, where the observation space includes the road adhesion coefficient, tire slip angle, vertical load, and lateral acceleration, and the action space is the parameters of the magic tire calculation formula; Step S4, defining the reinforcement learning stage reward function, error termination function and vehicle initialization function; Step S5: constructing an Actor-Critic neural network based on the state space and action space; Step S6: Establishing a deep reinforcement learning agent update and iteration mechanism; Step S7: Build and train a high-fidelity Carsim / Simulink joint deep reinforcement learning model to verify the accuracy of tire force estimation.
2. The method for online tire force estimation based on deep reinforcement learning according to claim 1, characterized in that: In step S1, the specific steps for designing the MPC controller are as follows: Step S11: Establish a two-degree-of-freedom dynamic model of the vehicle according to Newton's second law. The formula is as follows: Where m is the vehicle mass, v2 and v1 are the lateral and longitudinal velocities of the vehicle, K1 and K2 are the cornering stiffness of the front and rear axles, a and b are the distances from the center of mass of the vehicle to the front and rear axles, σ is the front wheel turning angle of the vehicle, and I zz is the moment of inertia, ω is the vehicle yaw rate, is the vehicle yaw angle angular acceleration; Step S12: Establish a vehicle lateral tracking error e d , lateral tracking error change rate Vehicle heading angle error and the rate of change of vehicle heading angle error The vehicle tracking error dynamics model for state information; Step S13: According to the vehicle tracking error dynamics model, select is the state variable, u(t)=σ is the control variable, satisfying the state space equation: in, is the state vector, f[·] is the nonlinear vector function; Step S14: According to step S13, the state space equation is obtained, and the working point (ξ(t), u(t)) of the vehicle tracking error dynamics model is linearized at time t to obtain a linear time-varying equation, which is as follows: in, is the current state vector, A d (t) is the state Jacobian matrix, B d (t) is the input Jacobian matrix, x(t) is the original state vector; Step S15: discretize the linear time-varying equation obtained in step S14 using the forward Euler method. The discretized equation at time k is: Among them, A d (k)=I+T′A(t), A d (k) is the discrete-time state matrix, A(t) is the continuous-time state matrix, and B d (k) = T′B(t), B d (k) is the discrete-time input matrix, B(t) is the continuous-time input matrix, is the control increment, T′ is the control step size, and I is the identity matrix; Step S16: establishing a two-degree-of-freedom linear time-varying discrete prediction model for the vehicle according to the discretization equation of step S15; Step S17: Design a cost function according to step S16, taking into account the error size and the size of the control amount in the path tracking process, achieving a good path tracking effect with the minimum control amount, and introducing weighted coefficients and relaxation factors to design the cost function J, and perform quadratic programming solution.
3. The method for online tire force estimation based on deep reinforcement learning according to claim 2, characterized in that: In step S12, the calculation formula of the vehicle tracking error dynamics model is as follows: in, is the state space equation of the vehicle tracking error dynamics model; u = delta_f, is the front wheel angle; is the system state matrix; is the input matrix; is the interference matrix, is the reference heading angular velocity.
4. The method for online tire force estimation based on deep reinforcement learning according to claim 2, characterized in that: In step S16, based on ζ(k+1|t)=[ξ(k|t)u(k-1|t)] T , the state space equation of the prediction model is obtained as: in, ζ(k|t) is the extended state vector, To expand the discrete state matrix, is the expanded discrete input matrix, Δu(k|t) is the control increment, η(k|t) is the predicted input vector, is the expanded output matrix, A k,t is the state matrix of the original linearized system, B k,t is the input matrix, I m is the m-dimensional unit matrix, n is the state dimension, and m is the control dimension; Assumptions The prediction interval is N p , the control interval is N c , and get the prediction model equation: Y(t)=ψζ(t|t)+ΘΔU(t); in, Predict output values for the system at multiple future moments, To reflect the impact of state variables on future output, is the future control increment, ΔU(t) is the dynamic response to the future output, η(t+f|t) is the system predicted output based on time t in the f-th step, for f-th power, used to recursively predict the state of the next f steps, f = 1, 2, ..., N p , is the input matrix treated as a constant.
5. The method for online tire force estimation based on deep reinforcement learning according to claim 2, characterized in that: In step S17, the cost function J is designed and solved by quadratic programming. The specific formula is as follows: min J=min[ΔU(t) T ,e]H t [ΔU(t) T ,e] T +G t [ΔU(t) T ,e]; Where ε is the relaxation factor, T is the matrix transpose, is the Hessian matrix, Q is the weighted matrix of the state variables, θ is the rolling prediction matrix, ρ is the proportional coefficient, G t =[2E(t) T Qθ t ] is the coefficient vector of the first-order term in the cost function, E(t)=ψ(t)ξ(t)-Y ref (t) is the initial error between the predicted output and the reference trajectory, ψ(·) is the state input mapping matrix, ξ(·) is the error state vector, and Y ref (·) is the reference trajectory output sequence, θ t To control the gain; The optimal sequence obtained in the control time domain is: ΔU t =[Δu(t),Δu(t+1),...,Δu(t+N c -1)] T , where Δu(t) is the current control increment.
6. The method for online tire force estimation based on deep reinforcement learning according to claim 1, characterized in that: In step S2, a tire model is designed based on the magic tire calculation formula, which is as follows: F y =-D*sin(C*atan(B*xE*atan(B*x-atan(B*x))))+Sv; Among them, F y is the tire lateral force, D is the peak factor, B is the stiffness factor, E is the curvature factor, C is the curve shape factor, Sv is the vertical drift of the curve, and x is the tire side slip angle.
7. The method for online tire force estimation based on deep reinforcement learning according to claim 1, characterized in that: In step S3, the observation space includes the road adhesion coefficient μ, the tire side slip angle α, the vertical load F z , vehicle lateral acceleration a y ,The action space is the five secondary parameters of the magic tire formula, which are C, B, D, E, and Sv.
8. The method for online tire force estimation based on deep reinforcement learning according to claim 1, characterized in that: In step S4, the calculation formula of the stage-by-stage reward function is as follows: Among them, reward is the reward value, force is the current tire force, and err is the absolute value of the error between the tire forces; The calculation formula of the error termination function is as follows: Among them, 1 means stop training, 0 means continue training, and err_rate represents the relative error; The calculation formula of the vehicle initialization function is as follows: initialSpeed=30+rand*100; Among them, initialspeed is the initial speed, and rand is a random number in [0,1].
9. The method for online tire force estimation based on deep reinforcement learning according to claim 1, characterized in that: In step S5, the Critic network structure includes an input layer, a hidden layer, and an output layer. The input layer includes an obsInLyr layer and an actInlyr layer, both of which have a dimension of 5. The hidden layer includes an add layer and multiple fully connected layers and a ReLu layer. Each fully connected layer contains 80 neurons, and QValLyr is the output layer. The Actor network structure includes an input layer, a hidden layer, and an output layer. The number of fully connected layer neurons in the hidden layer is 100, and the output layer includes a fully connected layer, a tanhLayer activation layer, and a scalingLayer scaling layer.
10. The method for online tire force estimation based on deep reinforcement learning according to claim 1, characterized in that: In step S6, a deep reinforcement learning agent update and iteration mechanism is established. The specific steps are as follows: Step S61: Through the vehicle initialization function, each training vehicle is put into a different state; Step S62: Update the agent. The steps are as follows: Step S621: Initialize the critic network Q(S,A;φ) and initialize the target critic network parameter φ t =φ, where S is the state space, A is the action space, φ is the Critic network parameter, t is the target critic network parameter; Step S622: Initialize the Actor network π(s;θ) and initialize the target Actor network parameters θ t =θ, where s is the current observation state; Step S623: Initialize the experience buffer R; Step S624, initializing random detection noise N; Step S625: Select action A based on the current state and the detected noise N t =π(S t ;θ)+N t , where A t is the action at time t, S t is the observed state at time t, N t is the detection noise at time t; Step S626: Execute action A t , observed reward value r t and the next observation state S t+1 ; Step S627: Store transition (S t ,A t ,r t ,S t+1 ) to the experience buffer R; Step S628: Randomly extract K experiences (S i ,A i ,r i ,S i+1 ), randomly sampled by MiniBatchSize, where S i is the state in the i-th experience, A i is the action in the i-th experience, r i is the reward in the i-th experience, S i+1 is the next state in the i-th experience, i is a positive integer; Step S629: Set the value function y i =r i +γQ t (S i+1 ,π t (S i+1 θ t );φ t ), where y i is the TD target value, γ is the discount factor, π t (·) is the current Actor strategy network, Q t (·) is the current Critic valuation network; Step S6210: Update the critic network using the minimum loss function. The loss function is as follows: Among them, K is the MiniBatcSize size, and M is the number of samples involved in gradient calculation; Step S6211: Update the Actor parameters using the strategy to maximize the expected discounted reward: Where A=π(S i ;θ), is the gradient operator with respect to the parameter θ, is the gradient operator for action A; Step S6212: Soft-update the target network parameters using the following smoothing factor τ:
Citation Information
Cited By
Vehicle tire force estimation method and device
CN122275910A