Vehicle roll stability control method based on deep reinforcement learning

Through the sliding mode adaptive active suspension control method based on deep reinforcement learning, the roll stability problem during vehicle high-speed steering is solved, the vehicle roll stability domain is widened and the occupant comfort is improved, and the traditional method is insufficient robustness under nonlinearity and disturbance is solved.

CN120245659APending Publication Date: 2025-07-04HEFEI UNIV OF TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510687562.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing vehicles are prone to excessive rolling or even rolling over when steering at high speeds, and traditional control methods are insufficient in robustness and adaptability under system nonlinearity and crosswind disturbance, resulting in poor control performance.

Method used

Adoptional active suspension control method of sliding mode based on deep reinforcement learning is adopted. By establishing a six-degree of freedom steering-roll vehicle dynamic model, the target body roll angle is calculated, and the sliding mode controller and intelligent body are designed. The sliding mode parameters are optimized in real time to output active roll moment, and the body is driven to track the target roll angle.

Benefits of technology

It significantly broadens the vehicle roll stability domain, reduces roll risk, improves occupants' lateral acceleration perception and driving comfort, and maintains accurate tracking of the target roll angle under complex disturbances, reduces jitter phenomenon, and improves system reliability and life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120245659A_ABST
    Figure CN120245659A_ABST
Patent Text Reader

Abstract

The invention discloses a sliding-mode adaptive vehicle roll stability control method based on deep reinforcement learning, which is applied to an active suspension vehicle and comprises the following steps: establishing a six-degree-of-freedom steering-roll vehicle dynamics model by combining a Doguff tire model, and calculating a target vehicle body roll angle based on a reverse roll mechanism. A sliding mode controller comprising a linear sliding mode surface and an exponential approaching rate is designed, an intelligent agent based on a soft behavior strategy gradient algorithm (SAC) is constructed, and the intelligent agent adaptively outputs optimal sliding mode parameters by sensing the vehicle state and the roll angle tracking condition. And finally, the sliding mode controller combines the optimal sliding mode parameters and the vehicle state to calculate the active roll-over moment and distributes the active roll-over moment to the main power of the left and right side suspensions, so that the precise regulation and control of the vehicle body posture are realized. According to the method, the roll stability domain during high-speed steering can be remarkably expanded, the lateral sensing acceleration of passengers is reduced, the high-robustness tracking effect is kept under complex disturbance, and meanwhile sliding mode buffeting is restrained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of vehicle control, and particularly relates to a method for controlling vehicle roll stability using an active suspension. Background Art

[0002] When a vehicle steers at high speed or makes an emergency turn, it often causes excessive body roll, and even leads to vehicle rollover, seriously endangering driving safety and reducing ride comfort. Compared with passive suspensions and semi-active suspensions, an active suspension has the ability to actively output driving force to adjust the body attitude, and can significantly improve the roll safety domain of the vehicle under extreme driving conditions. Existing roll control methods often focus on passive roll suppression or zero-roll control, and the system nonlinearity and crosswind disturbance during vehicle roll make the robustness and adaptability of traditional control methods significantly decline, resulting in poor control efficiency. Summary of the Invention

[0003] In order to overcome the deficiencies of the prior art, the present invention provides a method for controlling vehicle roll stability of a sliding mode adaptive active suspension based on deep reinforcement learning, aiming to achieve precise tracking of the target reverse roll angle, thereby significantly improving the roll stability domain and lateral comfort of the vehicle during high-speed steering.

[0004] The present invention adopts the following technical solutions to achieve the above invention objectives:

[0005] A method for controlling vehicle roll stability of a sliding mode adaptive based on deep reinforcement learning according to the present invention is characterized in that it is used in an active suspension vehicle and is carried out according to the following steps:

[0006] Step 1: Based on the Doguff tire model, establish a six-degree-of-freedom steering-roll vehicle dynamics model to describe the vertical vibration of the sprung and unsprung masses and the roll, yaw, and lateral motion of the vehicle;

[0007] Step 2: Calculate the target body roll angle based on the six-degree-of-freedom steering-roll vehicle dynamics model and the reverse roll mechanism;

[0008] Step 3: Design a sliding mode controller and calculate the output active roll moment to track the target body roll angle;

[0009] Step 4: Construct an agent based on the soft actor-critic algorithm, and design its state space, action space, reward function, and network structure and train to obtain the optimal agent, which is used to output the optimal sliding mode parameters according to the vehicle state and roll angle tracking situation;

[0010] Step 5: The sliding mode controller outputs the active roll moment according to the optimal sliding mode parameters and the vehicle state, and calculates the active forces of the left and right suspensions to drive the body to track the target body roll angle, thereby expanding the roll stability domain of the vehicle during high-speed steering.

[0011] The characteristics of a sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning according to the present invention also lie in that the step 2 includes:

[0012] Step 2.1: When the component of the vehicle body's gravitational moment is equal in magnitude and opposite in direction to the centrifugal moment, calculate the ideal body reverse roll angle according to Equation (1) :

[0013] (1)

[0014] In Equation (1), is the front wheel steering angle of the vehicle, is the longitudinal speed of the vehicle, 、 are the distances from the vehicle's center of mass to the front axle and the rear axle respectively, is the gravitational acceleration;

[0015] Step 2.2: Calculate the target body reverse roll angle according to Equation (2) :

[0016] (2)

[0017] In Equation (2), is the proportionality factor, is the set limit value of the reverse roll angle.

[0018] Furthermore, the step 3 includes:

[0019] Step 3.1: Design the sliding mode surface of the sliding mode controller according to Equation (3) :

[0020] (3)

[0021] In Equation (3), is the gain of the sliding mode surface, represents the roll angle tracking error, and , is the current body roll angle;

[0022] Step 3.2: Take the derivative of (3) to obtain Equation (4):

[0023] (4)

[0024] In Equation (26), 、 are the first derivative and the second derivative of the target roll angle with respect to time respectively, is the body roll angle speed; 、 respectively represent the first derivative and the second derivative of the roll angle tracking error with respect to time, represents the vehicle roll angle speed, represents the vehicle roll angle acceleration, and is obtained from Equation (5):

[0025] (5)

[0026] In Equation (5), is the total vehicle mass, is the body mass, is the distance from the roll center to the vehicle center of mass, is the half track width, is the yaw rate, is the body roll moment of inertia, 、 are the passive forces of the left and right suspensions respectively, 、 are the lateral forces of the left and right wheels respectively, is the active roll moment;

[0027] Step 3.3: Design a sliding mode controller using an exponential reaching law as in Equation (6):

[0028] (6)

[0029] In Equation (6), is the exponential gain, is the switching gain, and sgn() is the sign function.

[0030] Step 3.4: Combine Equation (4) and Equation (6) to calculate the control output of the sliding mode controller, i.e., the active roll moment :

[0031] (7).

[0032] Furthermore, the said Step 4 includes:

[0033] Step 4.1: Select the state space where, is the roll angle tracking error, is the change rate of the roll angle tracking error, is the vehicle lateral acceleration;

[0034] Step 4.2: Select the action space where, 、 、 are three sliding mode control parameters;

[0035] Step 4.3: Define the total reward function ;

[0036] Step 4.4: Design the network structure of the agent, including: Critic1 network, Critic2 network, target Critic1 network, target Critic2 network, and Actor network; among them, the network structures of the Critic1 network, Critic2 network, target Critic1 network, and target Critic2 network are exactly the same;

[0037] The Critic1 network, Critic2 network, target Critic1 network, and target Critic2 network all include: a first fully connected layer, a second fully connected layer, and an output layer connected in sequence; the output layer generates the value of the state-action pair;

[0038] The Actor network includes: a third fully connected layer, a fourth fully connected layer, and an action output layer connected in sequence, and the action output layer includes: a mean branch, a variance branch, and a Gaussian distribution sampling layer;

[0039] Among them, the mean branch includes: a first sub-fully connected layer and a mean output layer connected in sequence, the input end of the first sub-fully connected layer is connected to the output end of the fourth fully connected layer, and the mean output layer is used to output the first vector of the means of three sliding mode parameters;

[0040] The variance branch includes: a second sub-fully connected layer and a variance output layer connected in sequence, the input end of the second sub-fully connected layer is connected to the output end of the fourth fully connected layer, and the variance output layer is used to output the second vector of the variances of three sliding mode parameters;

[0041] The Gaussian distribution sampling layer constructs a Gaussian distribution using the first vector and the second vector of the active force and generates three sliding mode parameters.

[0042] Step 4.5: Train the agent under the input of a random front wheel steering angle excitation until the total reward converges to a high reward level, so as to obtain the optimal agent for offline optimization of the sliding mode parameters.

[0043] Furthermore, the step 4.3 includes:

[0044] Step 4.3.1: Construct the roll angle tracking error penalty term using Equation (8) :

[0045] (8)

[0046] In Equation (8), is the set limit value of the roll angle tracking accuracy;

[0047] Step 4.3.2: Construct a penalty term for stable tracking and chattering suppression using Equation (9) :

[0048] (9)

[0049] Step 4.3.3: Construct a penalty term for the sudden change of the active suspension actuation force using Equation (10) :

[0050] (10)

[0051] In Equation (10), is the set limit value for the sudden change of the active force;

[0052] Step 4.3.4: Construct the total reward function using Equation (11) :

[0053] (11)

[0054] In Equation (11), 、 、 are three weight coefficients.

[0055] Furthermore, the said Step 4.5 includes:

[0056] Step 4.5.1: Randomly initialize the parameters of the Critic1 network , the parameters of the Critic2 network ; and assign the parameters of the Critic1 network to the parameters of the first target Critic1 network , assign the parameters of the Critic2 network to the parameters of the second target Critic2 network ; randomly initialize the parameters of the Actor network ;

[0057] Step 4.5.2: Input the state at time into the Actor network, and output the action at time . After acting on the vehicle, obtain the state at time , and calculate the reward , at time . Thus, obtain a sample data and store it in the experience pool;

[0058] ​Step 4.5.3: Randomly extract a sample data from the experience pool , and and are input into the Critic1 network and the Critic2 network, and the first value of the state-action pair is output and the second value ;

[0059] Input into the Actor network, and output the action at time ;

[0060] Input and into the target Critic1 and target Critic2 networks, and output the first target value of the state-action pair and the second target value ;

[0061] Calculate according to Equation (12) the target value at time :

[0062] (12)

[0063] In Equation (11), is the discount factor;

[0064] Step 4.5.4: Respectively make the mean square error of with and to construct the Loss function, and thus use the gradient descent method to update the parameters of the Critic1 network , the parameters of the Critic2 network , and obtain the updated parameters of the Critic1 network and assign them to , the updated parameters of the Critic2 network and assign them to ;

[0065] Step 4.5.5: Use Equation (13) to update the parameters of the target Critic1 network and the of the target Critic2 network:

[0066] (13)

[0067] In Equation (13), is the smoothing factor, is the assignment symbol;

[0068] Step 4.5.6: According to , the parameters of the Actor network are updated using a delayed update mechanism to obtain the updated parameters of the Actor network and assign them to ;

[0069] Step 4.5.7: Iteratively train the agent according to the process of Step 4.5.2 - Step 4.5.6 until the total reward converges to a high reward level, thereby obtaining the optimal agent for offline deployment.

[0070] Further, the said Step 5 includes:

[0071] Step 5.1: Calculate the active roll moment according to Equation (14) :

[0072] ( 4)

[0073] In Equation (14), , , are three optimal sliding mode parameters output by the optimal agent according to the vehicle state;

[0074] Step 5.2: Calculate and according to Equation (15):

[0075] (15).

[0076] An electronic device of the present invention includes a memory and a processor, characterized in that the memory stores the program of the vehicle roll stability control method, and the processor implements active suspension reverse roll control when executing the program.

[0077] A computer-readable storage medium of the present invention stores a computer program, characterized in that the computer program implements the steps of the vehicle roll stability control method when executed by a processor.

[0078] Compared with the prior art, the beneficial effects of the present invention are reflected in:

[0079] 1. Based on the reverse roll mechanism, the present invention drives the vehicle body to tilt inward the bend through the active suspension, and uses the gravity moment to dynamically compensate the centrifugal moment, effectively suppressing the lateral load transfer and lateral acceleration during high-speed steering. Compared with the traditional zero-roll control method, the vehicle roll stability domain is significantly broadened, the rollover risk is reduced, and at the same time, the occupant's lateral acceleration perception is improved, and the driving comfort is enhanced.

[0080] 2. The present invention combines SAC deep reinforcement learning with sliding mode control. The SAC agent can dynamically optimize the sliding mode control parameters by real-time sensing of the system state, and can still maintain accurate tracking of the target roll angle under the conditions of strong nonlinear suspension dynamics, crosswind disturbance and model uncertainty, solving the problem of insufficient robustness of traditional sliding mode control under complex disturbances.

[0081] 3. The present invention introduces a control input smoothness constraint into the reinforcement learning reward mechanism. By suppressing the high-frequency switching characteristics of the sliding mode control, the mutation amplitude of the active suspension actuator force is reduced, the chattering phenomenon is effectively reduced, the mechanical loss of the actuator is avoided, and the system reliability and life are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 is the overall control strategy diagram of the present invention;

[0083] Figure 2 is the structure diagram of the Critic network of the SAC agent of the present invention;

[0084] Figure 3 is the structure diagram of the Actor network of the SAC agent of the present invention;

[0085] Figure 4 is the execution flow chart of the SAC algorithm of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] In this embodiment, a sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning is used for an active suspension vehicle, and the control framework is as Figure 1 shown, and includes the following steps:

[0087] Step 1: Establish a six-degree-of-freedom steering-roll vehicle dynamics model to describe the vertical vibrations of the sprung and unsprung masses and the roll, yaw and lateral motions of the vehicle, and select the Dugoff model to establish the coupling relationship between the vertical and lateral forces of the tires.

[0088] Step 1.1: Use Equation (1) to establish a six-degree-of-freedom vehicle dynamics model considering lateral, yaw, roll and vertical vibrations:

[0089] (1)

[0090] In Equation (12), is the body mass, is the total vehicle mass, , are the unsprung masses on the left and right sides respectively, , are the vehicle yaw moment of inertia and the body roll moment of inertia respectively, , They are the passive vertical forces of the left and right suspensions respectively, , They are the active forces of the left and right suspensions respectively, , They are the lateral forces of the front and rear wheels respectively, is the tire stiffness, , , , , They are the body displacement, the displacement of the unsprung mass on the left, the displacement of the unsprung mass on the right, the left road excitation input, and the right road excitation input respectively, , , They are the body acceleration, the acceleration of the unsprung mass on the left, and the acceleration of the unsprung mass on the right respectively, is the vehicle yaw angle, is the vehicle yaw angular acceleration, , They are the distances from the center of mass to the front and rear axles respectively, is the vehicle lateral acceleration, calculated according to Equation (2), is the body roll angle, is the body roll angular acceleration, is the half track, is the distance from the roll center to the vehicle center of mass, is the active roll moment.

[0091] (2)

[0092] In Equation (2), is the lateral displacement of the vehicle, is the longitudinal vehicle speed.

[0093] Step 1.2: Calculate the passive vertical forces of the suspension stiffness and damping on both sides according to Equation (3):

[0094] (3)

[0095] In Equation (3), is the stiffness coefficient of the suspension spring, is the damping coefficient of the suspension damper, , They are the displacements of the sprung mass on the left and right respectively, , They are the velocities of the sprung mass on the left and right respectively, , They are the velocities of the unsprung mass on the left and right respectively.

[0096] Step 1.3: Based on the small roll angle assumption, there is , the relationship between the sprung mass motion and the unsprung mass motion on both the left and right sides is expressed by Equation (4):

[0097] (4)

[0098] Step 1.4: Based on the assumption of the small front wheel angle, the front and rear wheel sideslip angles are calculated according to the bicycle steering model, as shown in Equation (5):

[0099] (5)

[0100] In Equation (5), and are the front and rear wheel tire sideslip angles respectively, and is the front wheel angle.

[0101] Step 1.5: Select the Dugoff tire model to characterize the coupling relationship between the tire lateral force and the vertical load, and calculate the tire lateral force according to Equation (6):

[0102] (6)

[0103] In Equation (6), is the road surface adhesion coefficient, is the tire slip ratio, is the tire sideslip angle, is the tire longitudinal stiffness, is the tire lateral stiffness, is the speed factor, is the tire vertical force.

[0104] Step 2: Based on the reverse roll mechanism, calculate the target body roll angle , and improve the vehicle roll stability and lateral comfort by actively tilting the body to generate a gravitational moment component opposite to the centrifugal moment.

[0105] Step 2.1: The centrifugal moment and the gravitational moment can be derived from the six-degree-of-freedom steering-roll vehicle dynamics model and are calculated according to Equation (7):

[0106] (7)

[0107] Step 2.2: When the vehicle is in steady-state driving, , calculate the ideal yaw acceleration response during vehicle steering according to Equation (8):

[0108] (8)

[0109] Step 2.3: Let = , combined with Equation (8), the ideal reverse roll angle calculation is obtained as shown in Equation (9):

[0110] (9)

[0111] Step 2.4: Considering the output limit of the active suspension, power consumption optimization, and the requirement for rapid adjustment of the roll angle, it is not necessary to track the ideal roll angle. The actual target roll angle to be tracked is calculated according to Equation (10):

[0112] (10)

[0113] In the formula, is the proportionality factor, is the set limit value of the active roll angle. In this example, , .

[0114] Step 3: Design a sliding mode controller. Its sliding mode surface is a linear combination of the roll angle tracking error and its change rate, and an exponential reaching law is adopted;

[0115] Step 3.1: Design the sliding mode surface of the sliding mode controller according to Equation (11) :

[0116] (11)

[0117] In Equation (3), is the gain of the sliding mode surface, represents the roll angle tracking error, ;

[0118] Step 3.2: Differentiate Equation (11) to obtain Equation (12):

[0119] (12)

[0120] In Equation (12), , are the first-order derivative and second-order derivative of the target roll angle with respect to time respectively, , represent the first-order derivative and second-order derivative of the roll angle tracking error with respect to time respectively. The vehicle roll angle acceleration is obtained from Equation (13):

[0121] (13)

[0122] Step 3.3: The sliding mode controller adopts an exponential reaching law as shown in Equation (14), which can respond quickly when the state deviates from the sliding mode surface and is small, enabling the state vector to move smoothly to the sliding surface, thereby reducing chattering:

[0123] (14)

[0124] In Equation (14), is the exponential gain, is the switching gain, and sgn() is the sign function, whose analytical formula is shown in Equation (15):

[0125] (15)

[0126] Step 3.4: By combining Equation (12) and Equation (14), the control output of the sliding mode controller, i.e., the active roll moment, can be solved as and is calculated according to Equation (16):

[0127] (16)

[0128] Step 4. Construct an agent based on Soft Actor-Critic (SAC), and design the state space, action space, reward function, and network structure of the SAC agent;

[0129] Step 4.1: Select the state space , where is the roll angle tracking error, is the change rate of the roll angle tracking error, is the lateral acceleration of the vehicle;

[0130] Step 4.2: Select the action space , , , are the three sliding mode control parameters.

[0131] Step 4.3: Define the reward function:

[0132] Use Equation (17) to construct the roll angle tracking error penalty term :

[0133] (17)

[0134] In Equation (17), is the selected limit value of the roll angle tracking accuracy. In this example, is taken.

[0135] When chattering occurs, the state vector shuttles back and forth on the sliding surface. When the roll angle approaches the target value smoothly, the directions of the roll angle error and the error change rate are opposite, i.e., < 0. Therefore, according to Equation (18), a penalty term for stable tracking and chattering suppression is constructed :

[0136] (18)

[0137] Considering the time lag of the actuator's power response, the sudden change of the power should be minimized as much as possible. The penalty term for the sudden change of the active suspension power is constructed using Equation (19):

[0138] (19)

[0139] In Equation (33), is the selected limit value of the active force mutation. In this example, 100 N is taken.

[0140] The total reward function is calculated as Equation (20):

[0141] (20)

[0142] In Equation (20), 、 、 are three weight coefficients. In this example, 1, 0.5, and 1 are taken respectively.

[0143] Step 4.4: Design the network structure of the agent, including: Critic1 network, Critic2 network, target Critic1 network, target Critic2 network, and Actor network; among them, the network structures of the Critic1 network, Critic2 network, target Critic1 network, and target Critic2 network are exactly the same;

[0144] The Critic1 network, Critic2 network, target Critic1 network, and target Critic2 network all include: a first fully connected layer, a second fully connected layer, and an output layer connected in sequence; the output layer generates the value of the state-action pair;

[0145] The Actor network includes: a third fully connected layer, a fourth fully connected layer, and an action output layer. The action output layer includes:

[0146] A mean branch, including a first sub-fully connected layer and a mean output layer connected in sequence. The input end of the first sub-fully connected layer is connected to the output end of the fourth fully connected layer, and the mean output layer is used to output the first vector of the means of the three sliding mode parameters;

[0147] A variance branch, including a second sub-fully connected layer and a variance output layer connected in sequence. The input end of the second sub-fully connected layer is connected to the output end of the fourth fully connected layer, and the variance output layer is used to output the second vector of the variances of the three sliding mode parameters;

[0148] Gaussian distribution sampling layer, which constructs a Gaussian distribution using the first vector and the second vector of the active force and generates three sliding mode parameters.

[0149] In this example, the structures of the Critic and Actor networks are shown in Figure 2 and Figure 3 . The input of the Critic network is 4 state variables and 3 action variables, and then the Q value is output through two hidden layers with 400 nodes and 300 nodes. The input of the Actor network is 4 state variables. After passing through two hidden layers with 200 and 200 nodes respectively, two branches are led out. Each passes through a hidden layer with 32 nodes to output the mean and standard deviation of the action. Finally, the sliding mode parameters are generated by Gaussian sampling from the mean and standard deviation.

[0150] Step 4.5: Train the agent under the input of random front wheel angle excitation until the total reward converges to a high reward level, so as to obtain the optimal agent for offline optimization of the sliding mode parameters. Its training framework is shown in Figure 4 .

[0151] Step 4.5.1: Randomly initialize the parameters of the Critic1 network , the parameters of the Critic2 network ; and assign the parameters of the Critic1 network to the parameters of the first target Critic1 network , assign the parameters of the Critic2 network to the parameters of the second target Critic2 network ; randomly initialize the parameters of the Actor network ;

[0152] Step 4.5.2: Input the state at time into the Actor network and output the action at time . Then, apply it to the vehicle and obtain the state at time . Calculate the reward , at time to obtain the reward , thus obtaining a sample data and storing it in the experience pool.

[0153] Step 5.3: Randomly extract a sample data from the experience pool, and input and into the Critic1 network and the Critic2 network, and output the first value and the second value of the state-action pair ;

[0154] Input it into the Actor network and output the action at ; time ;

[0155] Input and into the target Critic1 and target Critic2 networks, and output the first target value of the state-action pair value and the second target value ;

[0156] Calculate according to Equation (21) the target value at :

[0157] (21)

[0158] In Equation (21), is the discount factor, which is taken as 0.99 in this example.

[0159] Step 4.5.4: Take the mean square error of and and respectively to construct the Loss function, and then use the gradient descent method to update the parameters of the Critic1 network , the parameters of the Critic2 network , and obtain the updated parameters of the Critic1 network and assign them to , the updated parameters of the Critic2 network and assign them to ;

[0160] Step 4.5.5: Use Equation (13) to update the parameters of the target Critic1 network and the of the target Critic2 network:

[0161] (22)

[0162] In Equation (22), is the smoothing factor, which is taken as 0.001 in this example, is the assignment symbol.

[0163] Step 4.5.6: The Actor network adopts a delayed update mechanism, that is, after the Critic network is updated times, then according to use the policy gradient ascent method to update the parameters of the Actor network , and obtain the updated parameters of the Actor network and assign them to ;

[0164] Step 4.5.7: Iteratively train the agent according to the process of steps 4.5.2 - 4.5.6 until the total reward converges to a high reward level, thereby obtaining the optimal agent for offline deployment.

[0165] Step 5: The sliding mode controller calculates the active roll moment based on the sliding mode parameters output by the SAC agent and the vehicle state, and distributes the driving forces of the left and right active suspensions to drive the vehicle body to track the target roll angle, realizing the roll stability control under high-speed steering.

[0166] Step 5.1: The sliding mode controller calculates the active roll moment according to the optimal sliding mode parameters , , output by the SAC agent and the vehicle state , and the calculation method is shown in Equation (23):

[0167] (23)

[0168] Step 5.2: Since road surface fluctuations are not considered, it is optimal to equally and reversely distribute the active roll moment to the left and right active suspensions, which does not affect the vertical movement of the vehicle body and has the minimum power consumption. Therefore, the driving forces of the active suspensions are calculated according to Equation (24) and applied to the vehicle:

[0169] (24)

[0170] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0171] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.

Claims

1. A sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning, characterized in that, It is used in an active suspension vehicle and is carried out according to the following steps: Step 1: Based on the Doguff tire model, establish a six-degree-of-freedom steering-roll vehicle dynamics model to describe the vertical vibration of the sprung and unsprung masses and the roll, yaw, and lateral motion of the vehicle; Step 2: Calculate the target body roll angle based on the six-degree-of-freedom steering-roll vehicle dynamics model and the reverse roll mechanism; Step 3: Design a sliding mode controller and calculate the output active roll moment to track the target body roll angle; Step 4: Construct an agent based on the soft actor-critic algorithm, design its state space, action space, reward function, and network structure, and train to obtain the optimal agent, which is used to output the optimal sliding mode parameters according to the vehicle state and the roll angle tracking situation; Step 5: The sliding mode controller outputs the active roll moment according to the optimal sliding mode parameters and the vehicle state, and calculates the active forces on the left and right suspensions to drive the body to track the target body roll angle, thereby expanding the roll stability domain of the vehicle during high-speed steering.

2. The sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning according to claim 1, characterized in that, The said Step 2 includes: Step 2.1: When the component of the vehicle body's gravity moment is equal in magnitude and opposite in direction to the centrifugal moment, calculate the ideal reverse roll angle of the vehicle body according to Equation (1). : (1) In Equation (1), is the front wheel steering angle of the vehicle, is the longitudinal speed of the vehicle, , are the distances from the vehicle's center of mass to the front axle and the rear axle respectively, is the acceleration due to gravity; Step 2.2: Calculate the reverse roll angle of the target vehicle body according to Equation (2) :[[]]END]] (2) In Equation (2), is the scale factor, is the set reverse roll angle limit.

3. A sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning according to claim 1, characterized in that, The said Step 3 includes: Step 3.1: Design the sliding mode surface of the sliding mode controller according to Equation (3) :[[]]END]] (3) In Equation (3), is the gain of the sliding mode surface, represents the roll angle tracking error, and , is the current vehicle body roll angle; Step 3.2: Take the derivative of (3) to obtain Equation (4): (4) In Equation (26), and are the first derivative and second derivative of the target roll angle with respect to time, respectively, and is the vehicle roll angle speed; and represent the first derivative and second derivative of the roll angle tracking error with respect to time, respectively, represents the vehicle roll angle speed, represents the vehicle roll angle acceleration, and is obtained from Equation (5): (5) In formula (5), is the vehicle mass, is the body mass, is the distance from the roll center to the vehicle center of mass, is the half track width, is the yaw rate, is the body roll moment of inertia, and are the driving forces of the left and right suspensions respectively, and are the lateral forces of the left and right wheels respectively, is the active roll moment; Step 3.3: Design a sliding mode controller using an exponential reaching law as shown in Equation (6): (6) In Equation (6), is the exponential gain, is the switching gain, and sgn() is the sign function; Step 3.4: Combine Equation (4) and Equation (6), and thus calculate the control output of the sliding mode controller, i.e., the active roll moment, according to Equation (7). : (7)。 4. A sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning according to claim 1, characterized in that, The said Step 4 includes: Step 4.1: Select the state space , where is the roll angle tracking error, is the change rate of the roll angle tracking error, is the lateral acceleration of the vehicle; Step 4.2: Select the action space , where , , are three sliding mode control parameters; Step 4.3: Define the total reward function ; Step 4.4: Design the network structure of the agent, including: Critic1 network, Critic2 network, target Critic1 network, target Critic2 network, and Actor network; among them, the network structures of the Critic1 network, Critic2 network, target Critic1 network, and target Critic2 network are all exactly the same; The Critic1 network, Critic2 network, target Critic1 network, and target Critic2 network all include: a first fully connected layer, a second fully connected layer, and an output layer connected in sequence; the output layer generates the value of the state-action pair; The Actor network includes: a third fully connected layer, a fourth fully connected layer, and an action output layer, and the action output layer includes: a mean branch, a variance branch, and a Gaussian distribution sampling layer; Among them, the mean branch includes: a first sub-fully connected layer and a mean output layer connected in sequence, the input end of the first sub-fully connected layer is connected to the output end of the fourth fully connected layer, and the mean output layer is used to output the first vector of the means of the three sliding mode parameters; The variance branch includes: a second sub-fully connected layer and a variance output layer connected in sequence, the input end of the second sub-fully connected layer is connected to the output end of the fourth fully connected layer, and the variance output layer is used to output the second vector of the variances of the three sliding mode parameters; The Gaussian distribution sampling layer constructs a Gaussian distribution using the first vector and the second vector of the active force and generates three sliding mode parameters; Step 4.5: Train the agent under a random front wheel angle excitation input until the total reward converges to a high reward level, thereby obtaining the optimal agent for offline optimization of the sliding mode parameters.

5. A sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning according to claim 4, characterized in that, The said Step 4.3 includes: Step 4.3.1: Construct the roll angle tracking error penalty term using Equation (8) :[[]]END]] (8) In Equation (8), is the set limit value of the roll angle tracking accuracy; Step 4.3.2: Construct a penalty term for stable tracking and chattering suppression using Equation (9) :[[]]END]] (9) Step 4.3.3: Construct the active suspension actuator power mutation penalty term using Equation (10) :[[]]END]] (10) In formula (10), is the set limit value of the active force mutation; Step 4.3.4: Construct the total reward function using Equation (11) :[[]]END]] (11) In formula (11), , , are three weighting coefficients.

6. A sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning according to claim 5, characterized in that, The said Step 4.5 includes: Step 4.5.1: Randomly initialize the parameters of the Critic1 network , the parameters of the Critic2 network ; and assign the parameters of the Critic1 network to the parameters of the first target Critic1 network , and assign the parameters of the Critic2 network to the parameters of the second target Critic2 network ; randomly initialize the parameters of the Actor network ; Step 4.5.2: Input the state at time into the Actor network and output the action at time. After that, apply it to the vehicle and obtain the state at time . Calculate the reward at time , and obtain a sample data at time, and store it in the experience pool; ​ Step 4.5.3: Randomly draw a sample data from the experience pool , and input and into Critic1 network and Critic2 network, and output the first value and the second value ; Input into the input Actor network and output the action at the moment ; Input and into the target Critic1 and target Critic2 networks, and output the first target value of the value of the state-action pair and the second target value ; Calculate according to formula (12) The target value at the moment :[[]]END]] (12) In formula (11), is the discount factor; Step 4.5.4: respectively with and calculate the mean square error to construct the Loss function, and then use the gradient descent method to update the parameters of the Critic1 network, the parameters of the Critic2 network to obtain the updated parameters of the Critic1 network and assign them to ; Step 4.5.5: Update the target Critic1 network parameters using Equation (13) and those of the target Critic2 network :[[]]END]] (13) In formula (13), is the smoothing factor, is the assignment symbol; Step 4.5.6: According to , adopt a delayed update mechanism to update the parameters of the Actor network , and obtain the updated parameters of the Actor network and assign them to ; Step 4.5.7: Iteratively train the agent according to the process of Steps 4.5.2 - 4.5.6 until the total reward converges to a high reward level, thereby obtaining an optimal agent for offline deployment.

7. A sliding mode adaptive vehicle roll stability control method based on deep reinforcement learning according to claim 1, characterized in that, The said Step 5 includes: Step 5.1: Calculate the active roll moment according to Equation (14) :[[-END]] ( 4) In Equation (14), , , are the three optimal sliding mode parameters output by the optimal agent according to the vehicle state; Step 5.2: Calculate according to formula (15) and : (15)。 8. An electronic device, comprising a memory and a processor, characterized in that, The memory stores a program for implementing the vehicle roll stability control method according to any one of claims 1 - 7, and when the processor executes the program, active suspension reverse roll control is implemented.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the vehicle roll stability control method according to any one of claims 1 - 7 are implemented.

Citation Information

Cited By

  • Reinforcement learning semi-active suspension intelligent control method based on Bayesian optimization

    CN120792407A

  • Vehicle rollover prevention control method and device based on angle module, equipment and medium

    CN120886817A

  • Suspension control method and system based on reinforcement learning and multiple agents and vehicle

    CN121650389A

  • Suspension control method and system based on reinforcement learning and multi-agent, and vehicle

    CN121650389B

  • Industrial coal blending optimization method and system based on multi-agent collaboration and program product

    CN121707077A