Automobile semi-active suspension PID parameter adaptive adjustment and vibration control method and system based on reinforcement learning

By constructing an adaptive semi-active suspension PID controller using reinforcement learning, the problem of suspension performance degradation under nonlinear and time-varying conditions caused by traditional PID control methods is solved. This enables autonomous learning and multi-objective optimization of the suspension system, thereby improving vehicle ride comfort and handling stability.

CN121084104APending Publication Date: 2025-12-09ANHUI TUOSHENG AUTO PARTS CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511325412.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Traditional PID control methods cannot take into account various road conditions under nonlinear and time-varying conditions, resulting in a decline in suspension performance. Fuzzy PID control methods have poor generalization ability, and reinforcement learning can directly output damping force, which can easily lead to excessive damping force and lacks PID safety guarantee.

Method used

A semi-active suspension PID controller with adaptive adjustment is constructed using reinforcement learning. Combined with a variable structure PID controller and the TD3 algorithm, the suspension parameters are adjusted in real time to cope with different road conditions and vehicle load changes. Through state space and motion space design, ride comfort, driving safety and handling stability are optimized.

Benefits of technology

It significantly improves vehicle ride comfort and handling stability, reduces vehicle vertical acceleration and suspension travel, lowers tire dynamic load, and achieves autonomous learning and multi-objective optimization of the suspension system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121084104A_ABST
    Figure CN121084104A_ABST
Patent Text Reader

Abstract

The invention discloses an automobile semi-active suspension PID parameter adaptive adjustment and vibration control method and system based on reinforcement learning, and the method mainly comprises the following working steps: 1, constructing a reinforcement learning environment comprising a suspension dynamics high-fidelity simulation model and a PID controller; 2, defining a state space including a vehicle body vertical acceleration error e (t), an error integral term summa e (t) dt, an error differential term de (t) / dt, a suspension dynamic stroke zdef (t), a driving vehicle speed v (t) and road surface unevenness estimation xi (t); and 3, defining an action space including a proportion coefficient increment delta Kp, an integral coefficient increment delta Ki, a differential coefficient increment delta Kd and an output compensation amount delta e (t) of the PID controller. According to the method, the problem that the suspension performance is reduced due to the fact that control parameters of a traditional PID control method are fixed under the non-linear and time-varying working conditions is solved, and the riding comfort and the handling stability of the vehicle are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent control technology for automotive suspension, specifically relating to a method and system for adaptive adjustment of PID parameters and vibration control of semi-active automotive suspension based on reinforcement learning. Background Technology

[0002] For automobiles, the ability to quickly and accurately adjust the damping force of shock absorbers is crucial for improving ride comfort and handling stability. In existing technologies, traditional PID control methods suffer from poor adaptability due to fixed parameters. For example, on smooth roads, a smaller damping force is needed to ensure comfort, while on bumpy roads, a larger damping force is required to suppress high-frequency vibrations. Fixed PID controller parameters cannot accommodate various road conditions. Fuzzy PID control methods require pre-defined rules and pre-set parameter mapping tables for gain scheduling, resulting in poor generalization and an inability to cover changing road conditions caused by variations in road surface grade and dynamic load transfer. Self-tuning PID control methods adjust PID control parameters only based on the error amplitude, without introducing dynamic coupling in the state space, thus their control performance still needs improvement. Furthermore, the strong nonlinearity of the suspension system can also lead to a decrease in control performance. While reinforcement learning methods can directly output damping force to achieve suspension vibration control, this approach is prone to generating excessive damping force without the safety margin of PID control and increases actuator losses. Therefore, leveraging the powerful environmental interaction capabilities of reinforcement learning methods, developing a semi-active suspension control method and system that can learn autonomously, take into account multi-objective optimization, and retain the PID framework has broad application prospects. Summary of the Invention

[0003] This invention employs reinforcement learning to achieve adaptive adjustment of the parameters of the semi-active suspension PID controller, effectively overcoming the problem of suspension performance degradation caused by fixed control parameters under nonlinear and time-varying conditions in traditional PID control methods. This reduces vehicle vertical acceleration, suspension travel, and tire dynamic load, significantly improving vehicle ride comfort and handling stability.

[0004] The present invention discloses a method and system for adaptive adjustment of PID parameters and vibration control of semi-active suspension for automobiles based on reinforcement learning, the contents of which are as follows.

[0005] The constructed reinforcement learning environment includes:

[0006] The vehicle's vertical vibration model is simplified to a 1 / 4 scale semi-active suspension dynamic model, including the sprung mass m. s Unsprung mass m u The suspension spring stiffness is k s The equivalent passive damping coefficient of the vibration damper is c. s The tire stiffness is k tThe vertical displacement of the spring-loaded mass is z s The vertical displacement of the unsprung mass is z. u The road surface unevenness input is z r The damping force generated by the semi-active suspension magnetorheological damper is F. damp Therefore, the dynamic model of a 1 / 4 car semi-active suspension is as follows:

[0007]

[0008] The road surface unevenness input uses a filtered white noise model:

[0009]

[0010] In the formula, G0 is the road surface roughness coefficient, ω(t) is Gaussian white noise with a mathematical expectation of 0, and f0 is the lower cutoff frequency.

[0011] The calculation model for the damping force of a magnetorheological damper is as follows:

[0012]

[0013] In the formula, α is the damping force amplitude coefficient, β1 is the velocity term coefficient, β2 is the current term coefficient, γ is the linear damping coefficient, and I is the control current.

[0014] For tire stiffness, a tire hysteresis compensation model is established:

[0015]

[0016] In the formula, k t0 δ represents the static stiffness of the tire, η represents the hysteresis strength coefficient, and η represents the speed influence coefficient.

[0017] The ideal vibration state of a car is defined as a vertical acceleration of 0, i.e., complete elimination of car body vibration. Therefore, the vertical acceleration error e(t) of the car body is:

[0018] e(t) = 0 - z s (t)=-z s (t) (6)

[0019] The damping force adjustment of the magnetorheological vibration damper adopts a variable structure PID controller, with the proportional term of the controller providing the basic damping. In small error modes such as on smooth roads, the integral term continuously acts to eliminate steady-state errors; in large error modes such as on bumpy roads, the proportional gain of the controller is increased by 30% to 50%, the derivative term is activated to predict the vibration trend, and the integral term is deactivated to prevent the actuator from saturating due to the continuous accumulation of large errors.

[0020] The small error mode and the large error mode are separated by the human sensitivity threshold of 0.5g.

[0021] At time t, the variable structure PID controller model is:

[0022]

[0023] In the formula, F d_des K is the force output by the PID controller (i.e., the desired damping force). p1 and K i1 These are the proportional gain coefficient and integral gain coefficient in small error mode, respectively, K p2 and K d2 These are the proportional gain coefficient and the derivative gain coefficient under large error mode, respectively, and T is the control period.

[0024] The defined state space vector is:

[0025] s t =[e(t),∫e(t)dt,de(t) / dt,z def (t),v(t),ξ(t)] T (8)

[0026] Combining formula (6) and the suspension dynamic travel calculation formula z def =z s -z u ,have to:

[0027] s t =[-z s (t),∫(-z s (t))dt,d(-z s (t)) / dt,z s -z u ,v(t),ξ(t)] T (9)

[0028] The action space vector employs an incremental adjustment method to conform to the continuous and gradual change characteristics of the control parameters, avoiding system instability caused by abrupt parameter changes. Simultaneously, to improve control accuracy and prevent large control errors due to sudden changes in operating conditions, an output compensation term Δe(t) is added to the action space vector. Furthermore, Gaussian noise is added to the action space during the reinforcement learning model training phase.

[0029] The defined action space vector is:

[0030] a t =[ΔK p (t),ΔK i (t),ΔK d (t),Δe(t)] T (10)

[0031] The ranges for the proportional gain, integral gain, and derivative gain are set as follows: ΔK p ∝[-50,50],ΔK i ∝[-5,5],ΔK d ∝[-20,20].

[0032] After one time step, the adjusted PID control parameters are as follows:

[0033] K p (t+1)=K p (t)+ΔK p (t) (11)

[0034] K i (t+1)=K i (t)+ΔK i (t) (12)

[0035] K d (t+1)=K d (t)+ΔK d (t) (13)

[0036] Therefore, at time t+1, the target damping force output by the PID controller is:

[0037]

[0038] Therefore, the required damping force F of the magnetorheological damper at a specific time step under a specific operating condition can be calculated. d_des Considering the capacity constraints of the magnetorheological damper, the target damping force is limited based on the measured damping force range of the damper, i.e., F. d_des ∝[F d_min ,F d_max ].

[0039] Furthermore, the damping force of a magnetorheological damper depends not only on the input current of the magnetorheological fluid but also on the vertical vibration velocity of the suspension, exhibiting strong nonlinear characteristics. d_des The functional relationship between the control current and the vertical vibration velocity of the suspension is defined as follows:

[0040] Combining the target damping force of the shock absorber and the vertical vibration velocity of the suspension measured at the current time step, based on the pre-stored inverse model... By consulting the table, the required control current I can be obtained.

[0041] To quantify in state s t Next, execute action a t Then transition to the new state s t+1 The performance of the reward function is determined by the following:

[0042]

[0043] In the formula, ω1, ω2, ω3, and ω4 are the weight values ​​corresponding to the ride comfort index, driving safety index, handling stability index, and control smoothness index, respectively. Based on the preliminary research results, their values ​​are taken as 0.8, 0.5, 0.3, and 0.1, respectively.

[0044] The reinforcement learning algorithm used is the TD3 algorithm.

[0045] In the designed TD3 algorithm, the main Actor network is based on the state space vector s at time t. t Generate action vector a t That is, taking the state vector shown in equation (9) as input, the output is the action vector shown in equation (10):

[0046] a=μ(s|θ μ (16)

[0047] In the formula, μ is the principal deterministic strategy function, and θ μ The parameters of the main Actor network.

[0048] In the designed TD3 algorithm, the target Actor network is the same as the main Actor network, and the network parameters are θ. μ' Under the objective deterministic policy function μ', it generates action a' based on the state s' at the next time step:

[0049] a'=μ'(s'|θ μ' (17)

[0050] Adding Gaussian noise ε to the target Actor network improves the algorithm's robustness to noise and fluctuations in the target value.

[0051]

[0052] In the designed TD3 algorithm, the state-action pair (s, a) is used as input. Two independent Critic networks are responsible for evaluating the value of action a. The value estimate V of the dual Critic networks is calculated, and the minimum of the two values ​​is taken as the target V value to reduce the overestimation error of V and minimize the time difference error (TD error). The TD error e TD The calculation formula is:

[0053] e TD =r+γV(s')-V(s) (19)

[0054] In the formula, r is the instantaneous reward at the current moment, and γ is the discount factor. For the continuous control task of semi-active suspension damping force, its value is taken as 0.99 based on the results of previous research.

[0055] In the designed TD3 algorithm, the main Actor network structure adopts a multi-layer fully connected neural network. The tanh activation function is used in the output layer to constrain the ranges of the proportional gain, integral gain, and differential gain. Finally, gradient updates are performed by maximizing the action value function V(s,a) evaluated by the Critic network.

[0056]

[0057] In the designed TD3 algorithm, the target Actor network has the same structure as the main Actor network. It does not directly participate in the interaction but is only used to calculate the target V value. The parameter update mechanism is a soft update.

[0058] θ μ' ←τθ μ +(1-τ)θ μ' (twenty one)

[0059] In the formula, τ is the smoothing update coefficient of the target network parameters, which controls the speed at which the main Actor network parameters are synchronized to the target Actor network parameters. Based on the preliminary research results, its value is taken as 0.005.

[0060] In the designed TD3 algorithm, the Critic network is updated twice before the main Actor network and the target Actor network are updated to improve the stability of the policy.

[0061] The offline training and real-vehicle deployment process of the reinforcement learning-based adaptive adjustment and vibration control method for PID parameters of automotive semi-active suspension is as follows:

[0062] Step 1: Initialize the Actor network and Critic network (including the main network and the target network), as well as the experience replay buffer;

[0063] Step 2: Based on the established suspension dynamics model and PID controller model, obtain the state s at the current time step. t ;

[0064] Step 3: Based on the established Actor network, generate the action a at the current time step according to equation (16). t ;

[0065] Step 4: Execute the action and obtain the reward r at the current time step. t and the state s' in the next time step;

[0066] Step 5: Introduce Gaussian error and calculate the target motion. and the target V value;

[0067] Step 6: Calculate the TD error according to equation (19);

[0068] Step 7: Minimize the TD error and update the main Critic network;

[0069] Step 8: Maximize expected cumulative rewards and delay updating the main Actor network;

[0070] Step 9: Soft update the target network;

[0071] Step 10: State update, co-simulation loop;

[0072] Step 11: Save the parameters of the trained Actor network;

[0073] Step 12: Deploy the vehicle-mounted system using the TensorRT inference engine, load the reinforcement learning agent network parameters, set the decision cycle to 50ms, the PID control cycle to 10ms, and output the control current and target damping force in real time.

[0074] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0075] (1) The method and system for adaptive adjustment of PID parameters and vibration control of semi-active suspension of automobile based on reinforcement learning described in this invention can make full use of the autonomous exploration and learning capabilities of reinforcement learning. The designed PID controller can adapt to different road conditions in real time and automatically adjust the PID control parameters to cope with complex scenarios such as changes in vehicle load and vehicle speed.

[0076] (2) The method and system for adaptive adjustment of PID parameters and vibration control of semi-active suspension of automobiles based on reinforcement learning described in this invention comprehensively considers ride comfort index, driving safety index, handling stability index and control smoothness index when designing the reward function, which can improve the overall performance of automobiles; the designed state space and action space can effectively handle the nonlinear problems of the system.

[0077] (3) The method and system for adaptive adjustment of PID parameters and vibration control of semi-active suspension of automobile based on reinforcement learning described in this invention fully considers the needs of convenient engineering implementation. While retaining the traditional PID control framework, it adopts a mode of offline reinforcement learning training combined with online deployment, which can meet the real-time requirements of vehicle on-board while reducing training costs. Attached Figure Description

[0078] Figure 1 The flowchart shows a reinforcement learning-based method for adaptive adjustment of PID parameters and vibration control of semi-active suspension systems in automobiles.

[0079] Figure 2 This is an architecture for adaptive adjustment of PID parameters and vibration control of automotive semi-active suspension based on reinforcement learning.

[0080] Figure 3 This is a 1 / 4 scale model of a semi-active car suspension.

[0081] Figure 4 The Actor network structure for the TD3 algorithm of RL agents;

[0082] Figure 5 The Critic network structure for the TD3 algorithm of RL agents;

[0083] Figure 6 The structure of the TD3 algorithm for RL agents;

[0084] Figure 7 This document outlines the offline training and real-vehicle deployment process for a reinforcement learning-based adaptive adjustment and vibration control method for PID parameters in automotive semi-active suspension.

[0085] Figure 8 for Figure 7 The implementation process of the joint simulation loop;

[0086] Figure 9 The above are comparison results of the vehicle body vertical acceleration obtained by different control methods in the embodiments;

[0087] Figure 10 The above are comparison results of suspension dynamic travel obtained by different control methods in the embodiments;

[0088] Figure 11 The results show a comparison of tire dynamic loads obtained by different control methods in the examples. Detailed Implementation

[0089] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. The models and processes shown in the drawings are only used to explain the present invention and should not be construed as limiting the present invention. All other embodiments obtained based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0090] The proposed reinforcement learning-based semi-active suspension damper PID parameter adaptive adjustment and system control architecture is as follows: Figure 1 and Figure 2 As shown, the implementation method will be described in detail below.

[0091] The first step is to establish a semi-active suspension model for automobiles.

[0092] like Figure 3 As shown, the vertical vibration model of the whole vehicle is simplified to obtain a 1 / 4-inch semi-active suspension dynamic model. In the model, the sprung mass is denoted as m. s The unsprung mass is denoted as m. u The suspension spring stiffness is denoted as k.s The equivalent passive damping coefficient of the vibration damper is denoted as c. s Tire stiffness is denoted as k. t The vertical displacement of the spring-loaded mass is denoted as z. s The vertical displacement of the unsprung mass is denoted as z. u The road surface unevenness is input as z. r The damping force generated by the semi-active suspension magnetorheological damper is denoted as F. damp Therefore, the dynamic model of a 1 / 4 car semi-active suspension is derived as follows:

[0093]

[0094] The road surface unevenness input uses a filtered white noise model:

[0095]

[0096] In the formula, G0 is the road surface roughness coefficient, ω(t) is Gaussian white noise with a mathematical expectation of 0, and f0 is the lower cutoff frequency.

[0097] The calculation model for the damping force of a magnetorheological damper is as follows:

[0098]

[0099] In the formula, α is the damping force amplitude coefficient, β1 is the velocity term coefficient, β2 is the current term coefficient, γ is the linear damping coefficient, and I is the control current.

[0100] For tire stiffness, a tire hysteresis compensation model is established:

[0101]

[0102] In the formula, k t0 δ represents the static stiffness of the tire, η represents the hysteresis strength coefficient, and η represents the speed influence coefficient.

[0103] The ideal vibration state of a car is defined as a vertical acceleration of 0, i.e., complete elimination of car body vibration. Therefore, the vertical acceleration error e(t) of the car body is:

[0104] e(t) = 0 - z s (t)=-z s (t) (6)

[0105] The damping force adjustment of the magnetorheological vibration damper uses a variable structure PID controller, with the proportional term providing the basic damping. In small error modes such as on smooth roads, the integral term continuously acts to eliminate steady-state errors; in large error modes such as on bumpy roads, the proportional gain is increased by 30% to 50%, the derivative term is activated to predict vibration trends, and the integral term is deactivated to prevent the actuator from saturating due to the accumulation of continuous large errors.

[0106] The small error mode and large error mode defined above are separated by the human sensitivity threshold of 0.5g.

[0107] At time t, the variable structure PID controller model is:

[0108]

[0109] In the formula, F d_des K is the force output by the PID controller (i.e., the desired damping force). p1 and K i1 These are the proportional gain coefficient and integral gain coefficient in small error mode, respectively, K p2 and K d2 These are the proportional gain coefficient and the derivative gain coefficient under large error mode, respectively, and T is the control period.

[0110] The second step is to define the state vector, action vector, and reward function based on the semi-active suspension model.

[0111] Taking into account the vehicle's vertical acceleration error e(t), the integral term of the error ∫e(t)dt, the differential term of the error de(t) / dt, and the suspension dynamic travel z def Given the vehicle speed v(t) and the road roughness estimate ξ(t), the state space vector is defined as:

[0112] s t =[e(t),∫e(t)dt,de(t) / dt,z def (t),v(t),ξ(t)] T (8)

[0113] Combining formula (6) and the suspension dynamic travel calculation formula z def =z s -z u The state space vector is obtained as follows:

[0114] s t =[-z s (t),∫(-z s (t))dt,d(-z s (t)) / dt,z s -z u ,v(t),ξ(t)] T(9)

[0115] The action space vector adopts an incremental adjustment method to conform to the continuous and gradual change characteristics of the control parameters and avoid system instability caused by sudden parameter changes. The defined incremental adjustment of the PID control parameters includes the proportional coefficient increment ΔK. p Integral coefficient increment ΔK i , Differential coefficient increment ΔK d Meanwhile, to improve control accuracy and avoid large control errors caused by sudden changes in operating conditions, an output compensation term Δe(t) is added to the action space vector. Furthermore, Gaussian noise is added to the action space during the reinforcement learning model training phase.

[0116] Based on this, the defined action space vector is:

[0117] a t =[ΔK p (t),ΔK i (t),ΔK d (t),Δe(t)] T (10)

[0118] Considering the actual control requirements of semi-active suspension, the ranges for proportional gain, integral gain, and derivative gain are set as follows: ΔK p ∝[-50,50],ΔK i ∝[-5,5],ΔK d ∝[-20,20].

[0119] Therefore, after one time step, the adjusted PID control parameters are:

[0120] K p (t+1)=K p (t)+ΔK p (t) (11)

[0121] K i (t+1)=K i (t)+ΔK i (t) (12)

[0122] K d (t+1)=K d (t)+ΔK d (t) (13)

[0123] At time t+1, the target damping force output by the PID controller is:

[0124]

[0125] Therefore, the required damping force F of the magnetorheological damper at a specific time step under a specific operating condition can be calculated. d_des Considering the capacity constraints of the magnetorheological damper, the target damping force is limited based on the measured damping force range of the damper, i.e., F. d_des ∝[F d_min ,F d_max ].

[0126] Furthermore, the damping force of a magnetorheological damper depends not only on the input current of the magnetorheological fluid but also on the vertical vibration velocity of the suspension, exhibiting strong nonlinear characteristics. d_des The functional relationship between the control current and the vertical vibration velocity of the suspension is defined as follows:

[0127] Combining the target damping force of the shock absorber and the vertical vibration velocity of the suspension measured at the current time step, based on the pre-stored inverse model... By consulting the table, the required control current I can be obtained.

[0128] After defining the state vector and action vector, in order to quantize in state s t Next, execute action a t Then transition to state s t+1 The performance of the vehicle is determined by comprehensively considering minimizing the vertical acceleration of the vehicle body. Minimize suspension travel z def (t), Minimize tire dynamic load F dyn The reward function, designed to incorporate multi-objective constraints including (t), and penalizes excessively large parameter adjustments, is as follows:

[0129]

[0130] In the formula, ω1, ω2, ω3, and ω4 are the weight values ​​corresponding to the ride comfort index, driving safety index, handling stability index, and control smoothness index, respectively. Based on the preliminary research results, their values ​​are taken as 0.8, 0.5, 0.3, and 0.1, respectively.

[0131] The third step is to design reinforcement learning algorithms.

[0132] The reinforcement learning algorithm used is the TD3 algorithm, employing an Actor-Critic network architecture. Figure 4 and Figure 5 The diagrams show the designed Actor network and Critic network structures, respectively.

[0133] The Actor network structure is designed with 5 layers, and the input is the system state vector s. t The output is the action quantity a. t The Critic network structure is also designed with 5 layers, and the input includes the system state s.t and action a t That is, state-action pairs (s t ,a t The output is a value estimate V corresponding to the current state and action. Both the Actor and Critic networks are multi-layer fully connected neural networks, with the fully connected layers incorporating the ReLU activation function to introduce non-linearity. In the Actor network, the output layer uses the tanh activation function to constrain the action range.

[0134] The following is combined with Figure 6 The TD3 algorithm structure diagram shown below provides a detailed introduction to the setup and implementation process of the Actor network and Critic network.

[0135] In the designed TD3 algorithm, the main Actor network is based on the state space vector s at time t. t Generate action vector a t That is, taking the state vector shown in equation (9) as input, the output is the action vector shown in equation (10):

[0136] a=μ(s|θ μ (16)

[0137] In the formula, μ is the principal deterministic strategy function, and θ μ The parameters of the main Actor network.

[0138] The target Actor network is the same as the main Actor network, with network parameters θ. μ' Under the objective deterministic policy function μ', it generates action a' based on the state s' at the next time step:

[0139] a'=μ'(s'|θ μ' (17)

[0140] Adding Gaussian noise ε to the target Actor network improves the algorithm's robustness to noise and fluctuations in the target value.

[0141]

[0142] Using a state-action pair (s, a) as input, two independent Critic networks are responsible for evaluating the value of action a. The value estimate V is calculated using the dual Critic networks, and the minimum of the two estimates is taken as the target V value to reduce the overestimation error of V and minimize the time difference error (TD error). The TD error is e. TD The calculation formula is:

[0143] e TD =r+γV(s')-V(s) (19)

[0144] In the formula, r is the instantaneous reward at the current moment, and γ is the discount factor. For the continuous control task of semi-active suspension damping force, its value is taken as 0.99 based on the results of previous research.

[0145] Gradient updates are performed by maximizing the action value function V(s,a) evaluated by the Critic network:

[0146]

[0147] The target Actor network has the same structure as the main Actor network. It does not directly participate in the interaction; it is only used to calculate the target V value. Its parameter update mechanism is a soft update.

[0148] θ μ' ←τθ μ +(1-τ)θ μ' (twenty one)

[0149] In the formula, τ is the smoothing update coefficient of the target network parameters, which controls the speed at which the main Actor network parameters are synchronized to the target Actor network parameters. Based on the preliminary research results, its value is taken as 0.005.

[0150] To improve the stability of the control strategy, the main Actor network and the target Actor network are updated only after two Critic network updates are completed.

[0151] The fourth step is offline training and real-vehicle deployment.

[0152] After completing the design of the reinforcement learning environment (semi-active suspension simulation modeling and PID controller design), the definition of state vectors, action vectors and reward functions, and the design of the reinforcement learning algorithm, as follows: Figure 7 As shown, the proposed reinforcement learning-based semi-active suspension vibration control method is trained offline and deployed in a real vehicle. The specific process is as follows:

[0153] Step 1: Initialize the Actor network and Critic network (including the main network and the target network), as well as the experience replay pool;

[0154] Step 2: Based on the established suspension dynamics model and PID controller model, obtain the state s at the current time step. t ;

[0155] Step 3: Based on the established Actor network, generate the action a at the current time step according to equation (16). t ;

[0156] Step 4: Execute the action and obtain the reward r at the current time step. t and the state s' in the next time step;

[0157] Step 5: Introduce Gaussian error and calculate the target motion. and the target V value;

[0158] Step 6: Calculate the TD error according to equation (19);

[0159] Step 7: Minimize the TD error and update the main Critic network;

[0160] Step 8: Maximize expected cumulative rewards and delay updating the main Actor network;

[0161] Step 9: Soft update the target network;

[0162] Step 10: State update, co-simulation loop;

[0163] Step 11: Save the parameters of the trained Actor network;

[0164] Step 12: Deploy the vehicle-mounted system using the TensorRT inference engine, load the reinforcement learning agent network parameters, set the decision cycle to 50ms, the PID control cycle to 10ms, and output the control current and target damping force in real time.

[0165] In the next state, road surface information and vehicle vertical acceleration are measured again by sensors. Based on the new state vector and the saved reinforcement learning Actor network parameters, the PID controller updates the control parameters, calculates the new control current, and outputs the damping force. This process is repeated continuously, and the semi-active suspension magnetorheological damper continuously adjusts the damping force to achieve effective control of the vehicle's vertical vibration.

[0166] Taking a certain automobile as an example, the implementation of the proposed method is as follows.

[0167] The vehicle parameters and test conditions are shown in Table 1.

[0168] Table 1

[0169]

[0170] The parameters of the suspension and magnetorheological damper are shown in Table 2.

[0171] Table 2

[0172]

[0173] The vertical vibration control effect of the proposed method was verified according to the above implementation method. Tests were conducted on an ISO 8608 Class C random road surface for 10 seconds, with the vehicle controlled to pass over a speed bump at 5 seconds to evaluate transient impact conditions. The comparison results of vehicle body vertical acceleration, suspension dynamic travel, and tire dynamic load obtained using traditional PID control, fuzzy PID control, and the proposed reinforcement learning-modified PID control method (RL-PID method) are as follows: Figure 9 , Figure 10 and Figure 11 As shown, the relevant summary is as follows:

[0174] (1) For the vertical acceleration of the vehicle body, when using the traditional PID method, the impact peak is the highest and the recovery time is long; when using the fuzzy PID method, the impact response is improved by 31%, but the steady-state vibration is obvious; when using the RL-PID method, the impact peak is reduced by 33.7%, the vibration decay is the fastest, and the steady-state vibration can be more effectively suppressed.

[0175] (2) Regarding the suspension travel, under impact conditions, when using the traditional PID method, the maximum suspension travel is 0.092m (close to the limit block); the fuzzy PID method reduces the maximum travel by 18.5%, but it still exceeds the 0.08m safety threshold; when using the RL-PID method, the maximum travel is 0.068m, which is always below the safety threshold. Under steady-state conditions, the RL-PID method also achieves the optimal control effect.

[0176] (3) Regarding tire dynamic load, when using the traditional PID method, the tire dynamic load under impact conditions exceeds 1800N, resulting in the highest risk of ground lift-off. The fuzzy PID method reduces the tire dynamic load by 9.3%, but transient ground lift-off issues still exist. The tire dynamic load controlled by the RL-PID method is the smallest, maintaining the best ground contact performance. Under steady-state conditions, the RL-PID method also achieves the best control effect.

[0177] A comprehensive performance comparison was conducted, and the results are shown in Table 3.

[0178] Table 3

[0179]

[0180] The proposed method and system for adaptive adjustment of PID parameters and vibration control of semi-active suspension for automobiles based on reinforcement learning achieves effective improvements in ride comfort, driving safety, handling stability, and real-time response.

[0181] The above description is only one embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. All equivalent substitutions or modifications made under the concept of the present invention based on the description and drawings of the present invention, or direct / indirect applications in other related technical fields, should be covered within the scope of protection of the present invention.

Claims

1. A method for adaptive adjustment of PID parameters and vibration control of a semi-active automotive suspension based on reinforcement learning, characterized in that, The work includes the following steps: Step 1: Construct a reinforcement learning environment; Step 2: Define the reinforcement learning state space s t ; Step 3: Define the reinforcement learning action space a t ; Step 4: Design the reinforcement learning reward function r t ; Step 5: Offline training of the reinforcement learning agent; Step Six: Deploy reinforcement learning agents in actual vehicles.

2. The method for adaptive adjustment of PID parameters and vibration control of semi-active vehicle suspension based on reinforcement learning according to claim 1, characterized in that, The semi-active suspension uses a magnetorheological damper as the actuator, and the damping force is adjusted by controlling the input current of the magnetorheological damper.

3. The method for adaptive adjustment of PID parameters and vibration control of semi-active vehicle suspension based on reinforcement learning according to claim 1, characterized in that, The reinforcement learning environment includes a CarSim / Simulink co-simulation high-fidelity dynamics model, a PID controller, and a state observer.

4. The method for adaptive adjustment of PID parameters and vibration control of semi-active vehicle suspension based on reinforcement learning according to claim 1, characterized in that, The state space s t This includes the vertical acceleration error of the vehicle body e(t), the integral term of the error ∫e(t)dt, the differential term of the error de(t) / dt, and the suspension dynamic travel z. def (t), vehicle speed v(t), and road surface roughness estimation ξ(t).

5. The method for adaptive adjustment of PID parameters and vibration control of semi-active vehicle suspension based on reinforcement learning according to claim 1, characterized in that, The action space a t Including the PID controller proportional coefficient increment ΔK p Integral coefficient increment ΔK i , Differential coefficient increment ΔK d The output compensation amount Δe(t) is adjusted so that the proportional coefficient, integral coefficient and derivative coefficient satisfy the set constraints.

6. The method for adaptive adjustment of PID parameters and vibration control of semi-active vehicle suspension based on reinforcement learning according to claim 1, characterized in that, The reward function r t Including minimizing the vertical acceleration of the vehicle body Minimize suspension travel z def (t), Minimize tire dynamic load F dyn (t), and to prevent the control parameters from oscillating violently, and to punish excessive parameter adjustment actions.

7. The method for adaptive adjustment of PID parameters and vibration control of semi-active vehicle suspension based on reinforcement learning according to claim 1, characterized in that, The reinforcement learning algorithm is the dual-delay deep deterministic policy gradient algorithm TD3, in which both the Actor network and the Critic network are deep neural networks.

8. The method for adaptive adjustment of PID parameters and vibration control of semi-active vehicle suspension based on reinforcement learning according to claim 1, characterized in that, During the offline training phase, the optimal mapping strategy from state to action is learned by maximizing cumulative rewards.

9. The method for adaptive adjustment of PID parameters and vibration control of semi-active vehicle suspension based on reinforcement learning according to claim 1, characterized in that, During the actual vehicle deployment phase, the parameters of the PID controller are adjusted in real time by using coefficient increments and compensation outputs to control the input current of the magnetorheological damper, thereby adjusting the damping force and achieving vertical vibration control of the vehicle.

10. A reinforcement learning-based adaptive adjustment and vibration control system for PID parameters of a semi-active automotive suspension, comprising the reinforcement learning-based adaptive adjustment and vibration control method for PID parameters of a semi-active automotive suspension as described in any one of claims 1-9, characterized in that, It also includes the following modules: Module 1: Sensor module, which collects vehicle vertical acceleration, suspension displacement, and wheel speed; Module 2: State Observation Module, Calculating the State Vector s t This includes e(t), ∫e(t)dt, de(t) / dt, and z. def (t), v(t), and ξ(t); Module 3: Reinforcement Learning Decision Module, loading pre-trained agents and outputting control actions in real time, including ΔK. p ΔK i ΔK d and Δe(t); Module 4: PID controller execution module, which calculates the target current of the magnetorheological damper based on the updated parameters and adjusts the damping force in real time.

11. A reinforcement learning-based adaptive adjustment and vibration control system for PID parameters of a semi-active automotive suspension as described in claim 10, characterized in that, The sensor module acquires data in the CarSim / Simulink co-simulation environment.

12. A reinforcement learning-based adaptive adjustment and vibration control system for PID parameters of a semi-active automotive suspension as described in claim 10, characterized in that, The reinforcement learning decision module uses the TensorRT inference engine on the vehicle-mounted embedded platform.

Citation Information

Cited By

  • Active suspension full-band vibration suppression method

    CN121572753A

  • A full-band vibration suppression method for active suspension

    CN121572753B