A fully active suspension control method based on deep reinforcement learning and energy recovery

By optimizing suspension parameters through deep reinforcement learning, the contradiction between energy recovery and ride comfort in traditional suspension control methods has been resolved. This has enabled the suspension system to be optimized collaboratively in complex environments, thereby improving the vehicle's energy efficiency and comfort.

CN120116681BActive Publication Date: 2025-11-14HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510561828.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-11-14
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Traditional suspension control methods struggle to dynamically balance the conflict between energy recovery and ride comfort in complex and nonlinear environments, failing to simultaneously improve vehicle energy efficiency and ride comfort.

Method used

A fully active suspension control method based on deep reinforcement learning is adopted. By constructing a simulation dynamics model and a stochastic road surface model, the suspension state and action space are defined. The suspension parameters are optimized using a deep reinforcement learning network, and the network is trained in combination with the PPO algorithm to achieve the collaborative optimization of the suspension system.

Benefits of technology

To achieve good comfort and road adaptability of the suspension system under different road surface and driving conditions, while improving the overall energy efficiency of the vehicle, the electromagnetic energy recovery device converts vibration energy into electrical energy for storage, and dynamically adjusts suspension parameters to maximize energy recovery efficiency and ride comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120116681B_ABST
    Figure CN120116681B_ABST
Patent Text Reader

Abstract

This invention discloses a fully active energy recovery suspension control method based on deep reinforcement learning, comprising: 1. constructing an energy recovery fully active suspension simulation model and a simulated road surface model; 2. constructing state variables and motion variables, and processing the motion variables to control the suspension output control force; 3. designing a reward function; 4. constructing a policy-evaluation network and training the neural network based on the PPO algorithm to dynamically generate a suspension control strategy. This invention can improve vehicle performance under various road surface and driving conditions by optimizing the dynamic response of the suspension system, thereby achieving synergistic optimization of energy recovery and suspension performance to meet the multiple demands of intelligent vehicles for comfort, safety, and energy efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of fully active suspension control technology in intelligent driving, and is particularly suitable for vehicles with high requirements for driving smoothness and ride comfort. Background Technology

[0002] An active suspension system is a car suspension system that can adjust the suspension stiffness and damping in real time according to road conditions and driver needs. The active suspension system adjusts suspension stiffness and damping through actuators, and uses a controller to calculate the suspension system's control strategy in real time based on sensor signals, and controls the output state of the actuators.

[0003] Fully active suspension systems significantly improve vehicle vibration and bump response during driving by actively adjusting suspension forces, thus enhancing passenger comfort. Traditional passive and semi-active suspension systems, limited by their fixed parameters, struggle to maintain optimal performance across various road conditions. By introducing reinforcement learning algorithms, adaptive control in complex road conditions can be achieved, further improving ride comfort. Furthermore, with increasing demands for energy conservation and environmental protection, electromagnetic energy recovery technology is gradually becoming an important development direction for suspension systems. This technology converts the vibration energy of the suspension system into electrical energy for storage using electromagnetic devices, not only improving the overall energy efficiency of the vehicle but also providing new possibilities for intelligent control of the suspension system.

[0004] Traditional suspension control methods, such as PID control and fuzzy control, have limited performance in complex and nonlinear environments. Reinforcement learning, as an emerging intelligent control method, can effectively handle complex and variable control tasks through interactive learning with the environment. However, a significant technical contradiction exists between energy recovery and ride comfort in fully active suspension systems: to maximize vibration energy capture efficiency, suspension damping or stiffness needs to be increased to enhance energy conversion, but this leads to increased vertical vehicle acceleration and reduced ride comfort; to reduce road impact transmission, suspension damping or stiffness needs to be reduced to improve compliance, but this weakens the response efficiency of the energy recovery device. Traditional control methods struggle to dynamically balance this contradiction. Summary of the Invention

[0005] This invention addresses the shortcomings of existing technologies by proposing a fully active energy recovery suspension control method based on deep reinforcement learning. The aim is to improve vehicle performance under various road surfaces and driving conditions by optimizing the dynamic response of the suspension system, thereby achieving synergistic optimization of energy recovery and suspension performance to meet the multiple demands of intelligent vehicles for comfort, safety, and energy efficiency.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] The present invention provides an energy recovery fully active suspension control method based on deep reinforcement learning, characterized by the following steps:

[0008] Step 1: Construct a simulation dynamics model and a stochastic road surface model for the energy recovery 1 / 2 vehicle's four-degree-of-freedom fully active suspension;

[0009] Step 2: Define the suspension state parameter set , , This represents the total number of steps. For the first The normalized state variables of the step, and have ,in, For the first The car body acceleration of the step, , The first The front and rear wheel suspension dynamic deflection of the step. For the first The vehicle body displacement of the step, For the first The remaining battery power of the step;

[0010] Define the motion space set of the suspension ,in, Here, is the suspension stiffness, and is the suspension damping coefficient. , The magnitude of the external current connected to the motor. This refers to the PWM duty cycle of the H-bridge in the energy recovery circuit.

[0011] Step 3, construct the first reward number of steps ;

[0012] Step 4: Construct a deep reinforcement learning network, including an input layer with 5 neurons. The network consists of a fully connected hidden layer, an action network with an output layer of 4 neurons, and an input layer of 5 neurons. A value network consisting of a fully connected layer of hidden layers and an output layer of one neuron;

[0013] Step 5: The input is processed in the motion network to obtain the suspension in the first... Step movement volume The simulation dynamics model executes the first... Step movement volume After that, the first +1 step status Thus constituting the first 1 sample, denoted as This yields m samples, which are then stored in the sample pool D.

[0014] Step 6: Based on the sample pool D, train the deep reinforcement learning network using the PPO algorithm to obtain the optimal deep reinforcement learning model;

[0015] Step 7: Use the optimal action network in the optimal deep reinforcement learning model to process the real-time input state information and output corresponding control commands to achieve active control of the energy recovery fully active suspension.

[0016] The energy recovery fully active suspension control method based on deep reinforcement learning described in this invention is characterized in that, in step 3, the first step is to construct the first step using equation (1). reward number of steps ;

[0017] (1)

[0018] In equation (1), Let represent the energy recovery reward value at step i, and we have:

[0019] (2)

[0020] In equation (2), Indicates the battery capacity threshold. Let represent the energy recovery efficiency at step i, and we have:

[0021] (3)

[0022] (4)

[0023] In equation (3), The battery level at step i. This is the maximum usable battery capacity; when the remaining battery capacity is >90%, the energy recovery circuit opens and energy recovery stops.

[0024] In equation (4), To recover energy for the suspension in step i, Let be the total energy absorbed by the suspension in step i.

[0025] Furthermore, in step 5, the i-th step output value of the action network is obtained using equation (5). ;

[0026] (5)

[0027] In equation (5), These are the parameters of the action network; Indicates the first The output of the fully connected layer The corresponding activation value, Tanh is the output function, This is the bias matrix of the output layer. Let be the bias vector of the output layer, and use equation (6) to calculate the output of the first fully connected layer. Corresponding activation value The output of the k-th fully connected layer Corresponding activation value :

[0028] (6)

[0029] In equation (6), The bias matrix of the first fully connected layer, Let be the bias matrix of the k-th fully connected layer, and ReLU be the activation function. This is the bias vector of the first fully connected layer. Let K be the bias vector of the k-th fully connected layer. ;

[0030] Using equation (7), the suspension at the first... Step movement volume ,in, For the first The suspension stiffness of the step, For the first The suspension damping coefficient of the step, For the first The magnitude of the external current connected to the stepper motor; For the first The PWM duty cycle of the H-bridge in the energy recovery circuit of the step;

[0031] (7)

[0032] In equation (8), Indicates the first Random noise in the step.

[0033] Furthermore, step 6 includes the following steps:

[0034] Step 6.1: Calculate the first step using equation (8). Priority weight of each sample :

[0035] (8)

[0036] In equation (8), , , , There are 4 weighting coefficients;

[0037] If the first Vertical acceleration of the car body Exceeding the comfort threshold or the Front and rear wheel suspension dynamic deflection Approaching the mechanical limit or the first Step's battery power If the protection limit is reached, it will increase. The value of ; otherwise, Remain unchanged;

[0038] Step 6.2: Calculate the first step using equation (9). The probability of a sample being selected Thus, the probability of all samples being selected is obtained;

[0039] (9)

[0040] In equation (9), α is the priority intensity coefficient that controls the degree of sampling bias; The priority weight of the k-th sample;

[0041] Step 6.3: Define the current iteration number as c, and initialize c=1; define the update frequency as... The maximum number of iterations is Define the parameters of the action network in the c-th iteration as follows: Define the environment interaction parameters of the action network in the c-th iteration as follows: Let the parameters of the value network in the c-th iteration be denoted as... ;

[0042] Step 6.4: Calculate the value of the first iteration in the c-th iteration using equation (10). The probability ratio of the strategy of each sample :

[0043] (10)

[0044] In equation (10), Let be the policy output by the action network in the c-th iteration. The environment interaction strategy output by the action network in the c-th iteration; The strategy for the c-th iteration In state Select action at time The probability of; The strategy for the c-th iteration In state Select action at time The probability of;

[0045] Based on the probability that all samples are selected, a batch of samples is drawn from the sample pool D for the c-th iteration. Based on a batch of samples from the c-th iteration Using equation (11), the priority-weighted action objective function of the action network in the c-th iteration is constructed. :

[0046] (11)

[0047] In equation (11), The cropping threshold, For the first The expected value of a sample in the c-th iteration. For the clipping function, Indicates the first The advantage function of the sample in the c-th iteration is:

[0048] (12)

[0049] In equation (12), , This indicates that the value network in the c-th iteration respectively... and The predicted value;

[0050] Step 6.5: Construct the priority-weighted value function loss of the value network in the c-th iteration using equation (13). :

[0051] (13)

[0052] In equation (13), Indicates the discount factor. , This indicates that the value network in the c-th iteration respectively... and The predicted value;

[0053] Step 6.6: If and When the remainder of the ratio is 1, use equations (14) and (15) respectively to... , Update the parameters of the policy network at step c+1. And the parameters of the value network at step c+1 Otherwise, no update. , :

[0054] (14)

[0055] In equation (14), For learning rate, For policy network parameters The gradient operator;

[0056] (15)

[0057] In equation (15), Parameters of the value network The gradient operator;

[0058] Step 6.7: [The sentence is incomplete and requires more context to be translated accurately.] Assign to Then, make a judgment If the condition is met, the optimal deep reinforcement learning model has been obtained; otherwise, return to step 6.4 and execute sequentially.

[0059] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the energy recovery fully active suspension control method, and the processor is configured to execute the program stored in the memory.

[0060] The present invention discloses a computer-readable storage medium on which a computer program is stored, characterized in that the computer program, when executed by a processor, performs the steps of the energy recovery fully active suspension control method.

[0061] Compared with existing technologies, the present invention has the following beneficial effects:

[0062] 1. This invention innovatively applies deep reinforcement learning methods to fully active suspension control. The method is first trained extensively in a simulation environment, and then applied to HIL bench testing and real-vehicle testing after meeting the requirements. Because this application uses deep neural network reinforcement learning, the suspension can maintain good comfort and road adaptability under different road conditions while ensuring safety. Simultaneously, by introducing an electromagnetic energy recovery device, this invention can convert the vibration energy of the suspension system into electrical energy for storage, significantly improving the overall energy efficiency of the vehicle and achieving synergistic optimization of energy recovery and suspension performance.

[0063] 2. This invention assumes that the samples are independent and identically distributed when training the neural network. However, the data collected through reinforcement learning exhibits temporal correlation, and directly using this data for sequential training may lead to bias and instability in the policy optimization process. To address the issue of sample independence and improve sample efficiency, the PPO method of this invention creates a finite-sized trajectory pool to store the agent's experience sequences in the environment. Before each training step, a certain amount of fresh experience is collected based on the current policy and added to the trajectory pool. Subsequently, in each update cycle, PPO randomly selects multiple mini-batches from the trajectory pool to approximate the advantage function and adjust the policy parameters. This approach not only makes fuller use of the samples but also reduces the correlation between samples by randomly selecting mini-batches, helping to avoid local optima and promoting more robust policy learning, thereby ensuring the convergence and stability of the training process.

[0064] 3. This invention employs a multi-objective reward function design based on the PPO algorithm, taking the real-time energy recovery rate of the electromagnetic energy recovery device as one of the key optimization objectives, while also considering indicators such as vehicle vertical acceleration, suspension dynamic deflection (and vehicle vertical displacement). Under different driving conditions (such as bumpy roads, cornering, and straight-line driving), it can dynamically adjust suspension parameters to maximize energy recovery efficiency while ensuring comfort, significantly improving the overall performance of the vehicle.

[0065] 4. This invention boasts strong intelligence and adaptability. It can automatically adjust the suspension control strategy based on real-time road conditions and driving needs, adapting to different driving styles and environmental changes. For example, it prioritizes comfort and energy recovery on bumpy roads, prioritizes ride comfort when cornering, and maximizes energy recovery when driving straight. This intelligent control method not only enhances the driving experience but also reduces energy consumption, aligning with the trend of energy conservation and environmental protection. Attached Figure Description

[0066] Figure 1 A deep reinforcement learning control framework diagram for a fully active suspension system with energy recovery;

[0067] Figure 2 This is a schematic diagram of the four-quadrant operation principle of the motor;

[0068] Figure 3 This is a schematic diagram of an energy recovery circuit.

[0069] Figure 4 This is a schematic diagram of a deep reinforcement learning algorithm based on the PPO network. Detailed Implementation

[0070] The technical solutions will now be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention.

[0071] In this embodiment, an energy recovery fully active suspension control method based on deep reinforcement learning includes the following steps:

[0072] Step 1: Construct a simulation dynamic model and a stochastic road surface model for the energy recovery 1 / 2 vehicle four-degree-of-freedom fully active suspension, and select Class A, Class B, and Class C road surfaces as stochastic road surface models for low-speed, medium-speed, and high-speed conditions, respectively.

[0073] like Figure 1 As shown, the fully active suspension reinforcement learning control framework of this embodiment includes the following parts: a fully active suspension reinforcement learning controller body, a fully active suspension simulation model, a simulated road surface model, an electromagnetic vibration energy recovery device, state observations, fully active suspension action quantities, and rewards. The controller obtains state observations (suspension dynamic deflection, vehicle acceleration, and vehicle vertical displacement) from the fully active suspension system. It uses a controller strategy to control the suspension stiffness, suspension damping coefficient, motor current magnitude, and H-bridge PWM duty cycle of the energy recovery circuit for each state, supplying these parameters to the fully active suspension and its electromagnetic actuators. The suspension changes its current state quantity according to the active control force output by the controller. The actuator recovers energy based on the current output by the controller and the load resistance value, and generates a reward based on this state to evaluate the effectiveness of the strategy and update it accordingly.

[0074] The fully active suspension four-degree-of-freedom model can describe the vertical motion of the vehicle body, pitch motion, and the vertical motion of the front and rear wheels. It can more comprehensively reflect the dynamic characteristics of the vehicle in the vertical direction and avoid the complexity and computational burden caused by multi-directional coupling, focusing on the core vertical dynamics problem.

[0075] The main objective of this invention is to design a suspension control method for vehicles traveling on good road surfaces. Therefore, a random road surface model is adopted: the unevenness of random road surfaces is generally described by the statistical characteristic parameter of spatial frequency power spectral density and its time-domain form. The fitting expression is shown in equation (1):

[0076] (1)

[0077] in: Vertical displacement power spectral density ; Spatial frequency Define the number of wavelengths per meter of road surface; Reference spatial frequency , =0.1 ; Road surface roughness coefficient ; 𝑊 is the frequency exponent, usually 𝑊=2.

[0078] Road surface roughness is classified into eight levels according to the road surface power spectral density, as shown in Table 1. The table shows the road surface roughness of each level when φ=2. The upper and lower limits and their geometric mean are listed, along with 0.011. < 𝑛 <2.83 The root mean square value of road surface roughness within the range The geometric mean.

[0079] Table 1 Standard Classification Table of Road Surface Roughness

[0080]

[0081] Based on the requirements, Class A, Class B, and Class C road surfaces were selected as stochastic road surface models for low-speed, medium-speed, and high-speed operating conditions, respectively. These models were used to simulate highways, urban roads, and old roads with bumps, respectively.

[0082] Step 2: Considering the impact of different driving conditions on suspension control force, define the suspension's motion space and state space parameters based on the main performance evaluation indicators of the suspension system:

[0083] ① Vertical acceleration of the vehicle body, used to characterize the smoothness of the car's ride and the comfort of the passengers;

[0084] ② Suspension dynamic deflection is the amount of displacement change of the suspension system from the equilibrium position (static deflection position) to the maximum compression or extension position under dynamic conditions (such as when the vehicle is driving over bumpy roads or turning). Appropriate dynamic deflection helps reduce vibrations transmitted to the vehicle body and improve passenger comfort. If the dynamic deflection is too small, the vehicle may feel too "stiff" when encountering bumps, affecting the ride experience; while excessive dynamic deflection may make the vehicle feel loose and lack handling stability.

[0085] ③ Vertical displacement of the vehicle body. Changes in the vertical displacement of the vehicle body will affect the vehicle's center of gravity height and roll angle, thereby affecting ride comfort.

[0086] ④ Remaining battery power: The remaining battery power is used as a state parameter of the deep reinforcement learning network to achieve synergistic optimization of ride comfort and energy recovery;

[0087] Define the suspension state parameter set , , This represents the total number of steps. For the first The normalized state variables of the step, and have ,in, For the first The car body acceleration of the step, , The first The front and rear wheel suspension dynamic deflection of the step. For the first The vehicle body displacement of the step, For the first The remaining battery power of the step;

[0088] This invention, an energy-recovery fully active suspension, optimizes the existing passive suspension structure by placing a motor-type electromagnetic actuator inside the spring and outside the damper, with the same stroke as the spring and damper, forming a parallel structure. When the electromagnetic actuator acts as a motor and the suspension system is in active control mode, various sensors installed on the vehicle body collect signals such as vehicle displacement and acceleration and transmit them to the control system. After receiving the signals, the control system calculates the required current output to control the output of the Ampere force. The coil in the electromagnetic actuator generates the corresponding actuation force to counteract the road input excitation resistance, thereby suppressing vehicle vibration. When the electromagnetic actuator operates as a generator and the suspension system is in energy recovery mode, the actuator coil moves with the suspension, cutting magnetic field lines and generating current. This current is converted from alternating current to direct current through the energy recovery circuit and stored in the energy storage device, realizing the process of converting the mechanical energy of vibration into electrical energy.

[0089] The four-quadrant operation principle of the voice coil motor is as follows: Figure 2 As shown, the x-axis represents the direction of motor torque, and the y-axis represents the direction of rotor velocity. For a rotating motor, this represents torque and speed. The motor's torque and speed form an xoy plane coordinate system. When the motor torque and speed are both positive, the motor operates in the first quadrant, in forward rotation. Similarly, when the motor operates in the third quadrant, it operates in reverse rotation. When the motor torque is negative and the speed is positive, the motor operates in the second quadrant, in forward rotation. Similarly, when the motor operates in the fourth quadrant, it operates in reverse rotation. Vibration energy recovery utilizes the four-quadrant operation and regenerative braking principle of the motor. When the thrust of the electromagnetic actuator is in the same direction as the relative velocity between the mover and stator (the difference between the speed of the sprung mass and the speed of the unsprung mass), the motor is operating in the motoring state (quadrants I and III). The electromagnetic actuator dissipates electrical energy, and its actuation force simultaneously suppresses suspension vibration. When the thrust of the electromagnetic actuator is in the opposite direction to its velocity, the motor is operating in the generating state (quadrants II and IV). The suspension does work on the motor, and electrical energy flows to the power source through the actuator. Vibration energy is converted into electrical energy and collected by the energy storage section.

[0090] This energy-recovery fully active suspension uses a moving-coil voice coil linear motor as the electromagnetic actuator. This type of voice coil motor is a single-phase, two-pole device. Its drive method can be combined with an H-bridge circuit integrating inverter and rectifier. The drive circuit connects the motor to the DC-DC circuit, achieving four-quadrant operation through the inverter H-bridge drive circuit. This allows the motor to switch operating quadrants according to working conditions, realizing active control of the actuator and energy recovery. The recovered energy is stored in an energy storage module using supercapacitors as energy storage elements through the DC-DC control circuit. The principle is as follows: Figure 3 As shown.

[0091] The suspension stiffness, damping coefficient, motor current, and PWM duty cycle of the H-bridge in the energy recovery circuit are selected as the action space of the deep reinforcement learning network: the network outputs the suspension stiffness, damping coefficient, and motor current to control the active control force output by the suspension; and the network outputs the PWM duty cycle of the H-bridge in the energy recovery circuit to control the energy recovery efficiency.

[0092] Define the motion space set of the suspension ,in, Here, is the suspension stiffness, and is the suspension damping coefficient. , The magnitude of the external current connected to the motor. This refers to the PWM duty cycle of the H-bridge in the energy recovery circuit.

[0093] Step 3: Construct the first equation using equation (2). reward number of steps ;

[0094] (2)

[0095] In equation (2), Let represent the energy recovery reward value at step i, and we have:

[0096] (3)

[0097] In equation (3), Indicates the battery capacity threshold. Let represent the energy recovery efficiency at step i, and we have:

[0098] (4)

[0099] (5)

[0100] In equation (4), The battery level at step i. This is the maximum usable battery capacity; when the remaining battery capacity is >90%, the energy recovery circuit opens and energy recovery stops.

[0101] In equation (5), To recover energy for the suspension in step i, This represents the total energy absorbed by the suspension in step i.

[0102] Among them, negative rewards are given to the square values ​​of four state variables: vehicle acceleration, front and rear wheel suspension deflection, and vehicle displacement, respectively, to penalize states that affect ride comfort; the design of the energy recovery reward value enables the control method to prioritize energy recovery at low SOC and reduce energy recovery efficiency at high SOC, so as to achieve synergistic optimization of energy recovery and ride comfort.

[0103] Step 4: Construct a deep reinforcement learning network, including an input layer with 5 neurons. The network consists of a fully connected hidden layer, an action network with an output layer of 4 neurons, and an input layer of 5 neurons. A value network consists of a fully connected layer of hidden layers and an output layer with one neuron; the number of neurons in the input layer of the action network and the value network corresponds to the number of state parameters, and the number of neurons in the output layer of the action network corresponds to the number of action parameters.

[0104] Step 5: The input is processed in the action network, and the i-th step output value of the action network is obtained using equation (6). ;

[0105] (6)

[0106] In equation (5), These are the parameters of the action network; Indicates the first The output of the fully connected layer The corresponding activation value, Tanh is the output function, This is the bias matrix of the output layer. Let be the bias vector of the output layer, and use equation (7) to calculate the output of the first fully connected layer. Corresponding activation value The output of the k-th fully connected layer Corresponding activation value :

[0107] (7)

[0108] In equation (6), The bias matrix of the first fully connected layer, Let be the bias matrix of the k-th fully connected layer, and ReLU be the activation function. This is the bias vector of the first fully connected layer. Let K be the bias vector of the k-th fully connected layer. .

[0109] Using equation (8), the suspension at the first... Step movement volume ,in, For the first The suspension stiffness of the step, For the first The suspension damping coefficient of the step, For the first The magnitude of the external current connected to the stepper motor; For the first The PWM duty cycle of the H-bridge in the energy recovery circuit of the step;

[0110] (8)

[0111] In equation (8), Indicates the first Random noise in the step.

[0112] The simulation dynamic model executes the first Step movement volume After that, the first +1 step status Thus constituting the first Sample, denoted as This process yields m samples, which are then stored in the sample pool D.

[0113] Step 6, as follows Figure 4 The diagram shown illustrates the principle of a deep reinforcement learning algorithm based on the PPO network. Using this principle and a sample pool D, the deep reinforcement learning network is trained on simulated road surface models of grades A, B, and C under three different operating conditions: low speed, medium speed, and high speed, respectively, to obtain the optimal deep reinforcement learning model.

[0114] Step 6.1: Calculate the first step using equation (9). Priority weight of each sample :

[0115] (9)

[0116] In equation (8), , , , There are 4 weighting coefficients; through dynamic weighting It enhances the learning efficiency for critical states such as large-amplitude vibrations and high-energy recovery scenarios, which is different from traditional uniform sampling.

[0117] If the first Vertical acceleration of the car body Exceeding the comfort threshold or the Front and rear wheel suspension dynamic deflection Approaching the mechanical limit or the first Step's battery power If the protection limit is reached, it will increase. The value of ; otherwise, It remains unchanged.

[0118] Step 6.2: Calculate the first step using equation (10). The probability of a sample being selected Thus, the probability of all samples being selected is obtained;

[0119] (10)

[0120] In equation (10), α is the priority intensity coefficient that controls the degree of sampling bias; Let be the priority weight of the k-th sample.

[0121] Step 6.3: Define the current iteration number as c, and initialize c=1; define the update frequency as... The maximum number of iterations is Define the parameters of the action network in the c-th iteration as follows: Define the environment interaction parameters of the action network in the c-th iteration as follows: Let the parameters of the value network in the c-th iteration be denoted as... ;

[0122] Step 6.4: Calculate the value of the first iteration using equation (11) in the c-th iteration. The probability ratio of the strategy of each sample :

[0123] (11)

[0124] In equation (10), Let be the policy output by the action network in the c-th iteration. The environment interaction strategy output by the action network in the c-th iteration; The strategy for the c-th iteration In state Select action at time The probability of; The strategy for the c-th iteration In state Select action at time The probability of.

[0125] Based on the probability that all samples are selected, a batch of samples is drawn from the sample pool D for the c-th iteration. Based on a batch of samples from the c-th iteration Using equation (12), the priority-weighted action objective function of the action network in the c-th iteration is constructed. :

[0126] (12)

[0127] In equation (12), The cropping threshold, For the first The expected value of a sample in the c-th iteration. The pruning function is updated using weighted gradients, prioritizing the optimization of samples that significantly impact suspension performance (such as those experiencing severe vibrations or high energy recovery scenarios). Indicates the first The advantage function of the sample in the c-th iteration is:

[0128] (13)

[0129] In equation (13), , This indicates that the value network in the c-th iteration respectively... and The predicted value.

[0130] Step 6.5: Construct the priority-weighted value function loss of the value network in the c-th iteration using equation (14). :

[0131] (14)

[0132] In equation (13), Indicates the discount factor. , This indicates that the value network in the c-th iteration respectively... and The predicted value.

[0133] Step 6.6: For each , Perform the update to obtain the parameters of the policy network at step c+1. And the parameters of the value network at step c+1 Otherwise, no update. , :

[0134] (15)

[0135] In equation (15), For learning rate, For policy network parameters The gradient operator;

[0136] (16)

[0137] In equation (16), Parameters of the value network The gradient operator.

[0138] Step 6.7: [The sentence is incomplete and requires more context to be translated accurately.] Assign to Then, make a judgment Check if the condition is met. If it is met, it means that the optimal deep reinforcement learning model has been obtained; otherwise, return to step 6.4 and execute sequentially.

[0139] Step 7: Use the optimal action network in the optimal deep reinforcement learning model to process the real-time input state information and output corresponding control commands to achieve active control of the energy recovery fully active suspension.

[0140] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0141] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

Claims

1. A fully active suspension control method for energy recovery based on deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Construct a simulation dynamics model and a stochastic road surface model for the energy recovery 1 / 2 vehicle's four-degree-of-freedom fully active suspension; Step 2: Define the suspension state parameter set , , This represents the total number of steps. For the first The normalized state variables of the step, and have ,in, For the first The car body acceleration of the step, , The first The front and rear wheel suspension dynamic deflection of the step. For the first The vehicle body displacement of the step, For the first The remaining battery power of the step; Define the motion space set of the suspension ,in, For suspension stiffness, This is the suspension damping coefficient. The magnitude of the external current connected to the motor. This refers to the PWM duty cycle of the H-bridge in the energy recovery circuit. Step 3, construct the first reward number of steps ; Step 4: Construct a deep reinforcement learning network, including an input layer with 5 neurons. The network consists of a fully connected hidden layer, an action network with an output layer of 4 neurons, and an input layer of 5 neurons. A value network consisting of a fully connected layer of hidden layers and an output layer of one neuron; Step 5: The input is processed in the motion network to obtain the suspension in the first... Step movement volume The simulation dynamics model executes the first... Step movement volume After that, the first +1 step status Thus constituting the first Sample, denoted as This yields m samples, which are then stored in the sample pool D. Step 6: Based on the sample pool D, train the deep reinforcement learning network using the PPO algorithm to obtain the optimal deep reinforcement learning model; Step 7: Use the optimal action network in the optimal deep reinforcement learning model to process the real-time input state information and output corresponding control commands to achieve active control of the energy recovery fully active suspension.

2. The energy recovery fully active suspension control method based on deep reinforcement learning according to claim 1, characterized in that, In step 3, the first equation is constructed using equation (1). reward number of steps ; (1) In equation (1), Let represent the energy recovery reward value at step i, and we have: (2) In equation (2), Indicates the battery capacity threshold. Let represent the energy recovery efficiency at step i, and we have: (3) (4) In equation (3), The battery level at step i. This is the maximum usable battery capacity; when the remaining battery capacity is >90%, the energy recovery circuit opens and energy recovery stops. In equation (4), To recover energy for the suspension in step i, Let be the total energy absorbed by the suspension in step i.

3. The energy recovery fully active suspension control method based on deep reinforcement learning according to claim 2, characterized in that, In step 5, the i-th step output value of the action network is obtained using equation (5). ; (5) In equation (5), These are the parameters of the action network; Indicates the first The output of the fully connected layer The corresponding activation value, Tanh is the output function, This is the bias matrix of the output layer. Let be the bias vector of the output layer, and use equation (6) to calculate the output of the first fully connected layer. Corresponding activation value The output of the k-th fully connected layer Corresponding activation value : (6) In equation (6), The bias matrix of the first fully connected layer, Let be the bias matrix of the k-th fully connected layer, and ReLU be the activation function. This is the bias vector of the first fully connected layer. Let K be the bias vector of the k-th fully connected layer. ; Using equation (7), the suspension at the first... Step movement volume ,in, For the first The suspension stiffness of the step, For the first The suspension damping coefficient of the step, For the first The magnitude of the external current connected to the stepper motor; For the first The PWM duty cycle of the H-bridge in the energy recovery circuit of the step; (7) In equation (8), Indicates the first Random noise in the step.

4. The energy recovery fully active suspension control method based on deep reinforcement learning according to claim 3, characterized in that, Step 6 includes the following steps: Step 6.1: Calculate the first step using equation (8). Priority weight of each sample : (8) In equation (8), , , , There are 4 weighting coefficients; If the first Vertical acceleration of the car body Exceeding the comfort threshold or the Front and rear wheel suspension dynamic deflection Approaching the mechanical limit or the first Step's battery power If the protection limit is reached, it will increase. The value of ; otherwise, Remain unchanged; Step 6.2: Calculate the first step using equation (9). The probability of a sample being selected Thus, the probability of all samples being selected is obtained; (9) In equation (9), α is the priority intensity coefficient that controls the degree of sampling bias; The priority weight of the k-th sample; Step 6.3: Define the current iteration number as c, and initialize c=1; define the update frequency as... The maximum number of iterations is Define the parameters of the action network in the c-th iteration as follows: Define the environment interaction parameters of the action network in the c-th iteration as follows: Let the parameters of the value network in the c-th iteration be denoted as... ; Step 6.4: Calculate the value of the first iteration in the c-th iteration using equation (10). The probability ratio of the strategy for each sample : (10) In equation (10), Let be the policy output by the action network in the c-th iteration. The environment interaction strategy output by the action network in the c-th iteration; The strategy for the c-th iteration In state Select action at time The probability of; The strategy for the c-th iteration In state Select action at time The probability of; Based on the probability that all samples are selected, a batch of samples is drawn from the sample pool D for the c-th iteration. Based on a batch of samples from the c-th iteration Using equation (11), the priority-weighted action objective function of the action network in the c-th iteration is constructed. : (11) In equation (11), The cropping threshold, For the first The expected value of a sample in the c-th iteration. For the clipping function, Indicates the first The advantage function of the sample in the c-th iteration is: (12) In equation (12), , This indicates that the value network in the c-th iteration respectively... and The predicted value; Step 6.5: Construct the priority-weighted value function loss of the value network in the c-th iteration using equation (13). : (13) In equation (13), Indicates the discount factor. , This indicates that the value network in the c-th iteration respectively... and The predicted value; Step 6.6: If and When the remainder of the ratio is 1, use equations (14) and (15) respectively to... , Update the parameters of the policy network at step c+1. And the parameters of the value network at step c+1 Otherwise, no update. , : (14) In equation (14), For learning rate, For policy network parameters The gradient operator; (15) In equation (15), Parameters of the value network The gradient operator; Step 6.7: [The sentence is incomplete and requires more context to be translated accurately.] Assign to Then, make a judgment If the condition is met, the optimal deep reinforcement learning model has been obtained; otherwise, return to step 6.4 and execute sequentially.

5. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store programs that support the processor in executing any of the energy recovery fully active suspension control methods of claims 1-4, and the processor is configured to execute the programs stored in the memory.

6. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the energy recovery fully active suspension control method according to any one of claims 1-4.

Citation Information

Patent Citations

  • A hybrid electromagnetic suspension capable of realizing self-power supply and a control method thereof

    CN109080399A

  • Active control method and system based on variable suspension

    CN118832998A