A Reinforcement Learning-Based Semi-Active Suspension Intelligent Control Method Based on Bayesian Optimization

By constructing a suspension intelligent control strategy based on a Bayesian optimization-based reinforcement learning method, a nonparametric reward function and a three-layer deep neural network are used. This solves the multi-objective optimization problem of the suspension system under complex working conditions, improves the vehicle's handling stability and ride comfort, and solves the hyperparameter sensitivity problem of the suspension system.

CN120792407BActive Publication Date: 2025-12-02JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511261109.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-12-02
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing suspension system control strategies are unable to effectively cope with nonlinear characteristics and complex operating conditions, resulting in insufficient adaptability under multiple operating conditions. Furthermore, deep reinforcement learning methods suffer from hyperparameter sensitivity issues, making it difficult to achieve multi-objective collaborative optimization of vehicle safety, ride comfort, and handling stability.

Method used

A reinforcement learning method based on Bayesian optimization is adopted to construct an intelligent suspension control strategy through a nonparametric reward function and a three-layer deep neural network. The PPO algorithm and Bayesian optimization are combined to automatically search for hyperparameters and optimize the control strategy of the suspension system.

Benefits of technology

It achieves multi-objective collaborative optimization of the suspension system under complex working conditions, improves vehicle handling stability and ride comfort, reduces the randomness and inefficiency of hyperparameter tuning, and improves the adaptability and efficiency of the control strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120792407B_ABST
    Figure CN120792407B_ABST
Patent Text Reader

Abstract

This invention relates to the field of vehicle intelligent control technology and provides a reinforcement learning-based semi-active suspension intelligent control method based on Bayesian optimization. This method collects and processes vehicle state data, then uses the Proximal Policy Optimization (PPO) algorithm combined with a nonparametric reward function to design a deep reinforcement learning model. The optimal hyperparameter set of the model is searched using the Bayesian optimization algorithm, and the optimized hyperparameters are used to train the model to generate a semi-active suspension intelligent control strategy. The vehicle state reward is used to evaluate the action results. This invention achieves multi-objective collaborative optimization by using a nonparametric reward function in the PPO algorithm, utilizes a three-layer deep neural network to address complex control problems in a data-driven manner, introduces Bayesian optimization to automatically search for hyperparameters to improve efficiency and performance, and designs the optimization objective function as a combination of trapezoidal numerical integral and simple moving average to stabilize the reward value and accelerate convergence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of vehicle intelligent control technology, and in particular relates to a reinforcement learning-based semi-active suspension intelligent control method based on Bayesian optimization. Background Technology

[0002] With the continuous advancement of automotive technology, the boundaries of vehicle performance, comfort, and safety are constantly expanding, and the suspension system, as the key interface connecting the vehicle frame and wheels, plays a central role in this process. It mitigates the impact of uneven road surfaces and enhances vehicle handling, ensuring a smooth and stable driving experience. As the demand for improved vehicle performance increases, ride comfort and safety, directly related to the suspension system, are receiving increasing attention. Semi-active suspension systems strike a balance between the simplicity of passive suspension systems and the control adaptability of active suspension systems. Through intelligent control strategies, they adjust damping force according to real-time driving conditions, improving ride comfort and handling stability without consuming as much energy as active suspension systems. Their control effect is close to that of active suspension systems, and they have become one of the common technologies in intelligent vehicles.

[0003] Currently, the research and development goals of suspension system control strategies focus on effectively suppressing vehicle vertical vibrations in various driving environments while improving handling stability to ensure driving safety. Existing control strategies are mainly divided into two categories: methods based on classical control theory, which, while easy to implement basic control of the suspension system, are all based on the assumption of a linear time-invariant system. However, actual suspension systems exhibit nonlinear characteristics, and the difference between simplified models and real systems leads to insufficient adaptability under various operating conditions. Methods based on modern control theory, while capable of implementing more complex control logic, rely on complex algorithms and models. Even using high-order mathematical methods, it is difficult to fully describe the true dynamics of the suspension system, such as damping hysteresis, component wear, and inertial parameter fluctuations caused by changes in vehicle load. Therefore, they are unable to cope with diverse and complex operating conditions, limiting their application scope.

[0004] In recent years, the explosive development of artificial intelligence technology has brought innovation to the field of intelligent driving. Among them, the breakthrough of deep reinforcement learning technology has provided a new solution for optimizing the control strategy of semi-active suspension. However, the efficiency of deep reinforcement learning is highly dependent on the configuration of hyperparameters such as learning rate, neural network size, and discount factor, and suffers from hyperparameter sensitivity. To address this, this invention proposes a reinforcement learning-based intelligent control method for semi-active suspension based on Bayesian optimization. Summary of the Invention

[0005] The purpose of this invention is to provide a reinforcement learning-based semi-active suspension intelligent control method based on Bayesian optimization, which aims to solve the problems mentioned in the background art.

[0006] The objective of this invention is achieved through the following technical solution:

[0007] A reinforcement learning-based semi-active suspension intelligent control method based on Bayesian optimization includes the following steps:

[0008] Step 1: Data Acquisition and Processing;

[0009] Vehicle status data is collected in real time through a sensor network deployed on the vehicle, and the data from different sensors on the vehicle is preprocessed using a preset method.

[0010] Step 2: Design of a deep reinforcement learning model based on a nonparametric reward function;

[0011] A deep reinforcement learning model is constructed based on the proximal policy optimization algorithm; the preprocessed vehicle state data is used as the state space, the control current output of the semi-active suspension is used as the action space, and a nonparametric reward function is designed; the vehicle state data includes the vehicle vertical acceleration, vehicle roll angle, vehicle roll rate, and vehicle yaw rate.

[0012] Step 3: Parameter optimization of deep reinforcement learning model based on Bayesian optimization algorithm;

[0013] Construct a set of hyperparameters to be optimized for the near-end policy optimization algorithm and search for them using the Bayesian optimization method;

[0014] Step 4: Training a deep reinforcement learning model based on a parameterless reward function optimized by Bayes and generating a semi-active suspension intelligent control strategy;

[0015] Initialize the experience buffer, policy network parameters, and value network parameters. Use the hyperparameter set obtained by Bayesian optimization to train the deep reinforcement learning model, and obtain a parameter-free reward function-based deep reinforcement learning semi-active suspension intelligent control strategy.

[0016] Furthermore, in step 1, the data collected by the sensor is sixteen-dimensional data, including the vehicle's vertical acceleration. Vehicle roll angle Vehicle roll rate Vehicle yaw rate Four sets of suspension dynamic deflection Four sets of tire travel The relative displacement velocity of the four suspension sets Preprocess the multidimensional dynamic data collected by the sensors to ensure data accuracy and quality, establish a sensor type-physical quantity dimension mapping table, perform dynamic unit system conversion on the raw output data of each sensor, and unify the measurement standard.

[0017] Furthermore, step 2 includes the following sub-steps:

[0018] Step 2.1: Selection of state observations;

[0019] Define the state observations of the PPO algorithm as follows: ,in, express State observations at any given time, including the vehicle's vertical acceleration. Vehicle roll angle Vehicle roll rate Vehicle yaw rate ;

[0020] Step 2.2: Selection of motion intensity;

[0021] Define the amount of motion as ,in express The amount of motion at any given moment, i.e., the control current acting on the semi-active suspension. ;

[0022] Step 2.3: Setting the nonparametric reward function;

[0023] Define the nonparametric reward function: ,in, express The reward function of the environment at any given time; the vertical acceleration of the vehicle body included in the reward function. Used to stabilize vibrations in the vertical direction of the vehicle body; vehicle roll angle Used to ensure the vehicle's handling stability;

[0024] Step 2.4: Suspension physical limitations;

[0025] The constraints are that the suspension dynamic deflection does not exceed the maximum travel and the dynamic load of the vehicle tires is less than the static load;

[0026] Step 2.5: Network architecture setup;

[0027] The policy network and the value network use the same three-layer deep neural network, with 100 neurons in each layer, and the activation function is a linear rectified function.

[0028] Furthermore, step 3 includes the following sub-steps:

[0029] Step 3.1: Construct the hyperparameter set;

[0030] Define the candidate hyperparameter set as ,in Represents the candidate hyperparameter set. For the policy network learning rate, For the value network learning rate, As a discount factor, For the experience zone, For batch sampling scale, For the number of learning rounds, Weights for entropy loss;

[0031] Step 3.2: Construct the proxy function;

[0032] Selecting Gaussian process As a surrogate model, it is used to perform posterior prediction of the objective function, where the input and output are candidate points, respectively. mean of the objective function and standard deviation ;

[0033] Step 3.3: Select the acquisition function;

[0034] choose As a data acquisition function, it guides the optimization process's exploration behavior in the hyperparameter space; among which... Represents the desired improvement function; Represents the posterior distribution of a Gaussian process; Represents the expected value; This indicates the current optimal GP function value;

[0035] Step 3.4: Construct the optimization objective function;

[0036] The objective function is defined as the trapezoidal numerical integral based on the average reward of deep reinforcement learning, and its expression is: ;in, Represents the trapezoidal numerical integral function; This indicates the use of candidate hyperparameter sets during deep reinforcement learning training. The simple moving average of the post-system reward; This is a fixed incremental parameter used to shift the overall reward to the positive range; This represents the time window size for the moving average. To use candidate hyperparameter sets during deep reinforcement learning training t Time-based system rewards.

[0037] Furthermore, step 4 includes the following sub-steps:

[0038] Step 4.1: Initialize the optimal hyperparameter set obtained from the Bayesian optimization search;

[0039] Step 4.2: Initialize the experience buffer to store the trajectory data of the agent's interaction with the environment, including state, action, reward, next state, and termination signal.

[0040] Step 4.3: Randomly initialize policy network parameters With value network parameters ;

[0041] Step 4.4: Set the number of training rounds for the PPO algorithm and the maximum number of training steps in each round. T ;

[0042] Step 4.5: Perform training, using the current policy network to interact with the environment and execute the maximum number of steps in each training round. T ,exist Observe the current system status at all times According to the policy network The sampling action controls the current, and the damping force is obtained from the current and the compression speed of the shock absorber. The system then generates... Momentary reward feedback and the state in the next moment ; use the system's empirical data The advantage value is stored in the experience buffer and calculated using the generalized advantage estimation method. ;

[0043] Step 4.6: Process the collected trajectory data. In each round of optimization and updates, the trajectory data is randomly shuffled and adjusted according to the batch sampling size. The size is divided; the old policy network is used to calculate the action probability. Calculate the strategy ratio Then, calculate the objective function with clipping. ,in This is the cutting factor. The function restricts policy updates by pruning importance weights, and its specific definition is as follows:

[0044] ;

[0045] After one round of updates is completed, the old policy network parameters will be updated to the current policy network parameters;

[0046] Step 4.7: The value network updates by minimizing the mean squared error between the predicted state value and the actual return, defining the value loss function. ,in For the loss of value, For value network to state Value prediction;

[0047] Step 4.8: Repeat steps 4.5, 4.6, and 4.7; if the maximum number of iterations is reached, training ends; otherwise, return to step 4.5 to continue training.

[0048] Step 4.9: After step 4.8 is completed, return to step 4.2 to reinitialize the experience buffer and perform Bayesian optimization training until the maximum number of iterations of Bayesian optimization is reached. Stop training and obtain the optimal hyperparameter set. Based on the deep reinforcement learning model trained with the optimal hyperparameter set, generate a semi-active suspension intelligent control strategy.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] 1. This invention replaces traditional parametric design with a nonparametric reward function in the PPO algorithm, directly linking the vehicle's vertical acceleration (a ride comfort index) and the vehicle's roll angle (a handling stability index), reducing the impact of parameter uncertainty, enabling the reward mechanism to dynamically adapt to complex operating conditions, achieving multi-objective collaborative optimization, and solving the multi-objective collaborative optimization problem of vehicle safety, ride comfort, and handling stability.

[0051] 2. This invention employs a three-layer deep neural network to construct a policy network and a value network. It processes the vehicle state space (vehicle vertical acceleration, vehicle roll angle, etc.) and motion space (semi-active suspension control current) through a data-driven approach, thus eliminating the dependence of traditional control algorithms on high-precision mathematical modeling and effectively addressing complex control problems in large state spaces.

[0052] 3. This invention addresses the hyperparameter sensitivity issue in deep reinforcement learning by introducing Bayesian optimization to automatically search the hyperparameter set of the PPO algorithm, reducing the randomness and inefficiency of manual parameter tuning, saving time and improving performance.

[0053] 4. This invention designs the objective function as a combination of trapezoidal numerical integral based on the average reward of deep reinforcement learning and simple moving average, which helps to accelerate the model convergence process while ensuring the stability of the reward function value. Attached Figure Description

[0054] Figure 1 This is a flowchart of the method of the present invention.

[0055] Figure 2 This shows the change of roll angle over time for different suspension control strategies when the vehicle is driving at 80 km / h in a slalom course.

[0056] Figure 3 The vertical acceleration of the vehicle body corresponding to different suspension control strategies under the condition of straight-line driving on a Class B road at 50km / h. How it changes over time. Detailed Implementation

[0057] In order to provide a clearer understanding of the technical features, objectives and beneficial effects of the present invention, the technical solution of the present invention will now be described in detail below, but it should not be construed as limiting the scope of implementation of the present invention.

[0058] This invention provides a reinforcement learning-based semi-active suspension intelligent control method based on Bayesian optimization, the flowchart of which is shown below. Figure 1 As shown, the method includes the following steps:

[0059] Step 1: Data Acquisition and Processing;

[0060] Multidimensional vehicle state data is collected through sensor networks, and preprocessed and standardized to provide high-quality input for subsequent model training. Specifically, this includes:

[0061] Data Acquisition: A sensor network consisting of accelerometers, displacement sensors, and IMU modules deployed on the vehicle is used to collect 16-dimensional real-time vehicle status data, including: vertical acceleration of the vehicle body. Vehicle roll angle Vehicle roll rate Vehicle yaw rate Four sets of suspension dynamic deflection Four sets of tire travel The relative displacement speed of the four suspension sets .

[0062] Data Processing: Interpolation correction is used to preprocess data from different vehicle sensors to ensure the accuracy and applicability of the collected real-vehicle data. A sensor type-physical quantity dimension mapping table is established to dynamically convert the unit system of the raw output data from accelerometers, displacement sensors, and IMU modules, ensuring that the data has a consistent measurement standard and enhancing the adaptability of the algorithm.

[0063] Step 2: Design of a deep reinforcement learning model based on a nonparametric reward function;

[0064] A deep reinforcement learning model is constructed based on the Proximal Policy Optimization (PPO) algorithm. Preprocessed vehicle state data (vehicle vertical acceleration, vehicle roll angle, vehicle roll angular velocity, and vehicle yaw rate) are used as the state space, and the control current output of the vehicle's semi-active suspension is used as the action space. A nonparametric reward function is designed as an extension and generalization of the current mainstream parametric reward function. The core of the PPO algorithm includes a policy network (for outputting action decisions) and a value network (for evaluating state value), which together support the decision-making and optimization process of reinforcement learning.

[0065] The sub-steps of step 2 are as follows:

[0066] Step 2.1: Selection of state observations;

[0067] Considering the semi-active suspension intelligent control method focusing on vehicle handling stability and ride comfort indicators, the state observation of the PPO algorithm is defined as follows: ,in, express State observations at any given time, including the vehicle's vertical acceleration. Vehicle roll angle Vehicle roll rate Vehicle yaw rate .

[0068] Step 2.2: Selection of motion intensity;

[0069] Define the amount of motion as ,in express The amount of motion at any given moment, i.e., the control current acting on the semi-active suspension. .

[0070] Step 2.3: Setting the nonparametric reward function;

[0071] The main purpose of a semi-active suspension control system is to maintain the vehicle's vertical acceleration within a reasonable range under different driving conditions, while simultaneously improving the suppression of body roll. To this end, a non-parametric reward function is defined as follows: ,in, express The reward function of the environment at any given time; the vertical acceleration of the vehicle body included in the reward function. This is to stabilize the vertical vibration of the vehicle body, prevent road excitation from interfering with passengers, and ensure vehicle ride comfort; vehicle roll angle This is to ensure the vehicle's handling stability.

[0072] Step 2.4: Suspension physical limitations;

[0073] To avoid damaging vehicle components, two physical limitations must be considered: first, the suspension dynamic deflection must be within its maximum travel limit; second, the dynamic load on the vehicle tires should be less than the static load to ensure that the road surface and wheels maintain constant contact. When the input action exceeds the limits, only the maximum value input will be used.

[0074] Step 2.5: Network architecture setup;

[0075] The policy network and the value network use the same deep neural network structure. Considering the system complexity and real-time requirements, a three-layer network is used, with 100 neurons in each layer. The activation function is the rectified linear function (ReLU).

[0076] Step 3: Parameter optimization of deep reinforcement learning model based on Bayesian optimization algorithm;

[0077] Construct a set of hyperparameters to be optimized for the PPO algorithm and search for them using Bayesian optimization (Bayesian optimization is a model-based sequential optimization technique that can be used to tune the hyperparameters of any noisy black-box function).

[0078] The sub-steps of step 3 are as follows:

[0079] Step 3.1: Construct the hyperparameter set;

[0080] Define the candidate hyperparameter set as ,in Represents the candidate hyperparameter set. For the policy network learning rate, For the value network learning rate, As a discount factor, For the experience zone, For batch sampling scale, For the number of learning rounds, The weights are used to calculate the entropy loss.

[0081] Step 3.2: Construct the proxy function;

[0082] Selecting Gaussian process As a surrogate model, it is used to perform posterior prediction of the objective function, where the input and output are candidate points, respectively. mean of the objective function and standard deviation .

[0083] Step 3.3: Select the acquisition function;

[0084] choose As a data acquisition function, it guides the optimization process's exploration behavior in the hyperparameter space; among which... Represents the desired improvement function; Represents the posterior distribution of a Gaussian process; The expected value is represented by ; while the best GP function value obtained so far is denoted as . .

[0085] Step 3.4: Construct the optimization objective function;

[0086] The objective function is defined as the trapezoidal numerical integral based on the average reward of deep reinforcement learning, and its expression is: ;in, Represents the trapezoidal numerical integral function; This indicates the use of candidate hyperparameter sets during deep reinforcement learning training. The simple moving average of the post-system reward; This is a fixed incremental parameter used to shift the overall reward to the positive range; This represents the time window size for the moving average. To use candidate hyperparameter sets during deep reinforcement learning training t Time-based system rewards.

[0087] Step 4: Training a deep reinforcement learning model based on a parameterless reward function optimized by Bayes and generating a semi-active suspension intelligent control strategy;

[0088] Initialize the experience buffer, policy network parameters, and value network parameters. Use the hyperparameter set obtained by Bayesian optimization to train the deep reinforcement learning model, and obtain a parameter-free reward function-based deep reinforcement learning semi-active suspension intelligent control strategy.

[0089] The sub-steps of step 4 are as follows:

[0090] Step 4.1: Initialize the optimal hyperparameter set obtained from the Bayesian optimization search;

[0091] Step 4.2: Initialize the experience buffer to store trajectory data such as state, action, reward, next state and termination signal generated by the agent's interaction with the environment;

[0092] Step 4.3: Randomly initialize policy network parameters With value network parameters ;

[0093] Step 4.4: Set the number of training rounds for the PPO algorithm and the maximum number of training steps in each round. T ;

[0094] Step 4.5: Perform training, using the current policy network to interact with the environment and execute the maximum number of steps in each training round. T ,exist Observe the current system status at all times According to the policy network The sampling action controls the current, and the damping force is obtained by inputting the current and the compression speed of the shock absorber (obtained by looking up a table according to the suspension characteristics). The system then generates... Momentary reward feedback and the state in the next moment . The system's empirical data The advantage value is stored in the experience cache. The advantage value is calculated using the generalized advantage estimation method. This is used for subsequent updates to the policy network.

[0095] Step 4.6: Process the collected trajectory data. In each round of optimization and updates, the trajectory data is randomly shuffled and adjusted according to the batch sampling scale. Divide the data into segments based on size. Calculate action probabilities using the old policy network. Calculate the strategy ratio Then, calculate the objective function with clipping. ,in This is the cutting factor. The function limits policy updates to a "reasonable range" (not exceeding) by pruning importance weights. No less than Within this range, it is ensured that the difference between the new strategy and the old strategy is not too large, thereby improving sample utilization efficiency and training stability. Its specific definition is as follows:

[0096] ;

[0097] After one round of updates is completed, the old policy network parameters will be updated to the current policy network parameters.

[0098] Step 4.7: The value network updates by minimizing the mean squared error between the predicted state value and the actual return, defining the value loss function. ,in For the loss of value, For value network to state Value prediction.

[0099] Step 4.8: Repeat steps 4.5, 4.6 and 4.7; if the maximum number of iterations is reached, training ends; otherwise, return to step 4.5 to continue training.

[0100] Step 4.9: After step 4.8 is completed, return to step 4.2 to reinitialize the experience buffer and perform Bayesian optimization training until the maximum number of iterations of Bayesian optimization is reached. Stop training and obtain the optimal hyperparameter set. Based on the deep reinforcement learning model trained with the optimal hyperparameter set, generate a semi-active suspension intelligent control strategy.

[0101] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.

[0102] Example 1: The comprehensive control effect on handling stability and ride comfort was experimentally verified under the following conditions: 80 km / h slalom driving and 50 km / h straight driving on a Class B road. The Bayesian-optimized parameterless reward function reinforcement learning algorithm (the method proposed in this invention, BO-NRPPO) was compared with passive suspension, the parameterized reward function reinforcement learning algorithm (PRPPO), the Bayesian-optimized parameterized reward function reinforcement learning algorithm (BO-PRPPO), and the parameterless reward function reinforcement learning algorithm (NRPPO).

[0103] from Figure 2 As can be seen, the BO-NRPPO strategy performs best among all the comparison strategies and passive suspensions, effectively suppressing the peak change in roll angle. Its roll angle suppression effect is significantly better than other control methods. Specifically, compared with passive suspension, BO-NRPPO improves roll control performance by 9.6%; compared with BO-PRPPO, NRPPO, and PRPPO strategies, it improves by 3.09%, 15.6%, and 15.93%, respectively, demonstrating the significant advantage of the proposed method in terms of handling stability.

[0104] from Figure 3 As can be seen, in the ride comfort evaluation, although the root mean square value of acceleration of BO-NRPPO did not reach the lowest, its performance was improved by 9.75% compared to BO-PRPPO; compared to passive suspension, the improvement reached 32.29%. It is worth noting that although BO-NRPPO may not be the optimal strategy in terms of ride comfort, NRPPO is significantly inferior to BO-NRPPO in handling stability, even falling below passive suspension. This indicates that strategies without Bayesian optimization technology struggle to achieve an effective balance between the two objectives of ride comfort and handling stability. BO technology, by systematically exploring the hyperparameter space and combining it with the performance evaluation of the PPO algorithm, optimizes the configuration of key hyperparameters, thereby significantly improving the overall control performance of the BO-NRPPO strategy.

[0105] Based on a comprehensive analysis of various performance indicators, the BO-NRPPO strategy proposed in this invention demonstrates strong overall optimization and practical application potential under diverse driving conditions.

[0106] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the implementation of the present invention or the practicality of the patent.

Claims

1. A reinforcement learning-based semi-active intelligent suspension control method based on Bayesian optimization, characterized in that, Includes the following steps: Step 1: Data Acquisition and Processing; Vehicle status data is collected in real time through a sensor network deployed on the vehicle, and the data from different sensors on the vehicle is preprocessed using a preset method. Step 2: Design of a deep reinforcement learning model based on a nonparametric reward function; A deep reinforcement learning model is built based on the proximal policy optimization algorithm; The preprocessed vehicle state data is used as the state space, and the control current output of the semi-active suspension is used as the action space. A non-parametric reward function is designed. The vehicle state data includes the vehicle vertical acceleration, vehicle roll angle, vehicle roll rate, and vehicle yaw rate. Step 3: Parameter optimization of deep reinforcement learning model based on Bayesian optimization algorithm; Construct a set of hyperparameters to be optimized for the near-end policy optimization algorithm and search for them using the Bayesian optimization method; Step 4: Training a deep reinforcement learning model based on a parameterless reward function optimized by Bayes and generating a semi-active suspension intelligent control strategy; Initialize the experience buffer, policy network parameters, and value network parameters. Use the hyperparameter set obtained by Bayesian optimization to train the deep reinforcement learning model and obtain a parameter-free reward function-based deep reinforcement learning semi-active suspension intelligent control strategy. Step 2 includes the following sub-steps: Step 2.1: Selection of state observations; Define the state observations of the near-end policy optimization algorithm as follows: ,in, express State observations at any given time, including the vehicle's vertical acceleration. Vehicle roll angle Vehicle roll rate Vehicle yaw rate ; Step 2.2: Selection of motion intensity; Define the amount of motion as ,in express The amount of motion at any given moment, i.e., the control current acting on the semi-active suspension. ; Step 2.3: Setting the nonparametric reward function; Define the nonparametric reward function: ,in, express The reward function of the environment at any given time; the vertical acceleration of the vehicle body included in the reward function. Used to stabilize vibrations in the vertical direction of the vehicle body; vehicle roll angle Used to ensure the vehicle's handling stability; Step 2.4: Suspension physical limitations; The constraints are that the suspension dynamic deflection does not exceed the maximum travel and the dynamic load of the vehicle tires is less than the static load; Step 2.5: Network architecture setup; The policy network and the value network use the same three-layer deep neural network, with 100 neurons in each layer, and the activation function is a linear rectified function.

2. The reinforcement learning-based semi-active suspension intelligent control method based on Bayesian optimization according to claim 1, characterized in that, In step 1, the data collected by the sensor is 16-dimensional data, including the vehicle's vertical acceleration. Vehicle roll angle Vehicle roll rate Vehicle yaw rate Four sets of suspension dynamic deflection Four sets of tire travel The relative displacement velocity of the four suspension sets Preprocess the multidimensional dynamic data collected by the sensors to ensure data accuracy and quality, establish a sensor type-physical quantity dimension mapping table, perform dynamic unit system conversion on the raw output data of each sensor, and unify the measurement standard.

3. The reinforcement learning-based semi-active suspension intelligent control method based on Bayesian optimization according to claim 1, characterized in that, Step 3 includes the following sub-steps: Step 3.1: Construct the hyperparameter set; Define the candidate hyperparameter set as ,in Represents the candidate hyperparameter set. For the policy network learning rate, For the value network learning rate, As a discount factor, For the experience zone, For batch sampling scale, For the number of learning rounds, Weights for entropy loss; Step 3.2: Construct the proxy function; Selecting Gaussian process As a surrogate model, it is used to perform posterior prediction of the objective function, where the input and output are candidate points, respectively. mean of the objective function and standard deviation ; Step 3.3: Select the acquisition function; choose As a data acquisition function, it guides the optimization process's exploration behavior in the hyperparameter space; among which... Represents the desired improvement function; Represents the posterior distribution of a Gaussian process; Represents the expected value; This indicates the current optimal GP function value; Step 3.4: Construct the optimization objective function; The objective function is defined as the trapezoidal numerical integral based on the average reward of deep reinforcement learning, and its expression is: ;in, Represents the trapezoidal numerical integral function; This indicates the use of candidate hyperparameter sets during deep reinforcement learning training. The simple moving average of the post-system reward; This is a fixed incremental parameter used to shift the overall reward to the positive range; This represents the time window size for the moving average. To determine the system reward at time t after using a candidate hyperparameter set during deep reinforcement learning training.

4. The reinforcement learning-based semi-active suspension intelligent control method based on Bayesian optimization according to claim 1, characterized in that, Step 4 includes the following sub-steps: Step 4.1: Initialize the optimal hyperparameter set obtained from the Bayesian optimization search; Step 4.2: Initialize the experience buffer to store the trajectory data of the agent's interaction with the environment, including state, action, reward, next state, and termination signal. Step 4.3: Randomly initialize policy network parameters With value network parameters ; Step 4.4: Set the number of training rounds for the near-end policy optimization algorithm And the maximum training step size T in each round; Step 4.5: Perform training, using the current policy network to interact with the environment in each training round to execute the maximum step size T. Observe the current system status at all times According to the policy network The sampling action controls the current, and the damping force is obtained from the current and the compression speed of the shock absorber. The system then generates... Momentary reward feedback and the state at the next moment ; use the system's empirical data The advantage value is stored in the experience buffer and calculated using the generalized advantage estimation method. ; Step 4.6: Process the collected trajectory data. In each round of optimization and updates, the trajectory data is randomly shuffled and adjusted according to the batch sampling size. The size is divided; the old policy network is used to calculate the action probability. Calculate the strategy ratio Then, calculate the objective function with clipping. ,in This is the cutting factor. The function restricts policy updates by pruning importance weights, and its specific definition is as follows: ; After one round of updates is completed, the old policy network parameters will be updated to the current policy network parameters; Step 4.7: The value network updates by minimizing the mean squared error between the predicted state value and the actual return, defining the value loss function. ,in For the loss of value, For value network to state Value prediction; Step 4.8: Repeat steps 4.5, 4.6, and 4.7; training ends when the maximum number of iterations is reached. Otherwise, return to step 4.5 and continue training; Step 4.9: After step 4.8 is completed, return to step 4.2 to reinitialize the experience buffer and perform Bayesian optimization training until the maximum number of iterations of Bayesian optimization is reached. Stop training and obtain the optimal hyperparameter set. Based on the deep reinforcement learning model trained with the optimal hyperparameter set, generate a semi-active suspension intelligent control strategy.

Citation Information

Patent Citations

  • Automatic-driving intelligent vehicle trajectory tracking control strategy based on deep reinforcement learning

    CN110322017A

  • Vehicle driving cost evaluation method based on data driving scene

    CN113034210A