Reinforcement learning-MPC hybrid optimization tractor bumping test bench cooperative control method

Through the collaborative control method of reinforcement learning-MPC hybrid optimization, the problem that traditional control methods are difficult to achieve vibration tracking accuracy and load response speed under complex road conditions is solved, and the coordinated improvement of vibration tracking accuracy, load response speed and suspension damping stability is achieved, and the robustness of the system is enhanced.

CN120215266APending Publication Date: 2025-06-27HENAN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510353758.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The traditional tractor bump test bench control method is difficult to achieve synergistic improvements in vibration amplitude and frequency tracking accuracy, load dynamic response speed and suspension damping stability under complex road conditions, and the system has poor robustness and fault tolerance during extreme loads and complex road conditions.

Method used

The collaborative control method of reinforcement learning-MPC hybrid optimization is adopted. By establishing a coupled nonlinear state space model, a hierarchical MPC framework is designed, and the weight matrix of the objective function is optimized by combining the reinforcement learning algorithm to generate the optimal control sequence that meets the vibration amplitude and frequency constraints, and a dual-channel adaptive PID controller is built to enhance the robustness of the system.

Benefits of technology

The coordinated improvement of vibration amplitude and frequency tracking accuracy, load dynamic response speed and suspension damping stability is achieved, which enhances the system's robustness to complex road conditions and extreme loads, and ensures the accuracy of experimental data and the stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215266A_ABST
    Figure CN120215266A_ABST
Patent Text Reader

Abstract

The invention provides a reinforcement learning-MPC hybrid optimization tractor bumping test bed control method and system. The method comprises the following steps: establishing a coupled nonlinear state space model of a tractor bumping test bed system, designing a layered MPC framework based on the coupled nonlinear state space model, constructing a dual-channel adaptive PID controller, dynamically adjusting PID parameters of a main channel based on an improved LSTM neural network, and establishing a dual-channel adaptive PID controller based on an improved LSTM neural network. The secondary channel compensates the nonlinear hysteresis effect of the suspension system based on the H-infinity robust control theory; designing a multi-mode cooperation strategy; an intelligent fault-tolerant mechanism is introduced, multi-source sensor data are fused through a federated filtering algorithm, fault positioning and signal reconstruction are realized, and experimental data are ensured not to be tampered and fault tracing integrity is ensured in combination with a block chain technology. According to the embodiment of the invention, the robustness and control precision of the system to complex road conditions and extreme loads are improved through reinforcement learning optimization of MPC weight, dynamic adjustment of PID parameters by the LSTM neural network and FPGA hardware acceleration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the control of a tractor bump test bench, and particularly to a cooperative control method for a tractor bump test bench with hybrid optimization of reinforcement learning - MPC. Background Art

[0002] As an important mechanical equipment in agricultural production, the bump problem during the driving process of a tractor not only affects the comfort of the driver and passengers, but also reduces the operation accuracy and efficiency, and even causes damage to the mechanical structure. Moreover, tractors often face complex terrains and bumpy road conditions during farm operations, which pose high requirements for the suspension system and load response of tractors. In order to study the performance of the suspension system, vibration characteristics, and optimization control strategies of tractors, a tractor bump test bench has become an important research tool.

[0003] However, traditional control methods for tractor bump test benches mostly adopt single control strategies, such as PID control or MPC control, and it is difficult to achieve the coordinated improvement of vibration amplitude - frequency tracking accuracy, load dynamic response speed, and suspension damping stability under complex road conditions. Moreover, when the existing control methods face extreme loads and complex road conditions, the robustness and fault - tolerance ability of the system are poor, and it is difficult to ensure the accuracy of experimental data and the stability of the system. Summary of the Invention

[0004] Aiming at the above - mentioned technical problems, the present invention provides a cooperative control method for a tractor bump test bench with hybrid optimization of reinforcement learning - MPC. By combining the reinforcement learning algorithm and the MPC framework, it realizes the coordinated improvement of vibration amplitude - frequency tracking accuracy, load dynamic response speed, and suspension damping stability, and enhances the robustness of the system to complex road conditions and extreme loads. The specific technical solutions are as follows:

[0005] The present invention provides a cooperative control method for a tractor bump test bench with hybrid optimization of reinforcement learning - MPC, including:

[0006] Step S1, establish a coupled non - linear state - space model of the tractor bump test bench system, where the coupled non - linear state - space model includes a vibration simulation unit, a dynamic load unit, a suspension feedback unit, and a road surface random excitation unit;

[0007] Step S2, design a hierarchical MPC framework based on the coupled non - linear state - space model. The upper layer optimizes the weight matrix of the objective function through the reinforcement learning algorithm, and the lower layer generates an optimal control sequence that satisfies the vibration amplitude - frequency constraint, energy loss limit, and multi - unit cooperative boundary conditions through rolling - horizon prediction;

[0008] Step S3: Construct a dual-channel adaptive PID controller, including an improved LSTM neural network main channel and an H∞ robust sub-channel. The main channel predicts the vibration error trend based on the improved LSTM neural network and dynamically adjusts the PID parameters. The sub-channel, based on the H∞ robust control theory, decomposes the road surface excitation spectrum into a deterministic component and a random component, generates a feedforward compensation signal and an anti-interference correction amount respectively, and compensates the nonlinear hysteresis effect of the suspension system in real time.

[0009] Step S4: Design a multi-modal cooperation strategy.

[0010] Step S5: Introduce an adaptive mechanism through vibration energy dissipation modeling and spectral fault diagnosis.

[0011] Further, the above-mentioned step S1 includes:

[0012] The vibration simulation unit, based on the electro-hydraulic servo and piezoelectric ceramic composite drive principle, establishes a hybrid relationship model between high-frequency or low-frequency vibration amplitude-frequency and multi-source control signals.

[0013] The dynamic load unit adopts magnetorheological fluid adaptive loading technology to construct an asymmetric hysteresis model between the load force and the magnetic field strength.

[0014] The road surface random excitation unit, based on the fractional calculus theory, generates a time-varying road surface spectrum that conforms to the ISO 8608 standard and embeds it into the state equation.

[0015] The coupled nonlinear state space model is discretized after being verified by the Lyapunov stability criterion to form a multi-variable state space expression containing a time-delay compensation term.

[0016] Further, the objective function dynamically optimized by the reinforcement learning algorithm in the above-mentioned step S2 is:

[0017]

[0018] In the formula, J is the objective function, which is used to comprehensively evaluate the performance of the control system, φ energy (k) is the vibration energy dissipation penalty term, which is calculated by real-time power integration; α, β, γ are the weight coefficients dynamically adjusted by reinforcement learning, satisfying the Pareto optimal condition; N is the number of steps within the optimization time window, y(k) is the actual output vector of the system at time step k, y ref (k) is the reference output vector of the system at time step k, and Δu(k) is the change amount of the control input at time step k.

[0019] Furthermore, the improved LSTM neural network includes an input layer and an output layer. The input layer contains 12-dimensional features including vibration acceleration frequency domain features, load force gradient, suspension displacement phase difference, and environmental temperature and humidity. The output layer uses a hybrid activation function composed of a Swish function for the output proportional term, a Tanh function for the output integral term, and a Sigmoid function for the output differential term, and accelerates online convergence through a meta-learning framework.

[0020] Furthermore, step S4 above includes:

[0021] The vibration load unit uses fuzzy sliding mode control to switch the output weights of MPC and PID to suppress high-frequency chattering.

[0022] The dynamic load unit introduces a generalized predictive control algorithm to eliminate the hysteresis effect of magnetorheological fluid through multi-step prediction.

[0023] The suspension feedback unit fuses IMU data and digital twin simulation results to construct an adaptive Kalman filter to optimize the damping control quantity.

[0024] Furthermore, after step S6, there is also an intelligent fault tolerance mechanism for: when the root mean square value of vibration acceleration exceeds the ISO 2631-1 health threshold, automatically insert a low-frequency sine sweep signal for system self-check. Among them, the root mean square value of vibration acceleration is calculated based on the vibration acceleration frequency domain features through frequency domain to time domain conversion;

[0025] Adopt a federated filtering algorithm to fuse multi-source sensor data, locate and isolate fault nodes, and reconstruct control signals;

[0026] Design a control log storage and proof module based on blockchain to ensure the integrity of the non-tampering of experimental data and fault traceability.

[0027] Furthermore, the state estimation equation of the federated filtering algorithm is:

[0028]

[0029] Among them, is the estimated value of the i-th local filter, is the trace of its covariance matrix, ω i is the weight coefficient, the weight assigned to the estimated value of the i-th local filter, and the weight assignment satisfies the minimum mean square error criterion.

[0030] Furthermore, the method also includes: obtaining an initial control parameter set and optimizing the initial control parameter set through a quantum particle swarm algorithm; realizing a microsecond-level real-time control cycle on the FPGA hardware platform to ensure that the system dynamic response bandwidth ≥ 200Hz under complex working conditions.

[0031] The beneficial effects of the present invention are as follows: by combining the reinforcement learning algorithm and the MPC framework, the LSTM neural network dynamically adjusts the PID parameters, and the FPGA hardware acceleration is realized, so as to achieve the coordinated improvement of the vibration amplitude-frequency tracking accuracy, the load dynamic response speed and the suspension damping stability, and enhance the robustness of the system to complex road conditions and extreme loads. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art.

[0033] Figure 1 It is the overall architecture diagram of the collaborative control method of the tractor bump test bench with reinforcement learning-MPC hybrid optimization of the present invention;

[0034] Figure 2 is Figure 1 The hierarchical MPC framework control flow chart of the collaborative control method shown;

[0035] Figure 3 is Figure 1 The structure diagram of the dual-channel adaptive PID controller of the collaborative control method shown;

[0036] Figure 4 is Figure 1 The multi-modal collaborative strategy flow chart of the collaborative control method shown;

[0037] Figure 5 is Figure 1 The intelligent fault tolerance mechanism flow chart of the collaborative control method shown;

[0038] Figure 6 It is the FPGA hardware acceleration control flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in combination with specific embodiments and the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Therefore, the detailed description of the embodiments of the present invention provided in the following drawings is not intended to limit the scope of the claimed invention.

[0040] Figure 1 It is the overall architecture diagram of the collaborative control method of the tractor bump test bench with reinforcement learning-MPC hybrid optimization of the present invention. The collaborative control system of the tractor bump test bench with reinforcement learning-MPC hybrid optimization includes:

[0041] The reinforcement learning-MPC hybrid optimization module is used for hierarchical target weight optimization and generation of the rolling horizon control sequence;

[0042] Dual-channel adaptive PID controller module, including an improved LSTM neural network main channel and an H∞ robust sub-channel. The main channel predicts the vibration error trend based on the improved LSTM neural network and dynamically adjusts the PID parameters; the sub-channel, based on the H∞ robust control theory, decomposes the road surface excitation spectrum into a deterministic component and a random component, generates a feedforward compensation signal and an anti-interference correction amount respectively, and compensates the nonlinear hysteresis effect of the suspension system in real time;

[0043] Federated filtering fault-tolerant module for multi-source sensor data fusion and fault isolation;

[0044] Blockchain data storage and certification interface module for encrypted storage and integrity verification of control logs.

[0045] Based on the above collaborative control system of the tractor bump test bench with reinforcement learning-MPC hybrid optimization, the present invention designs a collaborative control method for the tractor bump test bench with reinforcement learning-MPC hybrid optimization, including the following steps:

[0046] Step S1: Establish a coupled nonlinear state space model of the tractor bump test bench system, which includes a vibration simulation unit, a dynamic load unit, a suspension feedback unit, and a road surface random excitation unit;

[0047] Among them, the vibration simulation unit, based on the electro-hydraulic servo and piezoelectric ceramic composite drive principle, establishes a hybrid relationship model of high-frequency or low-frequency vibration amplitude-frequency and multi-source control signals;

[0048] The dynamic load unit adopts magnetorheological fluid adaptive loading technology to construct an asymmetric hysteresis model of load force and magnetic field strength;

[0049] The road surface random excitation unit, based on the fractional calculus theory, generates a time-varying road surface spectrum that conforms to the ISO 8608 standard and embeds it into the state equation;

[0050] The coupled nonlinear state space model is discretized after being verified by the Lyapunov stability criterion to form a multi-variable state space expression including a time delay compensation term.

[0051] Step S2: Design a hierarchical MPC framework based on the coupled nonlinear state space model. The upper layer optimizes the weight matrix of the objective function through a reinforcement learning algorithm, and the lower layer generates an optimal control sequence that satisfies the vibration amplitude-frequency constraint, the energy loss limit, and the multi-unit collaborative boundary conditions through rolling horizon prediction.

[0052] Among them, as Figure 2As shown in the figure, it is the structure diagram of the hierarchical MPC framework. The upper layer optimizes the weight matrix of the objective function through the reinforcement learning algorithm. The lower layer generates the optimal control sequence through the rolling horizon prediction, and imposes the vibration amplitude-frequency constraint, the energy loss limit value and the multi-unit cooperation boundary condition, and outputs the generated optimal control sequence.

[0053] The objective function dynamically optimized by the reinforcement learning algorithm is:

[0054]

[0055] Among them, φ energy (k) is the vibration energy dissipation penalty term, which is calculated by real-time power integration; α, β, γ are the weight coefficients dynamically adjusted by reinforcement learning, satisfying the Pareto optimal condition; N is the number of steps within the optimization time window, y(k) is the actual output vector of the system at time step k, and y ref( k) is the reference output vector of the system at time step k, and Δu(k) is the change in the control input at time step k.

[0056] Specifically, J is the objective function, which is used to comprehensively evaluate the performance of the control system, including aspects such as tracking accuracy, smoothness of control input change, and vibration energy dissipation. By minimizing J, the optimal control of the system can be achieved; N is the number of steps within the optimization time window, that is, in the model predictive control (MPC) framework, the control performance in the next N steps is considered; α is the weight coefficient of the vibration tracking error, which is used to adjust the importance of the vibration amplitude-frequency tracking accuracy in the objective function. A larger α value means a greater penalty for the vibration tracking error, thereby improving the vibration amplitude-frequency tracking accuracy of the system; y(k) is the actual output vector of the system at time step k, including the outputs of the vibration simulation unit, the dynamic load unit, the suspension feedback unit, etc., such as vibration acceleration, load force, suspension displacement, etc.; y ref(k) is the reference output vector of the system at time step k, including target values such as the desired vibration amplitude-frequency, load dynamic response, and suspension damping stability; β is the weight coefficient for the change in the control input, used to adjust the importance of the change in the control input in the objective function. A larger β value indicates a greater penalty for the change in the control input, thus making the control input smoother, reducing the energy consumption and mechanical wear of the system; Δu(k) is the change in the control input at time step k, that is, Δu(k) = u(k) - u(k - 1), where u(k) is the control input vector at time step k; γ is the weight coefficient for vibration energy dissipation, used to adjust the importance of vibration energy dissipation in the objective function. A larger γ value indicates a greater penalty for vibration energy dissipation, thus prompting the system to reduce vibration energy and improve ride comfort and the lifespan of the mechanical structure; the matrix Q is a positive definite matrix, used to weight errors in different dimensions, reflecting the relative importance of each error term in the objective function; the matrix R is a positive definite matrix, used to weight changes in the control input in different dimensions, reflecting the relative importance of each change in the control input in the objective function.

[0057] Step S3: Construct a two-channel adaptive PID controller, including an improved LSTM neural network main channel and an H∞ robust sub-channel. The main channel predicts the trend of vibration error based on the improved LSTM neural network and dynamically adjusts the PID parameters; the sub-channel is based on the H∞ robust control theory, decomposes the road surface excitation spectrum into a deterministic component and a random component, generates a feedforward compensation signal and an anti-interference correction amount respectively, and compensates the nonlinear hysteresis effect of the suspension system in real time.

[0058] Among them, the improved LSTM neural network includes an input layer and an output layer. The input layer contains 12-dimensional features such as the vibration acceleration frequency domain characteristics, load force gradient, suspension displacement phase difference, and environmental temperature and humidity. The output layer adopts a mixed activation function composed of the Swish function for the proportional term, the Tanh function for the integral term, and the Sigmoid function for the differential term, and accelerates online convergence through a meta-learning framework.

[0059] Specifically, as Figure 3 shown, the structure diagram of the two-channel adaptive PID controller. The improvements of the LSTM neural network include:

[0060] Input feature fusion: Fuse heterogeneous data such as vibration acceleration frequency domain characteristics (such as multi-band energy distribution), load force gradient (dynamic load change rate), suspension displacement phase difference (time domain and frequency domain phase offset), and environmental temperature and humidity to improve the feature expression ability.

[0061] Among them, the vibration acceleration frequency-domain feature is the energy proportion decomposed into 6 frequency bands (0 - 5Hz, 5 - 10Hz, 10 - 20Hz, 20 - 30Hz, 30 - 50Hz, 50 - 100Hz) through FFT, with a total of 6 dimensions; the load force gradient is the instantaneous change rate of the dynamic load (1 dimension); the suspension displacement phase difference includes the phase difference between the time-domain displacement (1 dimension) and the excitation signal (1 dimension) and the frequency-domain phase offset (1 dimension); the environmental temperature and humidity include temperature (1 dimension) and humidity (1 dimension), totaling 12-dimensional features.

[0062] Hybrid activation function design: The output layer adopts a combination of the Swish function (smooth non-linearity), Tanh function (constraining the output range), and Sigmoid function (probabilistic adjustment) to enhance the network's adaptability to different vibration modes.

[0063] The meta-learning framework accelerates online convergence by introducing a meta-learning mechanism, using historical experimental data to pre-train network parameters, realizing fast online parameter tuning, and reducing the number of iterations.

[0064] Step S4: Design a multi-modal collaborative strategy.

[0065] Among them, through the designed multi-modal collaborative strategy, the global optimization of MPC is combined with the local regulation of the dual-channel PID to achieve the collaborative improvement of the vibration amplitude-frequency tracking accuracy, the dynamic response speed of the load, and the suspension damping stability.

[0066] Specifically, as Figure 4 shown, the multi-modal collaborative strategy includes:

[0067] The vibration load unit uses fuzzy sliding mode control to switch the output weights of MPC and PID, suppressing the high-frequency chattering phenomenon;

[0068] The dynamic load unit introduces a generalized predictive control algorithm to eliminate the magnetorheological hysteresis effect through multi-step prediction;

[0069] The suspension feedback unit fuses IMU data and digital twin simulation results to construct an adaptive Kalman filter to optimize the damping control quantity.

[0070] Among them, the suspension feedback unit provides basic data as state variables during the modeling stage; during the design stage of the multi-modal collaborative strategy, it serves as a carrier for intelligent algorithms to achieve closed-loop optimization, that is, the function of the suspension feedback unit in the multi-modal collaborative strategy is the practical application and enhancement of the ability of the suspension feedback unit in the coupled non-linear state space model, and the two form an "model-control" integrated design.

[0071] Step S5: Introduce an adaptive mechanism to model and diagnose spectrum faults using vibration energy dissipation.

[0072] By introducing an adaptive mechanism and using vibration energy dissipation modeling and spectrum fault diagnosis, the robustness of the system to complex road conditions and extreme loads is enhanced.

[0073] Optionally, after step S5, an intelligent fault tolerance mechanism is further included for:

[0074] When the root mean square value of vibration acceleration exceeds the ISO 2631-1 health threshold, a low-frequency sine sweep signal is automatically inserted for system self-check. Among them, the root mean square value of vibration acceleration is calculated based on the frequency domain characteristics of vibration acceleration after frequency domain to time domain conversion;

[0075] The federated filtering algorithm is used to fuse multi-source sensor data, locate and isolate the fault nodes, and reconstruct the control signal;

[0076] A control log storage and proof module based on blockchain is designed to ensure the integrity of experimental data non-tampering and fault traceability.

[0077] Optionally, the state estimation equation of the federated filtering algorithm is:

[0078]

[0079] Among them, is the estimated value of the i-th local filter, is the trace of its covariance matrix, ω i is the weight coefficient, the weight assigned to the estimated value of the i-th local filter, and the weight assignment satisfies the minimum mean square error criterion, that is, the weighted sum of all estimated values is optimized to reach the minimum error.

[0080] Specifically, when the root mean square value of vibration acceleration does not exceed the ISO 2631-1 health threshold, it operates normally. However, when the root mean square value of vibration acceleration exceeds the ISO 2631-1 health threshold, a low-frequency sine sweep signal is automatically inserted for system self-check. The federated filtering algorithm is used to fuse multi-source sensor data to locate the fault nodes. If the fault nodes are not located, continue to monitor. If the fault nodes are located, reconstruct the control signal, and then the control log storage and proof module based on blockchain is used to store the blockchain data to ensure the integrity of experimental data non-tampering and fault traceability.

[0081] Optionally, as Figure 6 shown, obtain the initial control parameter set, optimize the initial control parameter set through the quantum particle swarm algorithm, and implement a microsecond-level real-time control cycle on the FPGA hardware platform to ensure that the system dynamic response bandwidth ≥ 200Hz under complex working conditions.

[0082] Specifically, in the initial stage of the system, the initial control parameter set is pre-optimized by the quantum particle swarm algorithm to provide a high-performance initial configuration for the system and reduce the computational burden of online adjustment. After optimizing the initial control parameter set, the objective function is obtained. During the operation of the system, it is dynamically adjusted using reinforcement learning, the MPC framework, and an adaptive PID controller, etc. The optimized initial parameter set provides a good starting point for dynamic parameter adjustment, improving the overall control efficiency and stability.

[0083] The present invention combines the reinforcement learning algorithm and the MPC framework, dynamically adjusts the PID parameters by the LSTM neural network, and accelerates with FPGA hardware, achieving a coordinated improvement in the vibration amplitude-frequency tracking accuracy, the load dynamic response speed, and the suspension damping stability, and enhancing the robustness of the system to complex road conditions and extreme loads.

[0084] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts, such as making formal modifications to the technical solutions recorded in the following embodiments or making equivalent replacements for some of the technical features, fall within the protection scope of the present invention.

Claims

1. A reinforcement learning-MPC hybrid optimization tractor bump test bench collaborative control method, comprising the following steps: Step 1: Establish a coupled nonlinear state space model of the tractor bump test bench system, wherein the coupled nonlinear state space model includes a vibration simulation unit, a dynamic load unit, a suspension feedback unit and a road surface random excitation unit; Step 2: Based on the coupled nonlinear state space model, a hierarchical MPC framework is designed. The upper layer optimizes the weight matrix of the objective function through a reinforcement learning algorithm, and the lower layer generates an optimal control sequence that satisfies vibration amplitude-frequency constraints, energy loss limits, and multi-unit collaborative boundary conditions through rolling time domain prediction. Step 3: Build a dual-channel adaptive PID controller, including an improved LSTM neural network main channel and an H∞ robust sub-channel. The main channel predicts the vibration error trend based on the improved LSTM neural network and dynamically adjusts the PID parameters. The sub-channel is based on the H∞ robust control theory, which decomposes the road excitation spectrum into deterministic components and random components, generates feedforward compensation signals and anti-interference corrections, and compensates for the nonlinear hysteresis effect of the suspension system in real time. Step 4: Design a multimodal collaboration strategy; Step 5: Introduce an adaptive mechanism and use vibration energy dissipation modeling and spectrum fault diagnosis.

2. The tractor bump test bench collaborative control method based on reinforcement learning-MPC hybrid optimization according to claim 1 is characterized in that: The step one comprises: The vibration simulation unit is based on the principle of electro-hydraulic servo and piezoelectric ceramic composite drive, and establishes a mixed relationship model between high-frequency or low-frequency vibration amplitude and frequency and multi-source control signals; The dynamic load unit adopts magnetorheological fluid adaptive loading technology to construct an asymmetric hysteresis model of load force and magnetic field intensity; The road random excitation unit generates a time-varying road spectrum that complies with the ISO 8608 standard based on fractional calculus theory and embeds it into the state equation; The coupled nonlinear state space model is discretized after verification by the Lyapunov stability criterion to form a multivariable state space expression including a time delay compensation term.

3. The tractor bump test bench collaborative control method based on reinforcement learning-MPC hybrid optimization according to claim 2 is characterized in that: In step 2, the objective function of the dynamic optimization of the reinforcement learning algorithm is: Where J is the objective function, which is used to comprehensively evaluate the performance of the control system, φ energy (k) is the vibration energy dissipation penalty term, which is calculated by real-time power integration; α, β, γ are the weight coefficients dynamically adjusted by reinforcement learning to meet the Pareto optimality condition; N is the number of steps in the optimization time window, y(k) is the actual output vector of the system at time step k, y ref( k) is the reference output vector of the system at time step k, and Δu(k) is the change of the control input at time step k.

4. The tractor bump test bench collaborative control method based on reinforcement learning-MPC hybrid optimization according to claim 3 is characterized in that: The improved LSTM neural network includes an input layer and an output layer. The input layer contains 12-dimensional features including vibration acceleration frequency domain characteristics, load force gradient, suspension displacement phase difference and ambient temperature and humidity. The output layer adopts a mixed activation function composed of a Swish function output proportional term, a Tanh function output integral term and a Sigmoid function output differential term, and accelerates online convergence through a meta-learning framework.

5. The tractor bump test bench collaborative control method based on reinforcement learning-MPC hybrid optimization according to claim 4 is characterized in that: The fourth step comprises: The vibration load unit uses fuzzy sliding mode control to switch the output weights of MPC and PID to suppress high-frequency chattering; The dynamic load unit introduces a generalized predictive control algorithm to eliminate the hysteresis effect of magnetorheological fluid through multi-step prediction; The suspension feedback unit is used to fuse IMU data with digital twin simulation results and construct an adaptive Kalman filter to optimize the damping control amount.

6. The control method of the tractor bump test bench based on reinforcement learning-MPC hybrid optimization according to any one of claims 1 to 5, characterized in that: After step 5, an intelligent fault-tolerant module is also included, which is used to: When the root mean square value of the vibration acceleration exceeds the ISO 2631-1 health threshold, a low-frequency sine sweep signal is automatically inserted to perform a system self-check, wherein the root mean square value of the vibration acceleration is calculated based on the frequency domain characteristics of the vibration acceleration by converting the frequency domain to the time domain; A federated filtering algorithm is used to fuse multi-source sensor data, locate and isolate faulty nodes, and reconstruct control signals. Design a blockchain-based control log notarization module to ensure that experimental data cannot be tampered with and that fault traceability is complete.

7. The control method of the tractor bump test bench based on reinforcement learning-MPC hybrid optimization according to claim 6 is characterized in that: The state estimation equation of the federated filtering algorithm is: In the formula, is the estimated value of the i-th local filter, is the covariance matrix trace, ω i is the weight coefficient, which is the weight assigned to the i-th local filter estimate. The weight assignment satisfies the minimum mean square error criterion.

8. The control method of the tractor bump test bench based on reinforcement learning-MPC hybrid optimization according to claim 1 is characterized in that: An initial control parameter set is obtained and optimized by a quantum particle swarm algorithm; the method realizes a microsecond-level real-time control cycle on an FPGA hardware platform, ensuring that the system dynamic response bandwidth is ≥200 Hz under complex working conditions.

Citation Information

Cited By

  • Electric vehicle hybrid energy storage system energy management method fused with Hemma particle swarm optimization LSTM

    CN121019302A