Active vibration control method based on PPO-HAFxLMS algorithm

By combining the PPO-HAFxLMS algorithm with feedback channels and the PPO reinforcement learning algorithm, the step size and filter weights are dynamically adjusted, which solves the problems of insufficient adaptive step size adjustment and computational complexity in existing active vibration control systems, and achieves high-precision and high-real-time vibration control effect.

CN121979053APending Publication Date: 2026-05-05NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHEASTERN UNIV CHINA
Filing Date
2026-01-29
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing active vibration control technologies suffer from problems such as insufficient adaptive step size adjustment, lack of multi-channel control, and difficulty in balancing computational complexity and real-time performance when dealing with complex dynamic vibration sources, thus failing to meet the requirements for high-precision and high-real-time control.

Method used

An active vibration control method based on the PPO-HAFxLMS algorithm is adopted, which combines feedback channels and PPO reinforcement learning algorithm to dynamically adjust step size and filter weights, thereby optimizing the convergence speed and steady-state error of the vibration control system.

Benefits of technology

It achieves dynamic step size adjustment, improves vibration control accuracy and system stability, and can quickly respond to and adapt to changes in vibration sources, meeting the control requirements of high precision and high real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979053A_ABST
    Figure CN121979053A_ABST
Patent Text Reader

Abstract

The invention discloses a vibration active control method based on a PPO-HAFxLMS algorithm, and relates to the field of vibration active control. The method comprises the following steps: determining a target vibration frequency band needing to be suppressed and a vibration attenuation standard so as to determine configuration of related hardware in the vibration active control system; capturing an external disturbance signal received by the controlled object; and on the basis of the time domain data of the external disturbance signal, a required control voltage signal is generated in real time by using a PPO-HAFxLMS algorithm, and vibration suppression is realized. According to the method, on the basis of a traditional FxLMS algorithm, a feedback channel is added to form a HAFxLMS algorithm, filter coefficients and step length updating are optimized through the HAFxLMS algorithm, the step length and the weight of a filter are dynamically adjusted according to the correlation between error signals and input signals and the change of the environment through a PPO reinforcement learning algorithm, and therefore the convergence speed and the steady-state error of vibration control are optimized, and the stability of vibration control is improved. And the vibration suppression effect is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of active vibration control technology, specifically involving an active vibration control method based on the PPO (Proximal Policy Optimization) reinforcement learning algorithm and the HAFxLMS (Hybrid Adaptive FxLMS) algorithm. Background Technology

[0002] With the increasing demand for vibration control technology, vibration issues have gradually become a key consideration in many engineering applications. The effectiveness of vibration control not only directly affects the stability and service life of equipment but also the overall performance of the system. In many fields, vibration control has become a crucial factor in improving product quality and performance. Especially in the control of low-frequency vibrations, due to their complexity and difficulty, it has become one of the main challenges facing vibration control technology.

[0003] In the field of active vibration control, the FxLMS algorithm is widely used in active vibration control systems due to its simple structure, high efficiency, and ease of implementation. This algorithm updates filter coefficients by minimizing the error signal and has been applied in various fields such as vehicle vibration, industrial equipment, and building structures. However, it still has a key drawback—it uses a fixed step size to update the filter. While a larger step size can achieve fast convergence, it easily introduces a large steady-state error; a smaller step size can reduce the steady-state error, but it leads to a slower convergence speed, ultimately causing a decline in the performance of the active vibration control system. To address the fixed step size problem of the FxLMS algorithm, some improved solutions have emerged (such as adaptive control methods based on the LMS algorithm), but these solutions still suffer from insufficient step size adjustment and poor real-time performance, making it difficult to meet the control requirements of complex vibration scenarios.

[0004] Existing related patent technologies also have shortcomings: Chinese patent application CN120122516A, "An Active Vibration Control Method Based on Online Identification of Secondary Channel FxLMS Algorithm," collects error signals through the LMS algorithm and applies the trapezoidal algorithm to optimize the control effect to adapt to vibration sources under different working conditions. However, this method still relies on a fixed step size and lacks an adaptive adjustment mechanism for dynamically changing environments, which limits the adaptability and accuracy of the control algorithm. Chinese patent CN114690622B, "An Adaptive Ship Vibration Control Method Based on Differential Calculation," adjusts control parameters in real time through a differential algorithm to optimize the vibration suppression effect. However, this algorithm has high computational complexity, which can easily affect the real-time response and control accuracy of the control algorithm. Moreover, in high-frequency changing environments, the adaptability and robustness of the control algorithm are poor, making it difficult to transfer and apply to different dynamic vibration scenarios.

[0005] In summary, existing active vibration control technologies generally suffer from problems such as insufficient adaptive step size adjustment, lack of multi-channel control, and difficulty in balancing computational complexity and real-time performance when dealing with complex dynamic vibration sources. They cannot meet the control requirements of high precision and high real-time performance. Therefore, it is particularly necessary to research and develop an active vibration control algorithm that can be adjusted in real time. Summary of the Invention

[0006] Therefore, this invention provides a vibration active control method based on the PPO-HAFxLMS algorithm. This method, based on the traditional FxLMS (Filtered-x Least Mean Square) algorithm, adds a feedback channel to form the HAFxLMS (Hybrid Adaptive FxLMS) algorithm. Furthermore, it dynamically adjusts the step size and filter weights through the PPO reinforcement learning algorithm, thereby optimizing the convergence speed and steady-state error of the vibration control system and significantly improving the vibration suppression effect.

[0007] The technical solution of this invention is:

[0008] An active vibration control method based on the PPO-HAFxLMS algorithm, comprising the following steps:

[0009] The target vibration frequency band to be suppressed and the vibration attenuation criteria are determined, thereby determining the configuration of relevant hardware in the active vibration control system;

[0010] Capture external disturbance signals received by the controlled object;

[0011] Based on the time-domain data of external disturbance signals, the vibration active control system uses the PPO-HAFxLMS algorithm to generate the required control voltage signal in real time to achieve vibration suppression.

[0012] Optionally, according to the vibration active control method, the configuration of the relevant hardware includes the type, quantity, and layout of sensors and actuators; the sensors include an accelerometer installed near the external disturbance signal as a reference sensor and another accelerometer installed in the vibration suppression region as an error sensor; the actuators are installed in the vibration suppression region of the controlled object.

[0013] Optionally, according to the aforementioned active vibration control method, the method for capturing the external disturbance signal received by the controlled object is as follows: capturing the external disturbance signal received by the controlled object from the acceleration signal collected by the reference sensor.

[0014] Optionally, according to the vibration active control method, the vibration active control system includes: a channel P(z) between the external disturbance signal x(n) received by the controlled object at any n time and the vibration response signal measured by the error sensor; a secondary channel, i.e., a channel model S(z) between the control voltage signal and the vibration response signal collected by the sensor; a feedforward filter W1(z) for generating the feedforward control signal y1(n); and a feedback filter W2(z) for generating the feedback control signal y2(n); where z represents a complex variable in the Z-transform, y1(n) is used to deal with the vibration caused by the external disturbance, and y2(n) is used to correct the response of the vibration active control system by adjusting the error signal e(n) of the vibration response.

[0015] Optionally, according to the aforementioned active vibration control method, based on the time-domain data of the input vibration signal, the active vibration control system utilizes the PPO-HAFxLMS algorithm to generate the required control voltage signal in real time, thereby suppressing the vibration signal, including:

[0016] Step 4.1: Based on the external disturbance signal x(n) received by the controlled object at any time n, calculate the actual vibration response time-domain data d(n) of the controlled object using the HAFxLMS algorithm, and then calculate the vibration response error signal e(n). The calculation process is as follows:

[0017] (1)

[0018] (2)

[0019] In the formula, This represents the vibration control quantity applied to the controlled object by the actuation voltage signal y(n) via the secondary channel S(z);

[0020] Step 4.2: Based on the calculated error signal e(n), update the feedforward filter W1(z) and feedback filter W2(z) using the HAFxLMS algorithm to optimize the generation of the control voltage signal and reduce the vibration error of the controlled object;

[0021] Step 4.3: Use the HAFxLMS algorithm to generate the control voltage signal. During the generation of the control voltage signal, the step size of the feedforward filter W1(z) and the step size of the feedback filter W2(z) are adaptively optimized using the PPO reinforcement learning algorithm.

[0022] Optionally, according to the aforementioned active vibration control method, the feedforward filter W1(z) is updated using the HAFxLMS algorithm as follows:

[0023] (3)

[0024] In the formula, W1(n) is the coefficient of the feedforward filter at time n, and W1(n+1) is the coefficient of the feedforward filter at time n+1. is the step size parameter of the feedforward filter, used to control the rate at which the feedforward filter coefficients are updated; x′(n) is the signal after the input signal x(n) is identified by the secondary channel;

[0025] The feedback filter W2(z) is updated using the HAFxLMS algorithm as follows:

[0026] (4)

[0027] In the formula, W2(n) is the coefficient of the feedback filter at time n, and W2(n+1) is the coefficient of the feedback filter at time n+1; This is the step size parameter of the feedback filter, used to control the rate at which the feedback filter coefficients are updated; The actuation voltage signal, i.e., the control input y(n), is identified by the secondary channel S(z). The control input signal identified by the error sensor is numerically... .

[0028] Optionally, according to the aforementioned active vibration control method, the generation of the control voltage signal using the HAFxLMS algorithm includes:

[0029] First, based on the input signal x(n) and the feedforward filter W1(z), the feedforward control signal y1(n) is generated according to equation (5):

[0030] (5)

[0031] In the formula, for transpose;

[0032] Then, based on the error signal e(n) and the secondary channel S(z), the feedback control signal y2(n) is calculated according to equation (6):

[0033] (6)

[0034] In the formula, for transpose; This indicates the identification of the secondary channel S(z);

[0035] Then, the feedforward control signal y1(n) and the feedback control signal y2(n) are combined to obtain the total control voltage signal y(n) acting on the controlled object:

[0036] (7)

[0037] Finally, the total control voltage signal y(n) is processed by the FIR filter, drives the actuator to generate control force, acts on the controlled object, and generates the final control signal y′(n) through the secondary channel S(z):

[0038] (8)

[0039] In the formula, s1(n) is the impulse response of the secondary channel S(z).

[0040] Optionally, according to the aforementioned active vibration control method, in the process of generating the control voltage signal in step 4.3, the step size of the feedforward filter W1(z) and the step size of the feedback filter W2(z) are adaptively optimized using the PPO reinforcement learning algorithm, including: the state space s(n) is designed as follows:

[0041] (10)

[0042] In the formula, e(n-1) represents the historical vibration error;

[0043] The action space a(n) is defined as:

[0044] (11)

[0045] In the formula, Δμ1(n) and Δμ2(n) represent the adjustment amounts of the feedforward channel step size and the feedback channel step size generated by the PPO reinforcement learning algorithm, respectively, used to optimize the filter weight update formula:

[0046] (12)

[0047] In the formula, W 11 (n+1) and W 22 (n+1) are the coefficients of the feedforward filter and the feedback filter at time n+1 after PPO reinforcement learning compensation adjustment;

[0048] The reward function r(n) is designed as follows:

[0049] (14)

[0050] In the formula, λ is a trade-off factor used to balance error minimization and step size stability;

[0051] The optimization objective function of the PPO reinforcement learning algorithm is:

[0052] (16)

[0053] In the formula, Importance sampling ratio; ϵ is the advantage function, which measures the merits of the current action; ϵ is a pruning parameter used to limit the magnitude of policy updates.

[0054] The vibration active control method provided by this invention can be widely applied in various fields such as automobiles, industrial equipment, and building structures, and is especially suitable for the active control of low and medium frequency vibrations. Compared with the prior art, this invention has the following beneficial effects:

[0055] (1) Dynamic step size adjustment: Based on the correlation between the error signal and the input signal, the step size is dynamically adjusted according to the changes in the environment through the PPO reinforcement learning algorithm to optimize the vibration control effect and overcome the problems of slow convergence speed and large steady-state error caused by fixed step size.

[0056] (2) Improve vibration control accuracy: Based on the traditional FxLMS algorithm, this invention adds a feedback channel to form the HAFxLMS algorithm. By optimizing the filter coefficients and step size updates through the HAFxLMS algorithm, the vibration active control system can suppress vibration signals more accurately, especially in the suppression of low-frequency and mid-frequency vibrations.

[0057] (3) System stability and real-time performance: By using the PPO reinforcement learning algorithm to adjust the step size through feedback mechanism and reward signal, the system can respond quickly and adapt to different changes in vibration source, maintaining good stability and real-time performance. Attached Figure Description

[0058] Figure 1 This is a flowchart illustrating the active vibration control method based on the PPO-HAFxLMS algorithm in this embodiment.

[0059] Figure 2 This is a schematic diagram illustrating the principle of the active vibration control method based on the PPO-HAFxLMS algorithm in this embodiment.

[0060] Figure 3 This is a flowchart of the PPO reinforcement learning algorithm in this embodiment;

[0061] Figure 4 This is a graph showing the cumulative reward value change of the PPO-HAFxLMS algorithm in this implementation method;

[0062] Figure 5 This is a graph showing the variation of step sizes μ1 and μ2 in this embodiment;

[0063] Figure 6 This is a comparison graph of the mean square error variation curves of the method of the present invention and the existing methods. Detailed Implementation

[0064] To facilitate understanding of this application, a more comprehensive description of this application will be provided below with reference to the accompanying drawings.

[0065] Figure 1 This is a flowchart illustrating the active vibration control method based on the PPO-HAFxLMS algorithm in this embodiment. Figure 1 As shown, the vibration active control method based on the PPO-HAFxLMS algorithm includes the following steps:

[0066] Step 1: Obtain the vibration source of the controlled object, and determine the target vibration frequency band and vibration attenuation standard to be controlled.

[0067] The vibration source of the controlled object is typically identified through testing or theoretical analysis, and the target vibration frequency band that needs to be suppressed is determined. Based on the frequency range of the vibration, it can be divided into low frequency (0.1-100Hz), mid frequency (100Hz-1kHz), and high frequency (above 1kHz). In a preferred embodiment, the target vibration frequency band to be controlled is low frequency (0.1-100Hz). Simultaneously, a vibration attenuation standard needs to be set to ensure that the active vibration control system achieves effective vibration control within the specified frequency band.

[0068] Step 2: Determine the configuration of relevant hardware in the active vibration control system according to the target vibration frequency band to be controlled.

[0069] In the preferred embodiment, the target vibration frequency band to be controlled is low frequency. The relevant hardware includes two accelerometers: one as a reference sensor installed near the external disturbance signal, and the other as an error sensor installed in the vibration suppression region. An electromagnetic actuator is installed in the vibration suppression region of the controlled object. The configuration of the sensors and actuators is closely related to the characteristics of the controlled object. Specific selection methods are well-established in this field and will not be discussed in detail here.

[0070] Step 3: Capture the external disturbance signal experienced by the controlled object from the acceleration signal collected by the accelerometer installed near the external disturbance signal;

[0071] Step 4: Based on the time-domain data of the external disturbance signal, the vibration active control system uses the PPO-HAFxLMS algorithm to generate the required control voltage signal in real time, so that the signal vibration can be suppressed.

[0072] Figure 2 In the diagram, P(z) represents the path between the external disturbance signal x(n) and the vibration response signal measured by the error sensor; S(z) represents the secondary path model, i.e., the path model between the control voltage signal and the vibration response signal acquired by the sensor; x(n), d(n), y(n) and These represent the external disturbance signal, the actual vibration response time-domain data of the controlled object, the actuation voltage signal, and the vibration control quantity applied to the controlled object via the actuation voltage signal through the secondary channel, respectively. W1(z) and W2(z) are the feedforward filter and feedback filter of the PPO-HAFxLMS algorithm, respectively. The PPO-HAFxLMS algorithm adjusts the step size μ1(n) of the feedforward filter and the step size μ2(n) of the feedback filter based on the feedback error signal e(n) in each iteration, gradually bringing these two step sizes closer to their optimal values, thereby optimizing the convergence speed and steady-state error of the active vibration control system.

[0073] The specific calculation process is as follows:

[0074] Step 4.1: Based on the external disturbance signal x(n) received by the controlled object at any time n, calculate d(n) using the HAFxLMS algorithm, and then calculate the vibration response error signal e(n). The calculation process is as follows:

[0075] (1)

[0076] (2)

[0077] In the formula, d(n) represents the time-domain data of the actual vibration response of the controlled object.

[0078] Step 4.2: Based on the calculated error signal e(n), update the feedforward filter W1(z) and feedback filter W2(z) using the HAFxLMS algorithm to optimize the generation of the control voltage signal and reduce the vibration error of the controlled object;

[0079] Here, z represents the complex variable in the Z-transform, describing the behavior of filters W1(z) and W2(z) in the frequency domain and the response of the HAFxLMS algorithm. The filter update process aims to reduce the vibration error of the controlled object and optimize the generation of the control voltage signal. The update process of the feedforward filter W1(z) is as follows:

[0080] The feedforward filter is the part used to process the input signal and generate the initial control voltage signal. Its update formula is:

[0081] (3)

[0082] In the formula, W1(n) is the coefficient of the feedforward filter at time n, and W1(n+1) is the coefficient of the feedforward filter at time n+1. is the step size parameter of the feedforward filter, used to control the rate at which the feedforward filter coefficients are updated; x′(n) is the signal after the input signal x(n) is identified by the secondary channel. Using this formula, the coefficients of the feedforward filter are adjusted according to the current error signal to generate a more accurate control signal.

[0083] Similar to a feedforward filter, the HAFxLMS algorithm also updates the filter coefficients of the feedback filter W2(z) to further optimize the feedback portion of the control voltage signal. The update formula is:

[0084] (4)

[0085] In the formula, W2(n) is the coefficient of the feedback filter at time n, and W2(n+1) is the coefficient of the feedback filter at time n+1. This is the step size parameter of the feedback filter, used to control the rate at which the feedback filter coefficients are updated; Identification of y(n) through secondary channels The control input signal detected by the error sensor is... Figure 2 It can be seen that numerically By updating the coefficients of the feedback filter, the HAFxLMS algorithm can correct for the vibration errors that have already occurred.

[0086] Step 4.3: Use the HAFxLMS algorithm to generate the control voltage signal. During the generation of the control voltage signal, the step size of the filter is adjusted and optimized using the PPO reinforcement learning algorithm.

[0087] In this step, the HAFxLMS algorithm generates a control voltage signal y(n) and uses it for active vibration control. First, the HAFxLMS algorithm takes into account the input signal x(n) and the feedforward filter W1(z) to generate a feedforward control signal y1(n), calculated as follows:

[0088] (5)

[0089] In the formula, For the feedforward filter at time n The transpose of y1(n). This feedforward control signal y1(n) is mainly used to deal with external disturbance signals x(n), and directly affects the response of the HAFxLMS algorithm to external disturbances.

[0090] Then, the HAFxLMS algorithm calculates the feedback control signal y2(n), which is based on the error signal e(n) and the secondary channel S(z) to adjust the control signal. The secondary channel is the dynamic transmission path from the control input y(n) to the error signal measurement point. The formula for generating the feedback control signal is:

[0091] (6)

[0092] In the formula, Feedback filter at time n transpose; This represents the identification of the secondary channel S(z). Through a feedback mechanism, the HAFxLMS algorithm adjusts the feedback signal y2(n) based on the error signal and the identification of the secondary channel. The model was further optimized to improve vibration control performance.

[0093] The generated feedforward control signal y1(n) and feedback control signal y2(n) are finally synthesized to obtain the total control voltage signal y(n):

[0094] (7)

[0095] The control voltage signal y(n) is the final control signal applied to the controlled object, used to reduce the vibration error of the controlled object. The feedforward control signal mainly deals with external disturbances, while the feedback control signal adjusts according to the error of the controlled object. The combination of the two can control vibration more accurately.

[0096] Finally, the total control voltage signal y(n) is processed by the FIR filter, drives the actuator to generate control force, acts on the controlled object, and generates the final control signal y′(n) through the secondary channel S(z):

[0097] (8)

[0098] In the formula, s1(n) is the impulse response of the secondary channel S(z).

[0099] In summary, the feedforward control signal y1(n) is mainly used to address vibrations caused by external disturbances, while the feedback control signal y2(n) corrects the response of the active vibration control system by adjusting the vibration response error signal e(n). Through the combination of feedforward and feedback mechanisms, the active vibration control system can quickly adapt to different environmental changes and disturbance conditions, optimizing the vibration control effect. In the vibration control region, y′(n) is expected to have the same amplitude and frequency as the actual vibration response time-domain data d(n) of the controlled object, but with opposite phase, to achieve active vibration control.

[0100] To further improve the control accuracy and adaptability of the HAFxLMS algorithm, the PPO reinforcement learning method is introduced to adaptively optimize the step size, enabling it to dynamically adjust according to the environment, thereby enhancing the robustness and convergence performance of the high-vibration active control system under different operating conditions.

[0101] The process of the PPO reinforcement learning algorithm is as follows: Figure 3As shown, the appropriate setting of PPO reinforcement learning algorithm parameters is crucial to the convergence and control effect of the vibration active control system during training. To ensure the stability and efficiency of training, the parameters selected, based on the characteristics of the PPO reinforcement learning algorithm, are shown in Table 1. Specifically, the PPO reinforcement learning algorithm continuously adjusts the step size of the feedforward and feedback paths to maximize the vibration suppression effect in each round of training.

[0102] Table 1. PPO Reinforcement Learning Algorithm Parameter Settings

[0103] Controller parameters Setting value Banch size 20 Discount rate 0.99 Network update rate 0.01 Maximum number of training sets 500 Maximum number of training sessions per group 30 Action network learning rate 0.0005 Policy Network Learning Rate 0.001 Actor updates steps 20 Critic network updates steps 20

[0104] Training data was collected by building an experimental system. The experiment simulated various typical external disturbance environments, such as periodic excitation and impact excitation, to comprehensively cover possible vibration scenarios. During data acquisition, sensors were used to monitor the vibration response and error signals of the main system in real time, and these signals were collected as input-output data pairs. This data was used to generate training samples, forming a multi-dimensional time-series dataset.

[0105] To ensure the model's generalization ability, the entire dataset was divided into training and validation sets. 70% of the data was used for the training set, and the remaining 30% for the validation set. The training set was used to train the PPO reinforcement learning algorithm, optimizing the step size of the feedforward and feedback paths through multiple rounds of reinforcement learning training. The validation set was used to evaluate the performance of the trained model on unseen data, ensuring the algorithm's robustness and adaptability. The training and validation processes were alternated to avoid overfitting and ensure good control performance under different vibration conditions.

[0106] The improvement goal of the PPO-HAFxLMS algorithm is to optimize the policy through reinforcement learning, maximize the cumulative reward r, and minimize the mean square value J of the error signal. e ,Right now:

[0107] (9)

[0108] The optimization goal of the PPO-HAFxLMS algorithm is to further adjust the step size μ1 of the feedforward channel and the step size μ2 of the feedback channel using PPO reinforcement learning, so that the feedforward filter and the feedback filter can adapt to different vibration environments, thereby improving the convergence speed and vibration suppression capability.

[0109] To further optimize the vibration active control algorithm by combining PPO reinforcement learning, the adaptive vibration control problem is transformed into a Markov decision process (MDP), whose main elements include state space, action space and reward function.

[0110] The state at time n should include information related to the current vibration state of the controlled object. Considering the structure of the PPO-HAFxLMS algorithm, the state space s(n) can be designed as follows:

[0111] (10)

[0112] In the formula, e(n-1) represents the historical vibration error, and x′(n) is the signal after the input signal x(n) is identified by the secondary channel. Through the information in the state space s(n), the agent in the PPO reinforcement learning algorithm can fully understand the dynamic characteristics of the controlled object in order to make appropriate control decisions.

[0113] To dynamically adjust the filter weights and update the step size μ1 of the feedforward channel and the step size μ2 of the feedback channel to adapt to different vibration environments, the action space a(n) can be defined as:

[0114] (11)

[0115] In the formula, Δμ1(n) and Δμ2(n) represent the adjustment amounts of the feedforward channel step size and the feedback channel step size generated by the PPO reinforcement learning algorithm, respectively, used to optimize the filter weight update formula:

[0116] (12)

[0117] In the formula, W 11 (n+1) and W 22 (n+1) are the coefficients of the feedforward filter and the feedback filter at time n+1 after PPO reinforcement learning compensation adjustment.

[0118] Since the control objective of the PPO-HAFxLMS algorithm is to minimize the error signal e(n), the reward function r(n) can be designed as follows:

[0119] (13)

[0120] The reward function r(n) increases the reward value as the error signal e(n) decreases, thus prompting the agent in PPO reinforcement learning to learn a control strategy that effectively reduces vibration errors. Furthermore, a penalty term for step size variation is added to prevent excessive step size fluctuations.

[0121] (14)

[0122] In the formula, λ is a trade-off factor used to balance error minimization and step size stability.

[0123] The core of the PPO reinforcement learning algorithm is the optimization policy π. θTo maximize cumulative reward while limiting the magnitude of policy updates and ensuring training stability, the optimization objective function of the PPO reinforcement learning algorithm can be expressed as:

[0124] (15)

[0125] In the formula, The importance sampling ratio, The advantage function measures the merits of the current action, and ϵ is a pruning parameter used to limit the magnitude of policy updates. The strategy is restricted to the interval [1-ϵ, 1+ϵ] to prevent excessive changes in strategy.

[0126] In the active vibration control task, the PPO-HAFxLMS algorithm adjusts the step size μ1 of the feedforward channel and the step size μ2 of the feedback channel. Therefore, the objective function can be specified as:

[0127] (16)

[0128] This objective function ensures that the step size adjustment strategy does not change drastically in a single update, thereby improving the stability of the controlled object and the speed of vibration control.

[0129] During the testing of the PPO-HAFxLMS algorithm, the input signal x(n) consists of a basic swept frequency signal (covering 5–100Hz, amplitude 1–5 mm) superimposed with random Gaussian white noise (standard deviation 2 mm). Figure 4 The curves showing the changes in cumulative reward (Episode Reward) and average reward (Average Reward) of the agent in each round of interaction during the operation of the PPO-HAFxLMS algorithm are presented. It can be observed that as the training iterations progress, the average reward shows a stable upward trend, indicating that the policy is constantly learning and improving, gradually achieving better control performance.

[0130] Furthermore, the reward curve fluctuates significantly in the early stages, indicating that the strategy is still in the exploratory phase; however, the fluctuations gradually converge in the later stages, showing that the strategy is stabilizing and the policy network has completed effective learning. Overall, this verifies that the PPO-HAFxLMS algorithm has good convergence and the ability to improve the effectiveness of active vibration control.

[0131] Figure 5 illustrates the adaptive adjustment trajectories of the step size μ1 of the feedforward channel and the step size μ2 of the feedback channel during the operation of the PPO-HAFxLMS algorithm. For the feedforward channel, μ1 maintains strong volatility throughout the training process, with the step size continuously changing with the state of the controlled object and input disturbances, reflecting the feedforward channel's high sensitivity and rapid response to external disturbances. Its continuous dynamic adjustment helps to track non-stationary inputs in complex environments in real time, enhancing the feedforward prediction and disturbance suppression capabilities of the PPO-HAFxLMS algorithm. For the feedback channel, the change curve of μ2 shows a trend from volatility to convergence. Its step size change fluctuates slightly in the early stage of training, then gradually decreases and tends to stabilize, indicating that the strategy of the feedback channel gradually converges, and the adjustment of the error signal becomes more stable and accurate. This behavior demonstrates that the feedback path mainly undertakes the function of steady-state error correction and improving the stability of the PPO-HAFxLMS algorithm.

[0132] Figure 5 The graphs showing the variation of step sizes μ1 and μ2 in this embodiment demonstrate that, overall, the step size parameters exhibit good division of labor and task sensitivity under the control mechanism of the PPO-HAFxLMS algorithm: the feedforward step size is dynamically flexible, while the feedback step size tends to be stable. The former improves the response speed and adaptability of the PPO-HAFxLMS algorithm, while the latter enhances the steady-state control effect. The two complement each other effectively through dynamic adjustment, providing superior performance assurance for the PPO-HAFxLMS algorithm under complex operating conditions.

[0133] After adjusting the step size and filter coefficients, the PPO-HAFxLMS algorithm compares whether the error signal e(n) is effectively suppressed. As the optimization process progresses, the error e(n) gradually decreases and tends to a small, stable value. This means that the controlled object has approached the optimal control state, and the vibration control effect has been effectively improved.

[0134] Finally, the stable control signal y'(n) generated by the PPO-HAFxLMS algorithm is output to the actuator to regulate the vibration of the controlled object. At this point, the vibration control effect of the controlled object has reached the expected level, the error signal e(n) remains at a low level, and the vibration suppression achieves the best effect.

[0135] To quantify the actual effects of different control strategies, the mean square error (MSE) variation curves under different vibration control methods were compared based on experimental results, and the results are plotted as follows: Figure 6As shown, the PPO-HAFxLMS algorithm exhibits the best control performance, with faster convergence speed, significantly reduced steady-state value, and MSE curve converges to -19.89dB in a short time. Compared with HAFxLMS, ​​PPO-HAFxLMS can effectively suppress the vibration of the controlled object and demonstrates strong adaptive capability when facing complex external disturbances, with a convergence speed improvement of 19.8% compared to the HAFxLMS algorithm.

[0136] It should be understood that, inspired by the technical concept of this invention, those skilled in the art can make various improvements or modifications based on the above content without departing from the scope of this invention, and these modifications still fall within the protection scope of this invention.

Claims

1. A vibration active control method based on the PPO-HAFxLMS algorithm, characterized in that, The method includes the following steps: The target vibration frequency band to be suppressed and the vibration attenuation criteria are determined, thereby determining the configuration of relevant hardware in the active vibration control system; Capture external disturbance signals received by the controlled object; Based on the time-domain data of external disturbance signals, the vibration active control system uses the PPO-HAFxLMS algorithm to generate the required control voltage signal in real time to achieve vibration suppression.

2. The vibration active control method according to claim 1, characterized in that, The configuration of the relevant hardware includes the type, quantity, and layout of sensors and actuators; the sensors include an accelerometer installed near the external disturbance signal as a reference sensor and another accelerometer installed in the vibration suppression area as an error sensor; the actuators are installed in the vibration suppression area of ​​the controlled object.

3. The vibration active control method according to claim 2, characterized in that, The method for capturing the external disturbance signal received by the controlled object is as follows: capture the external disturbance signal received by the controlled object from the acceleration signal collected by the reference sensor.

4. The vibration active control method according to claim 1, characterized in that, The vibration active control system includes: a channel P(z) between the external disturbance signal x(n) received by the controlled object at any time n and the vibration response signal measured by the error sensor; a secondary channel model S(z) between the control voltage signal and the vibration response signal collected by the sensor; a feedforward filter W1(z) for generating the feedforward control signal y1(n); and a feedback filter W2(z) for generating the feedback control signal y2(n); where z represents a complex variable in the Z-transform, y1(n) is used to deal with the vibration caused by the external disturbance, and y2(n) is used to correct the response of the vibration active control system by adjusting the error signal e(n) of the vibration response.

5. The vibration active control method according to claim 4, characterized in that, Based on the time-domain data of the input vibration signal, the active vibration control system uses the PPO-HAFxLMS algorithm to generate the required control voltage signal in real time, thereby suppressing the vibration signal, including: Step 4.1: Based on the external disturbance signal x(n) received by the controlled object at any time n, calculate the actual vibration response time-domain data d(n) of the controlled object using the HAFxLMS algorithm, and then calculate the vibration response error signal e(n). The calculation process is as follows: (1) (2) In the formula, This represents the vibration control quantity applied to the controlled object by the actuation voltage signal y(n) via the secondary channel S(z); Step 4.2: Based on the calculated error signal e(n), update the feedforward filter W1(z) and feedback filter W2(z) using the HAFxLMS algorithm to optimize the generation of the control voltage signal and reduce the vibration error of the controlled object; Step 4.3: Use the HAFxLMS algorithm to generate the control voltage signal. During the generation of the control voltage signal, the step size of the feedforward filter W1(z) and the step size of the feedback filter W2(z) are adaptively optimized using the PPO reinforcement learning algorithm.

6. The vibration active control method according to claim 5, characterized in that, The feedforward filter W1(z) is updated using the HAFxLMS algorithm as follows: (3) In the formula, W1(n) is the coefficient of the feedforward filter at time n, and W1(n+1) is the coefficient of the feedforward filter at time n+1. is the step size parameter of the feedforward filter, used to control the rate at which the feedforward filter coefficients are updated; x′(n) is the signal after the input signal x(n) is identified by the secondary channel; The feedback filter W2(z) is updated using the HAFxLMS algorithm as follows: (4) In the formula, W2(n) is the coefficient of the feedback filter at time n, and W2(n+1) is the coefficient of the feedback filter at time n+1; This is the step size parameter of the feedback filter, used to control the rate at which the feedback filter coefficients are updated; The actuation voltage signal, i.e., the control input y(n), is identified by the secondary channel S(z). The control input signal identified by the error sensor is numerically... .

7. The vibration active control method according to claim 6, characterized in that, The generation of the control voltage signal using the HAFxLMS algorithm includes: First, based on the input signal x(n) and the feedforward filter W1(z), the feedforward control signal y1(n) is generated according to equation (5): (5) In the formula, for transpose; Then, based on the error signal e(n) and the secondary channel S(z), the feedback control signal y2(n) is calculated according to equation (6): (6) In the formula, for transpose; This indicates the identification of the secondary channel S(z); Then, the feedforward control signal y1(n) and the feedback control signal y2(n) are combined to obtain the total control voltage signal y(n) acting on the controlled object: (7) Finally, the total control voltage signal y(n) is processed by the FIR filter, drives the actuator to generate control force, acts on the controlled object, and generates the final control signal y′(n) through the secondary channel S(z): (8) In the formula, s1(n) is the impulse response of the secondary channel S(z).

8. The vibration active control method according to claim 7, characterized in that, In step 4.3, during the generation of the control voltage signal, the step size of the feedforward filter W1(z) and the step size of the feedback filter W2(z) are adaptively optimized using the PPO reinforcement learning algorithm. This includes: the state space s(n) is designed as follows: (10) In the formula, e(n-1) represents the historical vibration error; The action space a(n) is defined as: (11) In the formula, Δμ1(n) and Δμ2(n) represent the adjustment amounts of the feedforward channel step size and the feedback channel step size generated by the PPO reinforcement learning algorithm, respectively, used to optimize the filter weight update formula: (12) In the formula, W 11 (n+1) and W 22 (n+1) are the coefficients of the feedforward filter and the feedback filter at time n+1 after PPO reinforcement learning compensation adjustment; The reward function r(n) is designed as follows: (14) In the formula, λ is a trade-off factor used to balance error minimization and step size stability; The optimization objective function of the PPO reinforcement learning algorithm is: (16) In the formula, Importance sampling ratio; ϵ is the advantage function, which measures the merits of the current action; ϵ is a pruning parameter used to limit the magnitude of policy updates.

Citation Information

Patent Citations

  • An adaptive active vibration control method for marine diesel engines based on differential evolution algorithm

    CN114690622B

  • Active vibration control method based on secondary channel online identification FxLMS algorithm

    CN120122516A