Strategy-action active control method considering random time delay

The strategy-action active control method constructed by the deep reinforcement learning algorithm solves the problem of traditional methods' dependence on model and time-lag effect, and achieves efficient control effect under random time-lag and external perturbations.

CN120578072APending Publication Date: 2025-09-02GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510772107.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Existing active control methods rely heavily on the accuracy of structural dynamics models and are unable to adapt to changes in structural parameters and actuator time lag effects, resulting in weakening of control effects or even instability. Data-driven methods are prone to overfitting or generalization capabilities when training data is insufficient.

Method used

A deep reinforcement learning algorithm is used to construct a strategy-action active control method for random time delay consideration. By interacting with the environment, learning the optimal control strategy, using the SAC deep reinforcement learning algorithm to train the agent, adjust the control force in real time to adapt to random time delay and external perturbations.

Benefits of technology

Robust control in the case of random time delay and external perturbation is realized, reliance on precise mathematical models is avoided, and the applicability and stability of the control system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578072A_ABST
    Figure CN120578072A_ABST
Patent Text Reader

Abstract

A strategy-action active control method considering random time delay comprises the steps that a model of a controlled structure and the installation position of an actuator are selected, and a kinetic equation of a structure vibration control system is established; adding random time lag of the actuator, and building an interaction environment of an SAC deep reinforcement learning algorithm; setting a feedback function; determining a framework and network training parameters of a deep reinforcement learning neural network by adopting an SAC algorithm; training a deep reinforcement learning agent in an interaction environment by using an SAC algorithm, and performing strategy updating iteration through an experience pool playback mechanism of the SAC algorithm and a feedback function value to obtain a mature agent; a mature intelligent agent is embedded into the kinetic equation in the first step to serve as a controller, a control force signal is directly calculated, and vibration control over the structure under the random time delay condition is achieved. According to the method, the random time lag and external interference conditions of the actuator are accurately captured through interaction with the environment, the control strategy is adjusted in real time, and the method belongs to the field of structural vibration control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of structural vibration control, and in particular to a strategy-action active control method considering random time delay. Background Art

[0002] As a cutting-edge technology for suppressing structural dynamic responses, active structural control systems, through real-time feedback and adaptive adjustment mechanisms, can precisely match the vibration suppression effect expected by users while ensuring control robustness under complex working conditions. Most existing active control methods currently rely heavily on the accuracy of structural dynamic models. As building structures age, their structural parameters will have errors, which will also cause the mathematical model of the structure to deviate from the design values, which is challenging for traditional controllers. In addition, these methods require analytical solutions based on optimal control theory and a comprehensive system description. In some active control systems, including high-speed or low-speed dynamic systems, significant time lag effects may be exhibited. This can easily lead to the failure of active control effects and even the risk of system instability. With the significant development of structural buildings, system complexity has increased significantly, which also poses challenges to traditional solutions.

[0003] Conventional active controller designs, such as linear quadratic Gaussian control, proportional-integral-derivative control, sliding mode control, and fuzzy control, are typically based on deterministic mathematical models and idealized environmental assumptions. Their control parameters are fixed upon initialization and lack dynamic adaptability. In practical engineering, time-varying system disturbances and actuator lag (signal transmission and computational delays) can cause phase mismatch in the control force, leading to pole shifts in the closed-loop system to the right half plane, potentially inducing structural resonance and even instability. Meanwhile, data-driven approaches, such as fuzzy logic controllers and traditional neural network controllers, while capable of nonlinear mapping, are highly dependent on large-scale, high-quality training data for their performance. In the field of structural vibration control, limited by sensor density, environmental noise interference, and the difficulty of reproducing extreme operating conditions, training datasets often fail to cover all potential operating conditions. This can lead to model overfitting and insufficient generalization, particularly in high-dimensional nonlinear systems. Summary of the Invention

[0004] In order to overcome the problem of weakened control effect of the controller in the presence of external disturbances and random time delays, the purpose of the present invention is to provide a strategy-action active control method considering random time delays, and to use a deep reinforcement learning algorithm to construct an active control strategy that adjusts the control performance as the external environment changes, thereby ensuring the robustness of the active control algorithm.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A strategy-action active control method considering random time delay includes the following steps:

[0007] Step 1: Select the model of the controlled structure and determine the actuator installation position; establish the dynamic equation of the structural vibration control system;

[0008] Step 2: Add random time delay to the actuator to simulate the actual working conditions of the actuator and build an interactive environment for the SAC deep reinforcement learning algorithm; use the structural dynamic response collected by the sensor to set up an appropriate feedback function;

[0009] Step 3: Use the SAC deep reinforcement learning algorithm as the training algorithm to determine the framework of the deep reinforcement learning neural network and the network training parameters;

[0010] Step 4: Use the SAC deep reinforcement learning algorithm to train the deep reinforcement learning agent in the interactive environment built in step 2. Use the SAC deep reinforcement learning algorithm's experience pool replay mechanism and feedback function value to iterate the strategy to obtain a mature agent that meets user needs;

[0011] Step 5: Use the mature intelligent agent obtained in step 4 to embed the dynamic equation of the structural vibration control system obtained in step 1 as a controller, directly calculate the control force signal, and realize the vibration control of the structure under random time delay conditions.

[0012] As a preferred embodiment, in step 1, the dynamic equation of the structural vibration control system is:

[0013]

[0014] Among them, M, C are the structural mass and damping matrix respectively; F k is the structural resilience; and X vectors represent the acceleration, velocity, and displacement of the structure, respectively; U represents the seismic record and active control force, respectively; I and H represent the unit column vector and the active control device layout matrix, respectively.

[0015] As a preferred embodiment, in step 1, the Bouw-Wen model is used to describe the structural restoring force:

[0016]

[0017] R T is the restoring force of the structure; u, are the inter-layer displacement and inter-layer velocity of the structure; k is the initial stiffness of the structure; z is the hysteresis displacement; A, β, and γ are the adjustment parameters of the Bouw-Wen model; α is the ratio of the linear stiffness to the nonlinear stiffness of the structure; and n is the order that controls the size of the transition range from elastic to inelastic.

[0018] As a preferred embodiment, in step 2, a two-way communication environment is established between the controlled structure and the actuator obtained in step 1; the dynamic response of the structure collected by the sensor is used as the observation value of the intelligent agent, and the feedback function value is calculated; the feedback function is:

[0019]

[0020] Wherein, ζ=100 is the parameter of the feedback function.

[0021] As a preferred embodiment, in step three, the framework of the SAC deep reinforcement learning algorithm includes a policy network, an action network, an action value function, and a state value function.

[0022] As a preferred embodiment, in step five, the input of the feedback function is the structural response collected by the sensor, the input of the controller is the collected structural response and the feedback function value, and the output is the next step signal of the actuator, corresponding to the control force signal.

[0023] As a preferred embodiment, in step five, when the controlled structure is subjected to external loads and the actuator itself has random time delays, the controller can adjust the output of the control force in real time according to the dynamic response of the structure.

[0024] As a preferred embodiment, the control framework includes a controlled structure, an active control device, a sensor, and a SAC controller; the active control device is installed on the controlled structure and acts on the controlled structure through an actuator; the sensor is installed on the controlled structure to obtain the dynamic response of the structure; and the SAC controller adjusts the control force signal of the actuator.

[0025] The present invention has the following advantages:

[0026] In order to improve the robustness of the active control algorithm under random time delays and external disturbances, the present invention proposes a strategy-action active control method that takes random time delays into account. The algorithm does not rely on the dynamic mathematical model in the environment, but learns the optimal strategy through direct interaction with the environment. The early training realizes the iterative update of the strategy through interaction with the environment to obtain a mature intelligent agent, which can well solve the problem of random time delay in the actuator and ensure the robustness of the active control system during operation.

[0027] Traditional active controllers require precise mathematical models of the system. In the presence of external noise disturbances and arbitrary time delays in the actuator, control failures may occur. Traditional neural network controllers require the early collection of a large amount of data to train the neural network to obtain the optimal solution, and data acquisition may have certain limitations. The strategy-action active controller proposed in the present invention does not require a large amount of preliminary data during the training process, nor does it rely on precise mathematical models. Instead, it utilizes the self-learning function of deep reinforcement learning, sets a suitable feedback function, and uses the information obtained through interaction with the real working environment to update and iterate the control strategy, thereby obtaining a mature intelligent agent (controller). This ensures the robustness and applicability of the strategy-action active controller to a certain extent.

[0028] The present invention has the following advantages:

[0029] 1. Accurately capture the random time delay of the actuator and external interference through interaction with the environment, and adjust the control strategy in real time.

[0030] 2. It does not rely on the precise mathematical model of the controlled structure and does not require a large amount of preliminary data. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a framework diagram of the strategy-action active control method considering random time delay.

[0032] Figure 2 It is a flow chart of the strategy-action active control method considering random time delay.

[0033] Figure 3 Graphs of maximum inter-story displacements on different floors according to the embodiment method.

[0034] Figure 4 Graphs of maximum acceleration on different floors according to the embodiment method are shown in FIG.

[0035] Figure 5 is the first-layer restoring force hysteresis curve of the embodiment method.

[0036] Figure 6 is the top restoring force hysteresis curve of the embodiment method.

[0037] Figure 7 1 is a time history diagram of the inter-story displacement of the first layer of the embodiment method.

[0038] Figure 8 is a top-level acceleration time history diagram of the embodiment method. DETAILED DESCRIPTION

[0039] The present invention will be further described in detail below with reference to specific implementation methods.

[0040] The present invention proposes a strategy-action active control method considering random time delay, the control framework is as follows Figure 1 As shown, it includes a controlled structure, a structural dynamic response, a SAC-based controller, an active control device, and a sensor. The method of the present invention includes the following steps 1 to 5. The specific flow chart is shown in Figure 2 The specific steps are as follows.

[0041] (1) Step 1: Select the model of the controlled structure and determine the actuator installation location; establish the dynamic equation of the structural vibration control system. The details are as follows.

[0042] In this embodiment, a five-layer nonlinear model is used as the controlled structure. The structural mass of each layer is 400 kg, the initial stiffness of each layer is 200 MN / m, and the structural damping is calculated using sharp damping. An actuator is installed on each layer of the controlled structure, and the dynamic equation of the structural vibration control system is established:

[0043]

[0044] Among them, M, C are the structural mass and damping matrix respectively; F k is the structural resilience; and X vectors represent the acceleration, velocity, and displacement of the structure, respectively; U represents the seismic record and active control force, respectively; I and H represent the unit column vector and the active control device layout matrix, respectively.

[0045] In this embodiment, the Bouc-Wen model is used to describe the structural restoring force:

[0046]

[0047] R T represents the restoring force of the structure; u, are the inter-layer displacement and inter-layer velocity of the structure; k represents the initial stiffness of the structure; z represents the hysteresis displacement; A, β, and γ represent the adjustment parameters of the Bouc_Wen model, which affect the fullness of the hysteresis loop; α is the ratio of the linear stiffness to the nonlinear stiffness of the structure; and n is the order that controls the size of the transition range from elastic to inelastic.

[0048] (2) Step 2: Add random time delay to the actuator to simulate the actual working conditions of the actuator and build an interactive environment for the SAC deep reinforcement learning algorithm; use the structural dynamic response collected by the sensor to set up a suitable feedback function. The details are as follows.

[0049] In this embodiment, a random time delay of 0-10 times the sampling time is added to the actuator working environment, and a two-way communication environment is established between the controlled structure and the actuator obtained in step 1. The structural dynamic response collected by the sensor is used as the observation value of the intelligent agent, and the feedback function value is calculated. The feedback function established in this embodiment is:

[0050]

[0051] Wherein, ζ=100 is the parameter of the feedback function.

[0052] (3) Step 3: Use the deep reinforcement learning algorithm Soft Actor-Critic (SAC) algorithm as the training algorithm to determine the neural network framework and network training parameters of deep reinforcement learning. The details are as follows:

[0053] In this embodiment, the framework of the SAC deep reinforcement learning algorithm includes a policy network, an action network, an action value function, and a state value function, wherein the hyperparameter settings involved are shown in Table 1.

[0054] Table 1. Hyperparameter settings of the SAC algorithm

[0055]

[0056] (4) Step 4: Use SAC to train the deep reinforcement learning agent in the interactive environment constructed in Step 2. Through the experience pool replay mechanism of the SAC algorithm and the value of the feedback function, the strategy is updated and iterated to obtain a mature agent that meets user needs. The details are as follows.

[0057] In this embodiment, the deep reinforcement learning agent is used as a controller to interact with the environment built in step 2. During the training process, the agent's strategy is updated and iterated through the experience pool replay mechanism of the deep reinforcement learning algorithm and the value of the feedback function, so that the agent can still maintain good control performance in complex environments, thus obtaining a mature agent.

[0058] (5) Step 5: Use the mature intelligent agent obtained in Step 4 to embed the dynamic equation of the structural vibration control system obtained in Step 1 as a controller, directly calculate the control force signal, and realize the vibration control of the structure under random time delay conditions. The details are as follows.

[0059] In this embodiment, the mature agent obtained in step 4 is embedded in the dynamic equations of the structural vibration control system obtained in step 1 to serve as a controller. The input of the feedback function is the structural response collected by the sensor, the input of the controller is the collected structural response and the feedback function value, and the output is the next step signal of the actuator, which corresponds to the control force signal. When the controlled structure is subjected to external loads and the actuator itself has random time delays, the SAC controller (mature agent) can adjust the control force output in real time based on the dynamic response of the structure, thereby ensuring the robustness and efficiency of the active control system.

[0060] In this embodiment, the external seismic records of the building structure are based on the classic 1940 El Centro earthquake in the United States. Following the above steps, a strategy that considers random time delays—an active action controller—is derived. The following is an analysis of the control effect.

[0061] In this embodiment, the maximum inter-story displacement of the structure in the uncontrolled state is 0.1264m, and the maximum acceleration is 4.0426m / s 2 When the actuator has a random time delay of 0-10 times the sampling time, the design of the controller for the nonlinear structure presents certain challenges. When the steps of the present invention are used to design the active controller, the maximum inter-story displacement of the structure is 0.0139m and the maximum acceleration is 0.9153m / s. 2 The maximum displacement is reduced by 89.00%, and the maximum acceleration is reduced by 77.36%, which shows that the SAC control algorithm proposed in the present invention still has excellent control effect in the presence of random time delay in nonlinear structures.

[0062] Figure 3 and Figure 4 The comparison diagrams are respectively the maximum inter-story displacement and maximum acceleration of different floors. It can be seen that the SAC controller has excellent control effect under the condition of random time delay. Figure 5 and Figure 6 The hysteresis curves of the structural restoring force of the first and top floors of the controlled structure in the uncontrolled state and the controlled state of the SAC controller are given. Figure 7 and Figure 8 It is a time-history comparison diagram of the inter-story displacement of the first floor of the structure and a time-history comparison diagram of the acceleration of the top floor.

[0063] The present invention utilizes a mature intelligent agent trained in enhanced deep learning as a controller in an active control system, which can sensitively capture random time delays and effectively adjust the control strategy. It is not only applicable to random time delays. Other deep learning control algorithms involving external disturbances, actuator failures, or structural parameter errors should also be covered by the claims of the present invention.

[0064] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A strategy-action active control method considering random time delay, characterized in that: The steps include: Step 1: Select the model of the controlled structure and determine the actuator installation position; establish the dynamic equation of the structural vibration control system; Step 2: Add random time delay to the actuator to simulate the actual working conditions of the actuator and build an interactive environment for the SAC deep reinforcement learning algorithm; use the structural dynamic response collected by the sensor to set up an appropriate feedback function; Step 3: Use the SAC deep reinforcement learning algorithm as the training algorithm to determine the framework of the deep reinforcement learning neural network and the network training parameters; Step 4: Use the SAC deep reinforcement learning algorithm to train the deep reinforcement learning agent in the interactive environment built in step 2. Use the SAC deep reinforcement learning algorithm's experience pool replay mechanism and feedback function value to iterate the strategy to obtain a mature agent that meets user needs; Step 5: Use the mature intelligent agent obtained in step 4 to embed the dynamic equation of the structural vibration control system obtained in step 1 as a controller, directly calculate the control force signal, and realize the vibration control of the structure under random time delay conditions.

2. The strategy-action active control method considering random time delay according to claim 1, characterized in that: In step 1, the dynamic equation of the structural vibration control system is: Among them, M, C are the structural mass and damping matrix respectively; F k is the structural resilience; and X vectors represent the acceleration, velocity, and displacement of the structure, respectively; U represents the seismic record and active control force, respectively; I and H represent the unit column vector and the active control device layout matrix, respectively.

3. The strategy-action active control method considering random time delay according to claim 2, characterized in that: In step 1, the Bouw-Wen model is used to describe the structural restoring force: R T is the restoring force of the structure; u, are the inter-story displacement and inter-story velocity of the structure; k is the initial stiffness of the structure; z is the hysteresis displacement; A, β, and γ are the adjustment parameters of the Bouw-Wen model; α is the ratio of the linear stiffness to the nonlinear stiffness of the structure; and n is the order that controls the size of the transition range from elastic to inelastic.

4. The strategy-action active control method considering random time delay according to claim 3, characterized in that: In step 2, a two-way communication environment is established between the controlled structure and the actuator obtained in step 1. The dynamic response of the structure collected by the sensor is used as the observation value of the intelligent agent, and the feedback function value is calculated. The feedback function is: Wherein, ζ=100 is the parameter of the feedback function.

5. The strategy-action active control method considering random time delay according to claim 4, characterized in that: In step three, the framework of the SAC deep reinforcement learning algorithm includes the policy network, action network, action value function and state value function.

6. The strategy-action active control method considering random time delay according to claim 1, characterized in that: In step five, the input of the feedback function is the structural response collected by the sensor, the input of the controller is the collected structural response and the feedback function value, and the output is the next step signal of the actuator, which corresponds to the control force signal.

7. The strategy-action active control method considering random time delay according to claim 6, characterized in that: In step five, when the controlled structure is subjected to external loads and the actuator itself has random time delays, the controller can adjust the output of the control force in real time according to the dynamic response of the structure.

8. The strategy-action active control method considering random time delay according to claim 1, characterized in that: The control framework includes the controlled structure, active control devices, sensors, and SAC controllers. The active control devices are installed on the controlled structure and act on it through actuators. Sensors are installed on the controlled structure to obtain the structural dynamic response. The SAC controller adjusts the control force signal of the actuator.