Structure active control random time delay compensation method based on deep reinforcement learning
Through the combination of deep reinforcement learning and traditional controllers, the robustness problem of the active control system under random time lag and external disturbance is solved, and efficient control effect is achieved in complex environments.
Patent Information
- Application Number
- CN202510773607.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-08-15
AI Technical Summary
The existing active control methods rely highly on accurate structural dynamic models, which are difficult to adapt to complex operating conditions of nonlinear and time-varying parameters, and the random time-delay effect leads to a decrease in control accuracy. Traditional neural network compensators are prone to cause model overfitting and generalization capabilities to decrease in operating conditions in complex multi-degree of freedom systems.
Using a deep reinforcement learning method, train the agents through interaction with the environment, capture random time lag and external perturbations, use the SAC deep reinforcement learning algorithm to update the strategy, build a dual-channel control strategy to achieve dynamic compensation, and integrate traditional controllers to ensure robustness.
In complex environments, the efficient control performance of the active control system is realized, the robustness is improved, and the impact of random time lag and external disturbances on the control effect is reduced.
Smart Images

Figure CN120491478A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of structural vibration control, and in particular to a random time-delay compensation method for active structural control based on deep reinforcement learning. Background Art
[0002] Active structural control systems are the optimal method for suppressing structural responses, allowing for the adjustment of structural vibration control effects according to user expectations. However, current active control methods are highly dependent on the accuracy of structural dynamic models. Over the long term, building structures experience time-varying parameters due to factors such as material aging and the accumulation of load history. This causes the mathematical model to gradually deviate from the real physical system, leading to a degradation in controller robustness. Furthermore, traditional control law design relies on analytical solutions from optimal control theory, requiring the system state matrix to be completely known and strongly convex, making it difficult to adapt to complex operating conditions with nonlinear and time-varying parameters. Actuator signal transmission and computational delays introduce significant time lags, resulting in a phase mismatch between the active control force and the structural response. This can weaken the vibration suppression effect at best, or even lead to closed-loop system instability at worst. As modern buildings develop towards ultra-high, large-span, and intelligent structures, the system dimensionality and degree of nonlinearity are growing exponentially. Traditional linear control theories based on simplified models are no longer able to meet the demands of high-reliability control.
[0003] At present, most active control compensation methods, such as phase shifting method, Taylor series expansion, fuzzy network compensator, etc., have the following shortcomings: they are highly dependent on accurate dynamic models during design and use, but in actual applications, system model mismatch (such as time-varying material parameters, degradation of connection node stiffness, random environmental disturbances) and time-varying lag effects will lead to compensation phase offset and gain attenuation, and the control accuracy will be significantly reduced; although traditional neural network compensators can handle nonlinear coupling problems, their performance is limited by the completeness and representativeness of the training data set, especially in complex multi-degree-of-freedom systems, which can easily cause model overfitting and reduced working condition generalization ability, making it difficult to meet high real-time compensation requirements. Summary of the Invention
[0004] In response to the problem of controller performance degradation caused by random time delay, the purpose of the present invention is to provide a structural active control random time delay compensation method based on deep reinforcement learning, which combines the compensator and the initial controller to capture the disturbance in the environment through interaction with the environment, thereby dynamically compensating the control signal to maintain the robustness of the original controller under complex working conditions.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A method for compensating stochastic time delay in active structural control based on deep reinforcement learning, comprising the following steps:
[0007] Step 1: Select the model of the controlled structure and determine the installation position of the actuator in the active control device; establish the dynamic equation of the structural vibration control system;
[0008] Step 2: Determine the preliminary controller designed according to the mathematical model of the precise structure;
[0009] Step 3: Add random time delay to the active control device's actuator to simulate real-world actuator operation, establishing a two-way communication environment between the controlled structure and the actuator. Use the structural dynamic response captured by the sensor to set a feedback function that meets user needs.
[0010] Step 4: Determine the deep reinforcement learning algorithm, neural network framework, and network training parameters; build a training interface between the agent and the environment in step 3;
[0011] Step 5: Use the SAC deep reinforcement learning algorithm to train the intelligent agent. Based on the feedback function of the SAC deep reinforcement learning algorithm and the expected cumulative reward, the agent's online policy gradient update is performed to ultimately obtain a mature intelligent agent that meets user needs.
[0012] Step 6: The mature intelligent agent trained in step 5 is embedded into the dynamic equation of the structural vibration control system obtained in step 1 as a control compensator to perform compensation calculation of the control force; at the same time, the feedback function of the mature intelligent agent is calculated using the structural response, thereby adjusting the control compensation signal to achieve vibration control of the controlled structure under random time delay conditions.
[0013] As a preferred embodiment, in step one, the controlled structure adopts a shear model that considers boundary constraints and assumes that the stiffness of the horizontal floor is infinite; the vertical and rotational degrees of freedom of each node of the frame are ignored, and only the horizontal degrees of freedom of each layer are considered; the original finite element model is reduced in complexity through static condensation.
[0014] As a preferred embodiment, in step 2, a controller device is installed on the top of the controlled structure, and a linear quadratic Gaussian controller is designed as a preliminary controller under the ideal state of the actuator.
[0015] As a preferred embodiment, in step three, the dynamic system of the controlled structure obtained in step one is used to build an interactive environment that conforms to the actual working state; the structural dynamic response collected by the sensor is used as the observation value of the intelligent agent and the input of the feedback function.
[0016] As a preferred embodiment, in step four, the network training algorithm adopts the SAC deep reinforcement learning algorithm; the neural network framework, the number of neuron layers and the network training parameters are determined according to the complexity of the selected controlled structure model to establish the interaction between the intelligent agent and the environment in the deep reinforcement learning algorithm; the framework of the SAC deep reinforcement learning algorithm includes a policy network, an action network, an action value function and a state value function.
[0017] As a preferred approach, in step five, the interaction information between the agent and the environment is used for training and stored in the experience pool of the SAC deep reinforcement learning algorithm. By continuously iteratively updating and optimizing the agent's policy network, the agent can adjust the control compensation signal in real time according to environmental changes in complex environments, allowing the active control system to maintain good control performance. An active control system refers to a control system that integrates a compensator and a traditional controller.
[0018] As a preferred approach, in step six, a dual-channel control strategy is constructed in the system closed loop: the main channel adopts an LQG controller to achieve linear quadratic optimal tracking; the auxiliary channel deploys a SAC agent to adjust the output compensation control force in real time through online policy gradient updates.
[0019] As a preferred embodiment, the dynamic equation of the structural vibration control system is:
[0020]
[0021] Among them, M, K, C are the structural mass, stiffness and damping matrices respectively; vector X represents the acceleration, velocity and displacement of the structure respectively: U represents the seismic record and active control force, respectively; I and H represent the unit column vector and the active control device layout matrix, respectively.
[0022] As an optimization, the feedback function is:
[0023]
[0024] Among them, ζ = 1e4 is the parameter of the feedback function; M, K are the mass and stiffness of the structure respectively; vector X represents the acceleration and displacement of the structure respectively; M and K represent the mass and stiffness of the structure respectively.
[0025] As a preference, the control framework includes a controlled structure, a sensor, an active control device, a traditional controller, and an active control compensator based on deep reinforcement learning.
[0026] The principle of the present invention is:
[0027] To improve the robustness of active control systems against random time delays and external disturbances, this paper proposes a deep reinforcement learning-based random time delay compensation method for structured active control. This random time delay compensation method updates the compensator strategy based on interaction with the environment, utilizing the feedback function and cumulative reward expectation of the SAC deep reinforcement learning algorithm. This method dynamically compensates for external disturbances and random time delays, thereby ensuring effective control of the active control system during operation.
[0028] The present invention constructs a new control architecture driven by data-model hybrid and uses the intelligent agent in deep reinforcement learning as a compensator of the original controller. It can accurately capture the random time delay in the actual environment and compensate the original controller to ensure efficient control of the active control system. It is not only applicable to random time delay. Other deep learning adjustment controllers involving external disturbances, actuator failures or structural stiffness parameter perturbations should be covered by the scope of the claims of the present invention.
[0029] The present invention has the following advantages:
[0030] 1. The compensator is trained using the SAC deep reinforcement learning algorithm, which does not rely on an idealized precise dynamics model and does not require a large amount of preliminary data.
[0031] 2. Real-time interaction with the constructed real environment can accurately capture the random time delay and external interference between the controlled structure and the actuator.
[0032] 3. Utilizing the feedback function and cumulative reward expectation of the SAC deep reinforcement learning algorithm, the compensation force output of the compensator is adjusted in real time, thereby ensuring the control performance of the controller in complex environments to a certain extent and ensuring the robustness of the active control system. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 This is a framework diagram of the random time-delay compensation method for structured active control based on deep reinforcement learning.
[0034] Figure 2 This is a flow chart of the random time-delay compensation method for structured active control based on deep reinforcement learning.
[0035] Figure 3 is the maximum inter-story displacement of each layer of the structure of the embodiment.
[0036] Figure 4 is the maximum acceleration of each layer of the structure of the embodiment.
[0037] Figure 5 FIG. 1 is a time history diagram of the inter-story displacement response of the 15th floor of the structure of the embodiment.
[0038] Figure 6 4 is a time history diagram of the acceleration response of the first floor of the structure of the embodiment.
[0039] Figure 7 is a time history diagram of the control force of the embodiment. DETAILED DESCRIPTION
[0040] The present invention will be further described in detail below with reference to specific implementation methods.
[0041] A random time-delay compensation method for active structural control based on deep reinforcement learning. The overall control framework is as follows: Figure 1 As shown, it includes a controlled structure, a sensor, an active control device, a traditional controller, and an active control compensator based on deep reinforcement learning.
[0042] A method for random time-delay compensation of active structural control based on deep reinforcement learning includes steps 1 to 6. The specific flow chart is shown in Figure 2 , each step will be introduced in detail below.
[0043] (1) Step 1: Select the model of the controlled structure, determine the installation position of the actuator in the active control device, and establish the dynamic equation of the structural vibration control system. The details are as follows.
[0044] In this example, the controlled structure is a twenty-story frame structure with an active mass damper installed at the top. For ease of analysis, a shear model with boundary constraints is used, assuming infinite stiffness for the horizontal floor slabs. The vertical and rotational degrees of freedom at each frame node are ignored, and only the horizontal degrees of freedom at each floor are considered. The original finite element model was statically solidified to reduce its complexity, retaining only 20 translational degrees of freedom.
[0045] Table 1 lists the mass and stiffness parameters of the controlled structure. The dynamic equation of the structural vibration control system is established:
[0046]
[0047] Among them, M, K, C are the structural mass, stiffness and damping matrices respectively; vector X represents the acceleration, velocity and displacement of the structure respectively: U represents the seismic record and active control force, respectively; I and H represent the unit column vector and the active control device layout matrix, respectively.
[0048] Table 1 Structural parameter information
[0049]
[0050] (2) Step 2: Determine a preliminary controller designed based on the mathematical model of the precise structure. The details are as follows. In this embodiment, a controller device is installed on the top of the structure, and a Linear Quadratic Gaussian (LQG) controller is designed as the initial active controller under ideal actuator conditions.
[0051] (3) Step 3: Add random time delay to the active control device's actuator to simulate the actual actuator's operating conditions, establishing a two-way communication environment between the controlled structure and the actuator. Use the structural dynamic response collected by the sensor to set a feedback function that meets user needs. The details are as follows.
[0052] In this embodiment, a random time delay of 0-20 times the sampling time is set, and the dynamic system of the controlled structure obtained in step 1 is used to build an interactive environment that conforms to the actual working state; the structural dynamic response collected by the sensor is used as the observation value of the intelligent agent and the input of the feedback function. The feedback function established in this embodiment is:
[0053]
[0054] Wherein, ζ=1e4 is the parameter of the feedback function.
[0055] (4) Step 4: Determine the deep reinforcement learning algorithm, neural network framework, and network training parameters; build a training interface between the agent and the environment in step 3. The details are as follows.
[0056] In this embodiment, the network training algorithm uses the Soft Actor-Critical (SAC) algorithm. The neural network framework, number of neuron layers, and network training parameters are determined based on the complexity of the selected structural model to establish the interaction between the agent and the environment in the deep reinforcement learning algorithm. In this embodiment, the SAC deep reinforcement learning algorithm framework includes a policy network, an action network, an action-value function, and a state-value function. The hyperparameter settings of the SAC deep reinforcement learning algorithm are shown in Table 2.
[0057] Table 2 Hyperparameter settings of the SAC algorithm
[0058]
[0059] (5) Step 5: Use the SAC deep reinforcement learning algorithm to train the intelligent agent. Based on the feedback function of the SAC deep reinforcement learning algorithm and the expected cumulative reward, the intelligent agent is updated online with a gradient strategy, ultimately obtaining a mature intelligent agent that meets user needs. The details are as follows.
[0060] In this embodiment, SAC is used as the core training algorithm for the agent. Information about the interaction between the agent and the environment is used for training and stored in the SAC algorithm's experience pool. By continuously iteratively updating and optimizing the agent's policy network, the agent can adjust control compensation signals in real time according to environmental changes in complex environments, ensuring that the active control system maintains good control performance.
[0061] (6) Step 6: The mature agent trained in Step 5 is embedded into the dynamic equations of the structural vibration control system obtained in Step 1 as a control compensator to perform compensation calculations on the control force. Simultaneously, the feedback function of the mature agent is calculated using the structural response, thereby adjusting the control compensation signal to achieve vibration control of the controlled structure under random time delay conditions. The details are as follows.
[0062] In this embodiment, a fully trained, mature agent serves as the compensation module for the initial controller, LQG. A dual-channel control strategy is constructed within the closed-loop system: the primary channel employs an LQG controller for linear quadratic optimal tracking; the secondary channel deploys a SAC agent, which adjusts the output compensation control force in real time through online policy gradient updates. This ensures that the overall control system maintains the robustness and efficiency of the active control system in the operating environment, even with the influence of random time delays in the actuators themselves.
[0063] In this example, the building structure is subjected to external earthquake motion recorded from the 1995 Great Hanshin Earthquake in Japan. Based on the above steps, a deep reinforcement learning-based stochastic time-delay compensation method for active structural control is developed. Compared with the traditional LQG control algorithm that does not consider a compensator, the control effect is as follows.
[0064] To intuitively demonstrate the effectiveness of the present invention's deep reinforcement learning-based active structural control random time-delay compensation method, Table 3 shows the maximum interstory displacement and acceleration of the controlled structure under Kobe wave excitation without active control, LQG control, and with the addition of a SAC compensator. As can be seen from Table 3, compared to the uncontrolled structure, under random time-delay conditions, the maximum control force of the LQG control algorithm has reached the actuator threshold of 1e6, and its maximum interstory displacement has only decreased by 3.82%, but the acceleration has increased by 11.59. After adopting the SAC compensator, the maximum interstory displacement of the structure is reduced by 92.71%, and the maximum acceleration is reduced by 62.72%, with the maximum control force being 1.5% of the actuator threshold.
[0065] Table 3 Maximum response of the floor
[0066]
[0067] Figure 3 and Figure 4The following are comparison diagrams of the maximum inter-story displacement and maximum acceleration of different floors. It can be seen that the maximum inter-story displacement of the LQG controller under random time delay is relatively close to that of the uncontrolled state, and the active control system with the addition of the SAC compensator still has excellent control effect under random time delay.
[0068] Since the maximum inter-story displacement occurs at the 15th floor and the maximum acceleration is at the top floor of the structure, Figure 5 The comparison diagram of inter-story displacement of the 15th floor of the controlled structure under random time delay is given; Figure 6 A comparison diagram of the top acceleration history of the controlled structure under random time delay is given. Figure 7 It is a time history comparison chart of the control force of the active control device.
[0069] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.
Claims
1. A method for compensating random time delays in active structural control based on deep reinforcement learning, characterized in that: The following steps are involved: Step 1: Select the model of the controlled structure and determine the installation position of the actuator in the active control device; establish the dynamic equation of the structural vibration control system; Step 2: Determine the preliminary controller designed according to the mathematical model of the precise structure; Step 3: Add random time delay to the active control device's actuator to simulate real-world actuator operation, establishing a two-way communication environment between the controlled structure and the actuator. Use the structural dynamic response captured by the sensor to set a feedback function that meets user needs. Step 4: Determine the deep reinforcement learning algorithm, neural network framework, and network training parameters; build a training interface between the agent and the environment in step 3; Step 5: Use the SAC deep reinforcement learning algorithm to train the intelligent agent. Based on the feedback function of the SAC deep reinforcement learning algorithm and the expected cumulative reward, the agent's online policy gradient update is performed to ultimately obtain a mature intelligent agent that meets user needs. Step 6: The mature intelligent agent trained in step 5 is embedded into the dynamic equation of the structural vibration control system obtained in step 1 as a control compensator to perform compensation calculation of the control force; at the same time, the feedback function of the mature intelligent agent is calculated using the structural response, thereby adjusting the control compensation signal to achieve vibration control of the controlled structure under random time delay conditions.
2. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 1, characterized in that: In step 1, the controlled structure adopts a shear model that considers boundary constraints and assumes that the stiffness of the horizontal floor is infinite; the vertical and rotational degrees of freedom of each node of the frame are ignored, and only the horizontal degrees of freedom of each layer are considered; the original finite element model is reduced in complexity through static condensation.
3. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 1, characterized in that: In step 2, a controller device is installed on the top of the controlled structure, and a linear quadratic Gaussian controller is designed as a preliminary controller under the ideal state of the actuator.
4. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 1, characterized in that: In step three, the dynamic system of the controlled structure obtained in step one is used to build an interactive environment that conforms to the actual working state; the structural dynamic response collected by the sensor is used as the observation value of the intelligent agent and the input of the feedback function.
5. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 1, characterized in that: In step 4, the network training algorithm adopts the SAC deep reinforcement learning algorithm; according to the complexity of the selected controlled structure model, the neural network framework, the number of neuron layers and the network training parameters are determined to establish the interaction between the intelligent agent and the environment in the deep reinforcement learning algorithm; The framework of the SAC deep reinforcement learning algorithm includes policy network, action network, action value function and state value function.
6. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 1, characterized in that: In step five, the interaction information between the agent and the environment is used for training and stored in the experience pool of the SAC deep reinforcement learning algorithm. By continuously iteratively updating and optimizing the agent's policy network, the agent can adjust the control compensation signal in real time according to environmental changes in complex environments, allowing the active control system to maintain good control performance.
7. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 1, characterized in that: In step six, a dual-channel control strategy is constructed in the system closed loop: the main channel adopts the LQG controller to achieve linear quadratic optimal tracking; the auxiliary channel deploys the SAC intelligent agent to adjust the output compensation control force in real time through online policy gradient updates.
8. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 2, characterized in that: Dynamic equations of structural vibration control system: Among them, M, K, C are the structural mass, stiffness and damping matrices respectively; vector X represents the acceleration, velocity and displacement of the structure respectively: U represents the seismic record and active control force, respectively; I and H represent the unit column vector and the active control device layout matrix, respectively.
9. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 4, characterized in that: The feedback function is: Among them, ζ = 1e4 is the parameter of the feedback function; M, K are the mass and stiffness of the structure respectively; vector X represents the acceleration and displacement of the structure respectively.
10. The method for compensating for random time delays in active structural control based on deep reinforcement learning according to claim 1, characterized in that: The control framework includes the controlled structure, sensors, active control devices, traditional controllers, and active control compensators based on deep reinforcement learning.