Adaptive Physics-Informed Learning Method for Predicting Nonlinear Seismic Response of RC Structures

Through the adaptive physical information learning method, combined with deep network and adaptive weight optimization algorithm, the problem of combining physical laws with data-driven methods in nonlinear seismic response prediction of RC structure is solved, and high-precision and robust prediction effects are achieved.

CN118940610BActive Publication Date: 2025-07-04CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410930727.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2025-07-04
Estimated Expiration
2044-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively combine the laws of physics and data-driven methods in the nonlinear seismic response prediction of RC structures, resulting in inconsistent prediction results and degradation of performance in the case of scarcity of data.

Method used

Adaptive physical information learning method is adopted, and deep network is established based on physical constraints, and the output variables of four networks are defined. Adaptive weight optimization algorithm is used for step-by-step training and iterative optimization, and combined with automatic differential technology to ensure that the model follows physical equations.

Benefits of technology

It significantly improves the prediction accuracy and robustness of the nonlinear seismic response of RC structures, reduces training time and resource requirements, and can maintain high prediction performance under data scarcity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118940610B_ABST
    Figure CN118940610B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of seismic wave prediction, and specifically discloses an adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures, including the following: S1. Establish a deep network based on physical constraints; S2. Establish an adaptive weight optimization algorithm; S3. Redefine four networks according to the adaptive optimization algorithm, and each network corresponds to a corresponding output variable, namely u, v, z, ε; First, train the i-th network to obtain the network parameter θ i,0 , and then freeze the parameter to train the subsequent i+1 networks to obtain θ i+1,0 ; Then, train the four networks in sequence and retrain them again; Update the parameter until the target loss function of all networks reaches the preset value, and the final network parameter θ can be obtained after the j-th loop iteration i,j . By adopting the technical solution of the present invention, available measurement data can be utilized and combined with physical laws to model and predict structures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of seismic wave prediction, and particularly to an adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures. Background Art

[0002] In the field of structural engineering, accurately predicting the response of reinforced concrete (RC) structures under seismic actions is crucial for ensuring the safety and robustness of buildings. With the rapid development of machine learning techniques, data-driven methods have been widely applied to learn and predict structural behavior from monitoring data. However, there are several limitations and challenges in dealing with this problem in the prior art.

[0003] Traditional data-driven methods, such as certain types of neural networks, may perform well in fitting observed data but often neglect physical laws, resulting in prediction results that may be theoretically inconsistent or unreasonable. These intelligent models are somewhat similar to black boxes, making it difficult to intuitively evaluate the accuracy of their predictions, which is unacceptable for safety-critical engineering applications.

[0004] In addition, current mainstream intelligent models rely heavily on rich data to reveal potential patterns and usually face data scarcity problems in system-level experiments due to high costs and the complexity of data acquisition, which limits the training and validation of the models.

[0005] Therefore, there is an urgent need for an adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures that can utilize available measurement data and combine physical laws to model and predict structures. Summary of the Invention

[0006] The present invention provides an adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures, which can utilize available measurement data and combine physical laws to model and predict structures.

[0007] To solve the above technical problems, the present application provides the following technical solutions:

[0008] An adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures, comprising the following:

[0009] S1. Establish a deep network based on physical constraints;

[0010] S2. Establish an adaptive weight optimization algorithm;

[0011] S3. Redefine four networks according to the adaptive optimization algorithm, each network corresponding to a respective output variable, namely u, v, z, ε; first train the i-th network to obtain network parameters θ i,0, and then freeze the parameters to train the subsequent i + 1 networks to obtain θ i+1,0 ; then, train the four networks in sequence and retrain them again; update the parameters until the target loss functions of all networks reach the preset value, and the final network parameters θ are obtained after the j-th loop iteration i,j .

[0012] The principle and beneficial effects of the basic scheme are as follows: In this scheme, a deep learning model is constructed using physical laws and available measurement data, and this model can simultaneously learn data features and physical phenomena. The automatic differentiation technique is used to accurately calculate the time derivative of the model solution to ensure that the model strictly follows the physical equations during the training process.

[0013] Develop an adaptive weight optimization algorithm that can dynamically adjust the weights of different task losses to address the imbalance between physical loss and data loss. By introducing a measure of gradient imbalance and the Soft Adapt technique, the dependence of the algorithm on the initial network parameters is reduced, and the robustness and adaptability of the algorithm are improved.

[0014] Define four independent networks, and each network is optimized for a specific output variable. Adopt a strategy of step-by-step training and freezing parameters to train each network in sequence to ensure that each network can focus on its corresponding output variable during the training process. Through the way of loop iteration, continuously update the network parameters until all networks reach the preset performance standard on the target loss function.

[0015] This scheme significantly improves the prediction accuracy of the nonlinear seismic response of RC structures by combining physical information and data-driven methods. The adaptive weight optimization algorithm enables the model to better generalize to new datasets and maintain high prediction performance even in the case of scarce data. The application of the loop iteration training strategy and transfer learning reduces the time and resources required for model training, especially when dealing with large-scale datasets. Through the adaptive weight optimization algorithm, the model has better robustness to different data features and noises and can maintain stability in the variable seismic response prediction tasks. The adopted optimization algorithm and training strategy enable the model to quickly converge to the optimal solution and shorten the model training cycle.

[0016] In summary, this scheme achieves the effect of being able to use available measurement data and combine physical laws to model and predict structures.

[0017] Furthermore, the deep network in S1 is an improved Bouc-Wen model.

[0018] Furthermore, the deep network in S1 is described by the following dynamic equation for a single-degree-of-freedom system

[0019]

[0020] where \(u\), are the elastoplastic displacement, velocity, and acceleration of the system respectively; \(m\) is the mass of the system; \(F(t)\) is the restoring force; \(f(t)\) is the external excitation force; the structural restoring force of the improved Bouc - Wen model includes the elastic restoring force and the hysteretic restoring force, and its expression is

[0021] \(F(t)=F\) el +\(F\) ELA =\(\alpha ku(t)+(1 - \alpha)kz(t)\) (2)

[0022] \(F(t)\) represents the yield force of the system; in the formula, \(\alpha\) is the ratio of the post - yield to the initial linear elastic stiffness, \(k\) is the initial linear elastic stiffness of the structure, \(z\) is the non - linear hysteretic displacement, and the lateral displacement \(u\) has the following relationship:

[0023] where \(A\), \(\beta\), \(\gamma\), and \(n\) are shape hysteresis control functions; \(A\) takes 1; \(\beta\) and \(\gamma\) are control combinations of the stiffness and strength changes and the smoothness of the hysteresis; \(n\) controls the sharpness of the yield point; \(h(z)\) is the pinching effect function; \(v\) and \(\eta\) are the strength degradation and stiffness degradation parameters respectively;

[0024]

[0025] represents the energy dissipation magnitude

[0026]

[0027] represents the pinching effect, \(sgn(-)\) is the sign function

[0028] \(\zeta_1(s)=\zeta\)

[0029] \([1 - exp(-p\varepsilon)]\), \(\zeta_2(\varepsilon)=(\psi+\delta\) s \((\varepsilon))(\lambda+\zeta_1(\varepsilon))\) (5) ψ where \(\zeta_s\) represents the total slip; \(p\) is a constant controlling the initial decline rate of the slope; \(\psi\) is the degree parameter of the pinching behavior; \(\delta\psi\) represents the dispersion rate of the pinching phenomenon; \(\lambda\) controls the changes of \(\zeta_1(\varepsilon_n)\) and \(\zeta_2(\varepsilon_n)\); Equation (5) can be transformed into a set of first - order ODEs;

[0030]

[0031] The goal of PINN or PI - LSTM is to learn the network parameters \(\theta\) and physical parameters by simultaneously training the neural network and solving the underlying physical equations

[0032] ​L is the total loss function, L d is the data loss function, L i is a component of the physical loss function;

[0033]

[0034]

[0035] where

[0036]

[0037] The key step in calculating the physical loss function is to calculate the time derivative of the model solution, and automatic differentiation is used to solve the time derivative.

[0038] Furthermore, the adaptive weight optimization algorithm in S2 represents the gradient of the loss function, representing the network parameters at the (n + l)-th iteration.

[0039] Furthermore, the network parameters at the (n + l)-th iteration in S2 are calculated using Equation 10 and Equation 11, and are assumed to remain constant as the initial value during the calculation; the training is updated using the gradient descent method as follows:

[0040] is the learning rate;

[0041] Then, applying the second-order Taylor expansion to the loss function at n, we get:

[0042]

[0043] Applying (8) to (9), we get:

[0044]

[0045] where:

[0046]

[0047] Furthermore, the following is also included in S2:

[0048] An improved algorithm based on GradNorm; is the weight of the n-th epoch, with the initial value set to 1, including the following steps:

[0049] The first step: We introduce the measure of gradient imbalance

[0050]

[0051] The second step: Soft Adapt is introduced through the use of the softmax function to quantify the impact of initialization;

[0052]

[0053] β is a Bernoulli random variable

[0054] p(β = k) = p k (1 - p) 1-k k ∈ {0,1}, and random backtracking is used to determine whether to use the loss statistics (L i,0 , L a,0 corresponding to the initial parameters) to calculate the scaling ratio

[0055] Step 3: Use the moving average in two forms

[0056]

[0057] Dynamic adjustment is achieved through the annealing factor, that is, softmax is used for approximation in the early stage, and the gradient algorithm is directly used in the later stage, reducing the influence of the initial parameters on the learning of the neighboring regions while maintaining the accuracy of the gradient algorithm

[0058] Step 4: Inspired by the GradNorm method, design an additional loss L grad to learn the weights of different task losses, but it does not participate in the backpropagation gradient update of the network layer parameters, and the gradients of other tasks can be changed to the same value through regularization

[0059]

[0060] E task is the average gradient norm of all tasks during training, || ||2 is the L2 - norm, and η is used to represent the learning rate is the ratio of the loss of the i - th task (step t) to the initial loss is the relative learning rate of the i - th task

[0061] Step 5: Update the parameters by gradient descent

[0062]

[0063] The process of the algorithm is that in each round of training, first calculate the weights of different task losses through backpropagation, and then update the network parameters Description of the Drawings

[0064] Figure 1 It is a schematic diagram of a five - degree - of - freedom (5DOF) system constructed based on the Bouc - Wen model principle

[0065] Figure 2aRelated displays of the performance of six networks (PI-NN, PhyLSTM3, PI-LSTM, PI-LSTM-G, and PI-LSTM-A) for predicting the nonlinear displacement of the SDOF Bouc-Wen model, including box plots of regression analysis of the predicted displacement time history.

[0066] Figure 2b Diagram showing the comparison of loss functions for the predicted displacement.

[0067] Figure 3a Diagram comparing the displacement predicted by LSTM with the reference value.

[0068] Figure 3b Diagram showing the comparison of hysteretic dissipation.

[0069] Figure 3c Relevant performance display diagram of the PI-LSTM-A meta-model.

[0070] Figure 4 Comparison diagram of the prediction performance of the PI-LSTM-A and PI-LSTM methods with the PI-NN and PhyLSTM3 methods for the normalized nonlinear displacement.

[0071] Figure 5 Comparison of the hysteresis curves and numerical experiment results obtained by different machine learning methods. It can be seen that the method in this paper (PI-LSTM-A) is relatively close to the numerical method results schematic diagram.

[0072] Figure 6 Relevant diagram showing the applicability of the physics-based LSTM network (PI-LSTM-A) framework in transfer learning.

[0073] Figure 7 Schematic diagram showing the reduction of the number of iterations and operation time under different combinations of structural hysteretic parameters after adopting transfer learning.

[0074] Figure 8 Relevant diagram showing the number of epochs required for training and the initial loss value during retraining on the new dataset.

[0075] Figure 9 Comparison diagram of the hysteresis curves and numerical implementation results of the (PI-LSTM-A) method under different combinations of hysteretic parameters after adopting transfer learning.

[0076] Figure 10 Relevant diagram showing the number of epochs required for training and the initial loss value during retraining on the new dataset.

[0077] Figure 11 Relevant diagram showing the use of synthetic self-noise data.

[0078] Figure 12Related illustration for using ground motion data from the PEER NGA database.

[0079] Figure 13 Schematic diagram of the structural example used to verify the model calculation effect

[0080] Figure 14a Schematic diagram of the value of the model loss function for different hyperparameter values.

[0081] Figure 14b Box plot of the regression analysis of the prediction performance of the PI-LSTM-A method and other benchmark models for normalized nonlinear displacement

[0082] Figure 15 Comparison diagram of the predicted displacement time history at the top floor of different models under representative earthquakes.

[0083] Figure 16 Schematic diagram of the ground motion training result.

[0084] Figure 17 Related illustration of the pre-trained model.

[0085] Figure 18 Schematic diagram of the details of the optimization algorithm example. Specific implementation manner

[0086] The following is a further detailed description through specific implementation manners:

[0087] The adaptive physical information learning method for predicting the nonlinear seismic response of RC structures includes the following:

[0088] Step 1: Establishment of a deep network based on physical constraints

[0089] BWBN (Bouc-Wen-Baber-Noori) is an improved Bouc-Wen model, aiming to describe the following dynamic equation of a single-degree-of-freedom system:

[0090]

[0091] where are the elastoplastic displacement, velocity, and acceleration of the system respectively; m is the mass of the system; F(t) is the restoring force; f(t) is the external excitation force. The structural restoring force of the BWBN model includes an elastic restoring force and a hysteretic restoring force, and its expression is

[0092] F(t) = F el + F ELA = aku(t) + (1 - α)kz(t) (2)

[0093] F(t) represents the yield force of the system. In the formula, α is the ratio of the post-yield to the initial linear elastic stiffness, k is the initial linear elastic stiffness of the structure, z is the non-linear hysteretic displacement, and the lateral displacement u has the following relationship:

[0094]

[0095] where A, B, γ, and n are shape hysteresis control functions. A takes 1; β and γ are the control combinations of the stiffness and strength changes of the hysteresis and the smooth strength; n controls the sharpness of the yield point; h(z) is the pinching effect function; v and η are the strength degradation and stiffness degradation parameters respectively.

[0096] v(ε) = 1 + δ v ε, η(ε) = 1 + δ η ε (4)

[0097] represents the energy dissipation;

[0098] represents the pinching effect, and sgn(-) is the sign function;

[0099] ζ1(ε) = ζ s [1 - exp(-pε)], ζ2(ε) = (ψ + δ ψ (ε))(λ + ζ1(ε)) (5)

[0100] where

[0101] ζs represents the total slip; p is the constant controlling the initial descent rate of the slope; ψ is the pinching behavior degree parameter; δψ represents the dispersion rate of the pinching phenomenon; λ controls the changes of ζ1(εn) and ζ2(εn). Equation (1) can be converted into a set of first-order ODEs as follows:

[0102]

[0103] The goal of PINN or PI-LSTM is to learn the network parameters θ and physical parameters by simultaneously training the neural network and solving the underlying physical equations L is the total loss function, L d is the data loss function, L i is a component of the physical loss function.

[0104]

[0105] where

[0106]

[0107] A key step in calculating the physical loss function is to calculate the time derivative of the model solution. The finite difference method is simple and general, but it introduces approximation errors due to sparse sampling or rapid function changes, and performing multiple function samplings and difference calculations during the solution process consumes more computational resources. For these considerations, automatic differentiation (AD) is adopted, which has proven successful in physics-informed neural networks (PINNs) used for forward and inverse analysis of nonlinear ODEs. AD is part of various deep learning frameworks, enabling physics-informed neural networks to conveniently evaluate derivatives.

[0108] Step 2: Propose a new adaptive weight optimization algorithm

[0109] The adaptive weight optimization algorithm represents the gradient of the loss function and can represent the network parameters at the (n + 1)-th iteration. We use equations (10 - 11) for calculation and assume it remains constant as the initial value during the calculation. The training is updated using the gradient descent method as follows: is the learning rate. Then, apply the second-order Taylor expansion to the loss function at n to obtain:

[0110]

[0111] Applying (8) to (9), we get:

[0112]

[0113] where:

[0114]

[0115] Physical losses are typically used to ensure that the model adheres to physical laws and usually have a much larger value range than data losses. When introduced directly, training is quickly dominated by it. Physical losses produce non-zero gradients across the domain, while data losses are only effective near data points. This difference can lead to gradient expansion, complicating convergence. Additionally, gradient calculation is affected by the loss value and the domain size. Therefore, even if the loss value is small, the gradient contribution of physical losses over a wide range of values can still dominate training. This can cause the model to get stuck in a feasible but suboptimal local optimum where physical laws are satisfied but the performance at data points is lacking. Previous researchers aimed to balance the contributions of the terms in the PINN loss function and its gradients and proposed a series of different solutions from different perspectives, including learning rate annealing, GradNorm, SoftAdapt, and multi-objective loss balancing. Learning rate annealing is overly sensitive to gradient changes, resulting in large fluctuations of some loss terms around the optimum of the objective function and making it difficult to converge. Secondly, the algorithm generates networks to optimize more easily achievable results, ignoring potentially important but more difficult-to-optimize loss functions. GradNorm addresses specific challenges posed by traditional learning rate annealing. There are still some drawbacks. The first is initialization dependence: the performance of GradNorm depends largely on the initial parameters of the network, which are set manually based on experience. This is because it uses these initial parameters to calculate the initial training error, which is used for subsequent gradient normalization. If the initialization is not ideal, it may have an adverse impact on the final performance. The second is the inability to handle extreme class imbalance: although GradNorm can handle imbalance between tasks to some extent, if there is an extreme imbalance in the class distribution between tasks, GradNorm may not be sufficient to solve this problem.

[0116] Inspired by the above work, we propose an improved algorithm based on GradNorm. is the weight at the n-th epoch, with the initial value set to 1. The method we proposed is divided into five steps:

[0117] Step 1: We introduce a measure of gradient imbalance impact

[0118]

[0119] Step 2: Introduce Soft Adapt using the softmax function to quantify the impact of initialization;

[0120]

[0121] β is a Bernoulli random variable:

[0122] p(β = k) = p k (1 - p) 1-kFor \(k\in\{0,1\}\), random backtracking is used to determine whether to use the loss statistics (\(L\) i,0 , \(L\) d,0 corresponding to the initial parameters) to calculate the scaling ratio;

[0123] Step 3: Use the moving average in two forms;

[0124]

[0125] Dynamic adjustment is achieved through the annealing factor, that is, softmax is used for approximation in the early stage, and the gradient algorithm is directly used in the later stage, reducing the influence of the initial parameters on the learning of neighboring regions while maintaining the accuracy of the gradient algorithm

[0126] Step 4: Inspired by the GradNorm method, an additional loss \(L\) grad is designed to learn the weights of different task losses, but it does not participate in the backpropagation gradient update of the network layer parameters, and the gradients of other tasks can be changed to the same value through regularization:

[0127]

[0128] \(E\) task is the average gradient norm of all tasks during training, \(\|\cdot\|_2\) is the L2-norm, \(\eta\) is used to represent the learning rate, is the ratio of the loss of the \(i\)-th task (step \(t\)) to the initial loss is the relative learning rate of the \(i\)-th task;

[0129] Step 5: Update the parameters by gradient descent:

[0130]

[0131] The process of the algorithm is that in each round of training, first calculate the weights of different task losses through backpropagation, and then update the network parameters.

[0132] Obviously, compared with the annealing algorithm, our method imposes a certain form of upper bound constraint on the gradient change through the setting of the relative learning rate, effectively alleviating the problem of optimal value oscillation commonly existing in the annealing algorithm. In addition, combining softmax and random parameters reduces the influence of initialization dependence, thereby enhancing the robustness of the method.

[0133] Based on the above adaptive optimization algorithm, we redefine four networks, each network corresponding to a corresponding output variable, namely \(u\), \(v\), \(z\), \(\varepsilon\). First, train the \(i\)-th network to obtain the network parameters \(\theta\) i,0 , and then freeze the parameters to train the subsequent \(i + 1\) networks to obtain \(\theta\) i+1,0Then, train the four networks in sequence and retrain them again. Update the parameters until the target loss functions of all networks reach the preset values. In this way, the final network parameters θ are obtained after the j-th loop iteration. i,j 。

[0134] For further

[0135] Numerical verification 1: 5-degree-of-freedom (SDOF) system

[0136] Utilize a five-degree-of-freedom (SDOF) system constructed based on the Bouc-Wen model principle (such as Figure 1 ), and then execute the standardization procedure. The accumulated dataset is divided into three subsets: a training group containing 37 samples, a validation set containing 13 samples, and a test group containing 50 samples.

[0137] Figure 2a and Figure 2b Shows the performance of six networks (such as PI-NN, PhyLSTM3, PI-LSTM, PI-LSTM-G, and PI-LSTM-A) for predicting the nonlinear displacement of the SDOF Bouc-Wen model. These prediction schemes include pure data-driven LSTM and four physics-informed models: the classical PI-NN, the PhyLSTM3 metamodel, the proposed metamodel PI-LSTM in this paper, the GradNorm optimization model (denoted as: PI-LSTM-G), and the optimization model using the adaptive weight mechanism in this paper (denoted as: PI-LSTM-A).

[0138] We conduct a regression analysis on the predicted displacement time histories of all test datasets and summarize the correlation coefficients in the Figure 2a box plots. It can be observed that most of the correlation coefficients (denoted as γ) of PI-LSTM-A, PI-LSTM-A, and PI-LSTM are greater than 0.8, indicating very accurate predictions. Several key indicators reflecting the data distribution range and dispersion predicted by PI-LSTM-A, such as the median, shape, box height, etc., are better than other methods. Observing the outliers (representing the worst performance of the model corresponding to the sample γ), the value of the LSTM method is significantly lower, at 0.61, while that of PI-LSTM is 0.75 and that of PI-LSTM-A is 0.82. The predicted displacement time histories at the lowest γ in the example, as well as the corresponding correlation coefficients marked in the box plots of PI-LSTM (γ = 0.75) and LSTM (γ = 0.61) ( Figure 2a ). Even for cases where the prediction is not very satisfactory (e.g., γ = 0.75), the PI-LSTM method can still reasonably predict the displacement time history using very limited training data.

[0139] The γ = 0.75 predicted by PI-LSTM matches well with the reference values in terms of phase, amplitude, and residual drift. The displacements predicted using LSTM are also as Figure 3a shown. Although the predicted peak amplitudes and phases of the displacements match the reference values relatively well, LSTM is unable to accurately predict the residual drift (such as plastic deformation). These characteristics hidden in the physical laws are complex and difficult to discover solely through pure data-driven training.

[0140] Figure 3b Hysteretic dissipation is shown. It can be seen that there are obvious deviations in the traditional LSTM, while the PI-LSTM with physical information is in line with the actual situation. Figure 3c and Figure 3d Similar prediction performances are observed in the PI-LSTM-A metamodel in

[0141] This indicates that it is difficult to learn complex hysteretic behaviors purely from training data, especially when the available dataset is limited.

[0142] Figure 2b The loss functions of the predicted displacements are compared. It can be seen that the proposed adaptive schemes (PI-LSTM-A) and GradNorm (PI-LSTM-G) show a faster decline in loss, which means that the gradient-based weight optimization algorithm can reduce the training cost and improve the prediction accuracy to a certain extent. During the training process, we observed that the loss function initially decreased smoothly from the random initial state of the network. However, it later showed oscillations. These oscillations can be partially alleviated by reducing the learning rate, applying parameter regularization, or increasing the batch size. However, the non-convexity of the loss function makes all these strategies significantly slow down the training speed and often lead to a decrease in accuracy.

[0143] In Figure 4 the prediction performances of the PI-LSTM-A and PI-LSTM methods for the normalized nonlinear displacements are compared with the PI-NN and PhyLSTM3 as the benchmark models. Figure 4 shows that when epch = 200, there are significant deviations between the prediction results of all four methods (epch = 1000) and the curves of the numerical results. Despite specific deviations from the numerical results, the PI-LSTM and PI-LSTM-A show better prediction performances than the PI-NN and PhyLSTM3. The results are shown in Table 1.

[0144]

[0145]

[0146] Table 1 Comparison of the loss functions (MSE and MAE) and operation times of the five algorithms

[0147] Under this background, we explored the applicability of the physics-based LSTM network (PI-LSTM-A) framework ( Figure 6 ) in transfer learning: using a pre-trained neural network to make inferences on a new dataset. It is expected that since the initial state of the neural network is no longer randomly selected, the training should converge to the solution and parameters of the data faster. To this end, we utilized a pre-trained model ( Figure 17 ) to train on a dataset with degradation factors and degradation shrinkage factors. Subsequently, we studied how the loss evolves and how the training converges when the data features are different: with four different intensity degradation factors, stiffness degradation factors: under the same squeezing effect parameters, pinching slope, pinching intensity,

[0148] As shown in the figure, re-training on the new dataset only requires a few hundred epochs and the initial loss value is small. This demonstrates the advantage of PI-LSTM, that is, the cost of re-training on similar data is much lower than the classical method relying on forward simulation. This is an important result, illustrating the robustness of the proposed method.

[0149] Figure 8 , Figure 9 Analysis shows that after combining transfer learning, the training convergence speed is significantly accelerated. Similar observations have also been obtained in previous studies using various physical information models, which indicates that the PINN method has good adaptability to transfer learning. This finding shows that transfer learning can be used to improve efficiency and cost-effectiveness, enhance domain adaptability, strengthen the adaptability to scarce or cost-effective training data, and cultivate better generalization ability.

[0150] We further verified the framework of synthetic data generated by numerical methods; the data is divided into two categories. The first category ( Figure 11 ) uses synthetic white noise data, and the second category ( Figure 12 ) uses ground motions from the PEERNGA database.

[0151] We adopted four numerical algorithms, including the 5(4)-order explicit Runge-Kutta method (RK45), the 3(2)-order explicit Runge-Kutta method (RK23), the implicit Runge-Kutta method of family IIA of order 5 (Radau), and the implicit multi-step variable-order (1 to 5) method based on the backward differential formula of derivative approximation (BDF). For ground motion training ( Figure 16 ), the training converges faster on the dataset generated by high-order numerical methods. However, for band-limited white noise ( Figure 15) For it, it contains a large number of high-frequency components and random fluctuations. Higher-order numerical algorithms such as RK45 may be more vulnerable to the interference of these high-frequency noises, resulting in unstable results. On the other hand, lower-order algorithms such as RK23 may exhibit better noise suppression ability to a certain extent, resulting in smoother and more stable results. These phenomena highlight the importance of using appropriate numerical algorithms and orders according to the data type to pre-train the ANN.

[0152] Numerical verification 2, three-layer MRF structure.

[0153] We use PI-NN, PhyLSTM3, and LSTM as benchmark models to compare the prediction performance of the PI-LSTM-A (PI-LSTM-G) method for normalized non-linear displacement.

[0154] || ||2 is the L2 norm

[0155] Calculated by automatic differentiation (AD)

[0156] We perform a regression analysis on the predicted displacement time history of all test datasets and use the Figure 14b box plots in it to illustrate the correlation coefficient. Obviously, the correlation coefficient of PI-LSTM-A is mainly greater than 0.9. Several key indicators, such as the median and the shape and height of the box reflecting the data distribution range and dispersion degree, are better than other methods. These results emphasize their robustness and excellent prediction accuracy. Observing the outliers representing the worst model performance corresponding to γ of the samples, the value of the LSTM method is significantly lower than 0.62, while the value of PI-LSTM-A is 0.75.

[0157] For further illustration, we select Figure 14b this discrete point with γ = 0.62 in it as sample A. And the LSTM and PI-LSTM-A methods below the first quartile (Q1) are randomly selected as sample B, and the predicted displacement time history of the top floor under seismic action is plotted.

[0158] Figure 15 The predicted displacement time history of the top floor under two representative earthquakes is shown, highlighting the least satisfactory part of the correlation coefficient of LSTM ((γ = 0.65 for (a)-(c), γ = 0.62 for (d)-(f)), corresponding to PI-LSTM-G ((γ = 0.80 for (a)-(c), γ = 0.74 for (d)-(f)), corresponding to PI-LSTM-A ((γ = 0.83 for (a)-(c) and γ = 0.78 for (d)-(f)). The prediction curve of PI-LSTM-A with γ = 0.83 for (a)-(c) is very close to the experimental data illustrating the amplitude, phase, and residual drift of plastic deformation.

[0159] Even in the worst-case scenario, such as γ = 0.78 in (d)-(f), PI-LSTM-A reasonably demonstrates its ability to predict displacement time histories given limited training data. The PI-LSTM-A meta-model also has a similar performance level. Figure 15 The normalized structural displacements predicted by the LSTM shown exhibit reasonable predictions of the displacement peak amplitude and phase; however, it fails to accurately predict the residual drift (e.g., plastic deformation), which shows a significant deviation from the experimental results, indicating that it is difficult to learn complex hysteretic behavior purely from limited training data.

[0160] The above are only embodiments of the present invention. The invention is not limited to the fields involved in this embodiment. Common knowledge such as the specific structures and characteristics known in the art is not described in detail here. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention belongs before the filing date or the priority date, are able to learn all the existing technologies in this field, and have the ability to apply conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, combine their own abilities to improve and implement this solution. Some typical well-known structures or well-known methods should not become obstacles for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can still be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners described in the specification can be used to interpret the content of the claims.

Claims

1. An adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures, characterized in that: It includes the following content: S1. Establish a deep network based on physical constraints; S2. Establish an adaptive weight optimization algorithm; S3. Four networks are redefined according to the adaptive optimization algorithm, and each network corresponds to a corresponding output variable, namely u, v, z, ε ; First, train the i-th network to obtain the network parameter θ i,0 , and then freeze the parameters to train the subsequent i+1 networks to obtain θ i+1,0 ; Then, train the four networks in turn and retrain them again; update the parameters until the target loss functions of all networks reach the preset values, and the final network parameter θ is obtained after the j-th loop iteration i,j ; Among them, the deep network in S1 is an improved Bouc-Wen model; The deep network in S1 is described by the following dynamic equation for a single-degree-of-freedom system: where u, are the elastoplastic displacement, velocity, and acceleration of the system, respectively; m is the mass of the system; F(t) is the restoring force; f(t) is the external excitation force; the structural restoring force of the improved Bouc-Wen model includes the elastic restoring force and the hysteretic restoring force, and its expression is F(t) = F el + F ELA = αku(t) + (1 - α)kz(t) (2) F(t) represents the yield force of the system; in the formula, α is the ratio of the post-yield to the initial linear elastic stiffness, k is the initial linear elastic stiffness of the structure, z is the nonlinear hysteretic displacement, and the lateral displacement u has the following relationship: Where A, β, γ, and n are shape hysteresis control functions; A takes 1; β and γ are control combinations of the stiffness and strength changes and the smoothness of the strength of the hysteresis; n controls the sharpness of the yield point; h(z) is the pinching effect function; ν and η are the strength degradation function and the stiffness degradation parameter, respectively.

2. The adaptive physical information learning method for predicting the nonlinear seismic response of RC structures according to claim 1, characterized in that: v(ε) = 1 + δ v ε, η(ε) = 1 + δ η ε (4) Indicates the energy consumption Indicates the pinching effect, and sgn(-) is the sign function ζ1(ε) = ζ s [1 - exp(-pε)], ζ2(ε) = (ψ + δ ψ (ε))(λ + ζ1(ε)) (5) ζs represents the total slip; p is a constant controlling the initial descent rate of the slope; ψ is the pinching behavior degree parameter; δψ represents the dispersion rate of the pinching phenomenon; λ controls the changes of the parameters ζ1(εn) and ζ2(εn); Equation (5) can be converted into a set of first-order ODEs; The goal of PINN or PI-LSTM is to learn the network parameters θ and physical parameters by simultaneously training the neural network and solving the underlying physical equations. L is the total loss function, L d is the data loss function, L i is a component of the physical loss function; Where The key step in calculating the physical loss function is to calculate the time derivative of the model solution, and automatic differentiation is used to solve the time derivative.

3. The adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures according to claim 2, wherein: The adaptive weight optimization algorithm in S2 represents the gradient of the loss function and represents the network parameters at the (n + 1)-th iteration.

4. The adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures according to claim 3, characterized in that: The network parameters at the (n + 1)-th iteration in S2 are calculated using Equation 10 and Equation 11, and are assumed to be constant as the initial value during the calculation process; the training is updated using the gradient descent method as follows: is the learning rate; Then, apply the second-order Taylor expansion to the loss function at n to obtain: Applying (8) to (9), we get: Where:

5. The adaptive physics-informed learning method for predicting the nonlinear seismic response of RC structures according to claim 4, characterized in that: The following content is also included in S2: An improved algorithm based on GradNorm; is the weight of the n-th epoch, and the initial value is set to 1, including the following steps: Step 1: We introduce the influence of the measure of gradient imbalance The second step: Introduce Soft Adapt through the softmax function to quantify the influence of the initialization; β is a Bernoulli random variable p(β = k) = p k (1 - p) 1-k k ∈ {0, 1}, randomly backtrack to decide whether to use the loss statistics (L i,0 , L d,0 corresponding to the initial parameters) to calculate the scaling ratio, The third step: Use the moving average in two forms; Dynamic adjustment is achieved through the annealing factor, that is, softmax is used for approximation in the early stage, and the gradient algorithm is directly used in the later stage, reducing the influence of the initialization parameters on the learning in the neighboring region while maintaining the accuracy of the gradient algorithm; Step 4: Inspired by the GradNorm method, design an additional loss L grad to learn the weights of different task losses, but it does not participate in the backpropagation gradient update of the network layer parameters, and the gradients of other tasks can be changed to the same value through regularization: E task is the average gradient norm of all tasks during training, || ||2 is the L2-norm, and η is used to represent the learning rate, is the ratio of the loss of the i-th task, i.e., at step t, to the initial loss is the relative learning rate of the i-th task; The fifth step: Update the parameters through gradient descent: The process of the algorithm is that in each round of training, first calculate the weights of different task losses through backpropagation, and then update the network parameters.

Citation Information

Patent Citations

  • Multivariate parameter coupling data-knowledge dual-drive seismic inversion method

    CN118295016A

  • Methods, systems, and computer readable media for causal training of physics-informed neural networks

    WO2023172408A2