Big data and random forest based most unfavorable ground motion selection method and system
Patent Information
- Application Number
- CN202410115939.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-01-26
AI Technical Summary
[0006]针对现有技术的以上缺陷或改进需求,本发明提供了一种基于大数据与随机森林的最不利地震动选取方法及系统,用于解决现有结构抗震分析中没有充分考虑地震动不确定性的缺点
[0017]1.本发明提出的方法可以在少量的计算的前提下获取到结构在数据集中所有地震记录下的地震响应,并依次计算出结构地震动需求参数DM指标。通过地震动需求参数DM指标排序,可以判断结构在海量地震动记录下的包络式响应,充分考虑了地震动的不确定性。
Smart Images

Figure CN118033729B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of earthquake measurement technology, and more specifically, relates to a method and system for selecting the most unfavorable ground motion based on big data and random forest. Background Technology
[0002] Earthquakes can cause severe casualties and economic losses. To ensure structural safety, nonlinear time-history analysis is used in civil engineering to calculate the seismic response of structures. This calculation requires selecting earthquake time-history records as input and evaluating the structure based on the time-history analysis results. Therefore, ensuring that the selected earthquake records have greater destructive potential and are more representative has become a major challenge in the field of seismic engineering.
[0003] Initially, methods for selecting seismic ground motion records were based on setting earthquakes and their intensities. For example, given the magnitude, epicentral distance, and peak value of a ground motion, the records were directly filtered from a seismic ground motion record database. However, this method resulted in seismic responses with high dispersion. To account for the uncertainty of ground motions, existing technologies propose the most unfavorable design ground motion, which involves selecting as many ground motions as possible for nonlinear time history calculations and then identifying the ground motion record with the greatest destructive potential based on the calculation results.
[0004] The relationship between seismic intensity parameters (IM) and seismic demand parameters (DM) can reflect the destructive potential of seismic motion. Seismic intensity parameters (IM) can be easily obtained from the seismic motion process, but seismic demand parameters (DM) require nonlinear time history calculations using the finite element method, which suffers from problems such as excessive calculation time and difficulty in convergence.
[0005] Many machine learning and deep learning-based methods have provided solutions for the rapid prediction of structural seismic response. However, because seismic motion has strong uncertainty, and many current seismic wave selection methods are limited by the computational cost of nonlinear time history analysis, they can only select seismic motion through the IM index, which cannot reflect the true response characteristics of the structure and does not fully consider the uncertainty of seismic motion. Summary of the Invention
[0006] In view of the above-mentioned defects or improvement needs of existing technologies, this invention provides a method and system for selecting the most unfavorable ground motion based on big data and random forest, which is used to solve the shortcomings of existing structural seismic analysis that does not fully consider the uncertainty of ground motion.
[0007] To achieve the above objectives, according to one aspect of the present invention, a method for selecting the most unfavorable ground motion based on big data and random forest is provided. The method includes: S1: performing nonlinear time history analysis on a finite element model of the structure under test using a preset number of ground motions as input to obtain the ground motion demand parameter DM of the structure under test; and mapping the ground motion intensity parameter IM of the ground motions to the ground motion demand parameter DM to obtain a training dataset; S2: normalizing the training dataset to obtain an initial dataset, and dividing the initial dataset into a training set and a test set; S3: dividing the training set into multiple parts, using each part as a validation set in turn, and using the rest as the training set for cross-validation, and then using the predicted value D obtained from each cross-validation... The maximum value of the average correlation coefficient between M and the true value DM is used as the objective function of the Bayesian optimization algorithm; S4: Construct the acquisition function of the Bayesian optimization algorithm and construct a surrogate model of the Bayesian optimization algorithm based on the objective function; S5: Preset multiple sets of random forest hyperparameters, and the acquisition function collects the random forest hyperparameters as input to the surrogate model for iterative optimization to obtain the optimal random forest hyperparameters, and then obtain the optimized random forest model; S6: Train the optimized random forest model using training sets respectively; S7: Input the undetermined ground motion into the trained optimized random forest model in sequence, sort the obtained ground motion demand parameters DM, and the ground motion corresponding to the maximum ground motion demand parameter DM is the target ground motion.
[0008] Preferably, step S1 further includes dimensionality reduction of the seismic intensity parameter IM.
[0009] Preferably, Pearson correlation coefficient is used for correlation analysis to achieve dimensionality reduction of the seismic intensity parameter IM. The formula for calculating the Pearson correlation coefficient is as follows:
[0010]
[0011] Where, ρ x,y Let x be the correlation coefficient between two IM parameters x and y, and Cov(·) be the covariance. x Let σ be the mean of the IM parameter x in the dataset. y This represents the mean of the IM parameter y in the dataset.
[0012] Preferably, step S1 further includes normalizing the seismic intensity parameter IM and the seismic demand parameter DM.
[0013] Preferably, step S6 further includes testing the trained optimized random forest model using a test set to obtain test ground motion demand parameters DM, comparing the test ground motion demand parameters DM with the ground motion demand parameters DM in the test set, and improving the random forest model if the error is greater than a preset value.
[0014] Preferably, the root mean square error or coefficient of determination is used to compare the test ground motion demand parameter DM with the ground motion demand parameter DM in the test set.
[0015] This application further provides a system for selecting the most unfavorable ground motion based on big data and random forest. The system includes: an analysis module for performing nonlinear time history analysis on a finite element model of the structure under test, using a preset number of ground motions as input, to obtain the ground motion demand parameter DM of the structure under test, and mapping the ground motion intensity parameter IM of the ground motions to the ground motion demand parameter DM to obtain a training dataset; a normalization module for normalizing the training dataset to obtain an initial dataset, and dividing the initial dataset into a training set and a test set; and an objective function acquisition module for dividing the training set into multiple parts, using each part as a validation set in turn, and using the rest as the training set for cross-validation, and comparing the predicted value DM obtained from each cross-validation with the true value DM. The average value of the coefficients is used as the objective function of the Bayesian optimization algorithm; the surrogate model acquisition module is used to construct the acquisition function of the Bayesian optimization algorithm and construct the surrogate model of the Bayesian optimization algorithm based on the objective function; the iterative optimization module is used to preset multiple sets of random forest hyperparameters, and the acquisition function collects the random forest hyperparameters as the input of the surrogate model for iterative optimization to obtain the optimal random forest hyperparameters, and then obtain the optimized random forest model; the training module is used to train the optimized random forest model with training sets respectively; the target ground motion identification module is used to input the ground motion to be determined into the trained optimized random forest model in sequence, sort the obtained ground motion demand parameters DM, and the ground motion corresponding to the maximum ground motion demand parameter DM is the target ground motion.
[0016] In summary, compared with the prior art, the method and system for selecting the most unfavorable ground motion based on big data and random forest provided by this invention have the following advantages:
[0017] 1. The method proposed in this invention can obtain the seismic response of a structure under all seismic records in the dataset with minimal computation, and sequentially calculate the structural ground motion demand parameter (DM index). By sorting the DM indexes, the envelope response of the structure under massive ground motion records can be determined, fully considering the uncertainty of ground motion.
[0018] 2. This invention employs a Bayesian-optimized random forest algorithm to predict the structural seismic demand parameter (DM). Experimental results demonstrate that the proposed method exhibits better generalization performance compared to other models, accurately predicting the DM index given the IM index.
[0019] 3. This invention considers the influence of various seismic IM indices on structural DM response indices, uses correlation analysis to reduce the dimensionality of IM indices, and balances the performance of machine learning models with the uncertainty of seismic action. Attached Figure Description
[0020] Figure 1 This is a schematic diagram illustrating the steps of the method for selecting the most unfavorable ground motion based on big data and random forest in an embodiment of this application;
[0021] Figure 2 This is a flowchart illustrating the method for selecting the most unfavorable ground motion based on big data and random forest in an embodiment of this application.
[0022] Figure 3 This is a schematic diagram of the finite element model of the seismic isolation continuous beam bridge used in the embodiments of this application, wherein (1) is the overall structural model of the bridge and the arrangement of nodes, (2) is the finite element model of the seismic isolation bearing used by the bridge - the LRB bearing, and the structural hysteresis curve of the LRB bearing, and (3) is the finite element model of the bridge pier.
[0023] Figure 4 This is a schematic diagram of the finite element model of the reinforced concrete frame structure used in the embodiments of this application, wherein (1) shows the front view of the frame structure, (2) shows the top view of the frame structure, and (3) shows the cross-sectional dimensions and reinforcement of the beams and columns of the frame structure.
[0024] Figure 5 This is a schematic diagram of the training and validation of the random forest model in the embodiment of this application, wherein (1) is the random forest algorithm model, and (2) is the cross-validation performance evaluation and optimization result validator;
[0025] Figure 6 This is a scatter plot of the verification of the random forest in the embodiment of this application, wherein (1) is the inter-story drift angle of the building under 70 gal, (2) is the inter-story drift angle of the building under 400 gal, (3) is the maximum support displacement of the bridge under 70 gal, (4) is the maximum support displacement of the bridge under 400 gal, (5) is the pier bottom curvature of the bridge under 70 gal, and (6) is the pier bottom curvature of the bridge under 400 gal.
[0026] Figure 7This is a distribution diagram of the predicted results of the structural DM index in the embodiments of this application, wherein (1) is the shelf displacement angle of the building structure under 70 gal, (2) is the shelf displacement angle of the building structure under 400 gal, (3) is the support displacement of bridge node 58 under 70 gal, (4) is the support displacement of bridge node 58 under 400 gal, (5) is the pier bottom curvature of bridge node 58 under 70 gal, and (6) is the pier bottom curvature of bridge node 58 under 400 gal. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0028] To achieve rapid and accurate selection of the most unfavorable design ground motion for a structure and address the shortcomings of existing methods that do not fully consider the uncertainty of ground motion, this invention provides a method for selecting the most unfavorable ground motion based on big data and random forests, using the Stanford Earthquake Dataset (STEAD) and two specific structures. Figure 1 and Figure 2 As shown, the main steps are as follows: steps S1 to S7.
[0029] S1: Based on the finite element model of the structure under test, a preset number of ground motions are used as input for nonlinear time history analysis to obtain the ground motion demand parameter DM of the structure under test. The ground motion intensity parameter IM of the ground motion is matched one-to-one with the ground motion demand parameter DM to obtain the training dataset.
[0030] For the structural type of the structure to be tested, a finite element model is built using Opensees software.
[0031] A small number of ground motions are subjected to nonlinear time history analysis. Based on the results of the nonlinear time history analysis, the response of the structure under test, i.e., the ground motion demand parameter DM, is extracted. The ground motion intensity parameter IM is extracted from the ground motions. The ground motion intensity parameter IM and the ground motion demand parameter DM are combined in a one-to-one correspondence to construct the training dataset required for subsequent processing.
[0032] In this design, nonlinear beam-column elements are used to model the material, taking into account its nonlinearity. For specific structural components, the nonlinear response process of the material is customized to fully consider the realistic representativeness of the finite element model.
[0033] The seismic motion was considered as ground acceleration time history input to the structure for calculation. The ground peak ground acceleration (PGA) of the seismic motion was adjusted to 70 gal and 400 gal, corresponding to two different waterproofing levels.
[0034] After calculating the nonlinear time history response of the structure, the seismic motion demand parameter DM is extracted. For bridge structures, the seismic motion demand parameter DM to be extracted is the maximum pier base curvature φ of the pier. max With the maximum support displacement δ max For building structures, the seismic demand parameter DM that needs to be extracted is the maximum inter-story drift angle θ of the structure. max .
[0035] In a further optimized scheme, correlation analysis and determination of seismic intensity measures (IM) indices are performed:
[0036] Considering the strong correlation between different structure types and the ground motion intensity parameter IM, excessive collinear ground motion intensity parameter IM may affect the performance of the random forest model. Therefore, it is necessary to reduce the dimensionality of the ground motion intensity parameter IM through correlation analysis.
[0037] In a further optimized scheme, the correlation analysis uses the Pearson correlation coefficient, calculated as follows:
[0038]
[0039] Where, ρ x,y Let x be the correlation coefficient between two IM parameters x and y, and Cov(·) be the covariance. x Let σ be the mean of the IM parameter x in the dataset. y Let y be the mean of the IM parameter in the dataset.
[0040] In this embodiment, modeling and verification were performed on a reinforced concrete seismic isolation continuous beam bridge and a six-story RC frame structure using numerical simulation methods. The models were built using OpenSees, and the finite element diagrams of the bridge structure and the building structure are attached. Figure 3 Appendix Figure 4 .
[0041] For bridge models, attached Figure 3 Figure (1) shows the overall structural model of the bridge and the arrangement of its nodes. This finite element model is a simplified version of the actual bridge structure. Figure 3 (2) shows the finite element model of the LRB bearing used in the bridge and the structural hysteresis curve of the LRB bearing; Appendix Figure 3 (3) shows the finite element model of the bridge pier.
[0042] For the structural model, attached Figure 4 Image (1) shows the main view of the frame structure, with appendix... Figure 4 Image (2) shows a top view of the frame structure, with appendix... Figure 4 Figure (3) shows the cross-sectional dimensions and reinforcement details of the beams and columns of the frame structure. The beams and columns of this frame are all reinforced concrete structures with a concrete grade of C35 and a steel reinforcement grade of HRB400. The concrete constitutive model uses Concrete02 and is simulated using Opensees' nonlinear beam-column elements.
[0043] Considering the applicability of the method to different structural types, models were created for both building and bridge structures. Nonlinear dynamic finite element analysis was used to calculate the seismic response of a small number of structures and extract the ground motion demand parameter DM.
[0044] Considering structural nonlinearity, a self-developed nonlinear constitutive calculation program in OpenSee was used for analysis. Seismic action was considered as ground acceleration time history input to the structure for calculation. The peak ground acceleration (PGA) was adjusted to 70 gal and 400 gal, corresponding to two different flood control levels for minor and major earthquakes, respectively. After calculating the structural nonlinear time history response, the structural DM index was extracted. For bridge structures, the parameter to be extracted is the maximum pier base curvature φ. max With the maximum support displacement δ Max There are two bridge structures; for building structures, the parameter that needs to be extracted is the maximum inter-story drift angle θ. max .
[0045] The IM and DM indices of the structure are calculated. Using the IM index as input and the DM index as label, a training dataset for the random forest model is constructed. It should be noted that the number of seismic waves initially selected for the nonlinear time history calculation in constructing the training dataset is not a fixed value. Based on experiments and conservative estimates, this invention considers extracting 500 seismic ground motions for calculation to be reasonable.
[0046] Correlation analysis and determination of seismic intensity measures (IM) indices: First, 25 IM indices frequently discussed and used in engineering were extracted, and the selected IM indices are shown in Table 1. After correlation analysis, 10 IM indices were selected, namely PGD, RMSA, and I... A ,VSI,SMA,SMV,I V ,S V (T1),T90,FAS.
[0047]
[0048] Table 1
[0049] S2: Normalize the training dataset to obtain an initial dataset, and divide the initial dataset into a training set and a test set.
[0050] In a further optimized scheme, after obtaining the ground motion intensity parameter IM and ground motion demand parameter DM data, considering that the ground motion intensity parameter IM and ground motion demand parameter DM are usually of a very small order of magnitude, in order to improve the generalization performance of the model, a logarithmic operation is performed on the ground motion intensity parameter IM and ground motion demand parameter DM.
[0051] To eliminate the dimensional differences between the seismic intensity parameter IM and the seismic demand parameter DM, a maximum-minimum normalization (MPN) method is used to normalize the data. The MPN formula is as follows:
[0052]
[0053] Where, m i Let be the i-th sample value of feature m, min(m) be the minimum value of feature m, and max(m) be the maximum value of feature m. These are the features after min-max normalization.
[0054] In a further optimized approach, the initial dataset is divided into a supervised learning training set and a test set.
[0055] In this embodiment, the IM and DM indices have the following characteristics: under the International System of Units (SI), their orders of magnitude are typically in the range of 10. -2 The following points should be noted: The difference in magnitude between IM and DM is very large. The structural response calculated by different seismic waves with the same PGA may differ by more than 3-5 orders of magnitude. In engineering, it is believed that there is a linear logarithmic correlation between IM and DM. Therefore, logarithmic operations are performed on both IM and DM indices to eliminate the impact of the large difference in magnitude on model performance.
[0056] A structure typically has different DM (damming) indices, such as the maximum pier base curvature and maximum support displacement of a bridge structure. To facilitate comparison of seismic failure potentials using DM indices, the calculated structural DM indices are used to calculate the following formula:
[0057]
[0058] To standardize the model's input, min-max normalization is used. value.
[0059] In this embodiment, the machine learning network model upon which the structural DM index prediction model is based during the training and optimization phases mainly includes: the random forest algorithm model. Figure 5(1) Bayesian optimization hyperparameter generator, cross-validation performance evaluator, and optimization result validator are attached. Figure 5 In (2), the steps of the prediction model in the training and optimization phases are as follows.
[0060] S3: Divide the training set into multiple parts, and use each part as the validation set in turn, while using the rest as the training set for cross-validation. Take the maximum value of the average correlation coefficient between the predicted value DM and the true value DM obtained from each cross-validation as the objective function of the Bayesian optimization algorithm.
[0061] This application uses a random forest model to screen earthquake motions. The first step is to optimize the hyperparameters of the random forest. In a further preferred scheme, the Bayesian optimization algorithm is used to optimize the hyperparameters of the random forest.
[0062] Specifically, the training set is divided into multiple parts, and each part is used as the validation set in turn, while the rest are used as the training set for cross-validation. The cross-validation method can be k-fold cross-validation; more specifically, in this embodiment, 5-fold cross-validation is used. The correlation coefficient (R0) between the predicted value DM and the true value DM obtained from each cross-validation is calculated. 2 The average value is used as the objective function f(x) of the Bayesian optimization algorithm.
[0063] S4: Construct the acquisition function of the Bayesian optimization algorithm and construct the surrogate model of the Bayesian optimization algorithm based on the objective function.
[0064] The acquisition function of the Bayesian optimization algorithm is constructed using the upper confidence bound (UCB) algorithm.
[0065] Based on the above objective function, a surrogate model for the Bayesian optimization algorithm is constructed.
[0066] S5: Multiple sets of random forest hyperparameters are preset. The acquisition function collects the random forest hyperparameters as input to the surrogate model for iterative optimization to obtain the optimal random forest hyperparameters, thereby obtaining the optimized random forest model.
[0067] This application pre-sets multiple sets of random forest hyperparameters as the collection objects of the collection function. Each time, the collection function collects a set of random forest hyperparameters, inputs them into the surrogate model for calculation, and updates the posterior distribution of the surrogate model. The collection function then collects the next set of random forest hyperparameters and updates the posterior distribution of the surrogate model. This process is repeated until the optimal random forest hyperparameters are finally determined.
[0068] In this embodiment, four hyperparameters are selected for optimization and the optimization range is determined. These four hyperparameters are the number of base learners, the minimum classification tree, the maximum number of features, and the maximum tree depth.
[0069] S6: Train the optimized random forest model using the training set respectively.
[0070] The optimal hyperparameters of the random forest are used to learn from the training set. The random forest continuously splits and generates multiple regression trees by partitioning the feature variables, transforming multiple weak regressors into strong regressors. The random forest is an ensemble learning model composed of multiple base models, and its predicted value can be expressed as the sum of the results of multiple base models, as follows:
[0071]
[0072] Where N is the hypothesis space of the regression trees, K is the number of regression trees in the model, and q i For the i-th sample, input features, f k (q i ) represents the prediction result for the k-th tree.
[0073] Step S6 further includes testing the trained optimized random forest model using a test set to obtain test ground motion demand parameters DM. The test ground motion demand parameters DM are then compared with the ground motion demand parameters DM in the test set. If the error is greater than a preset value, the random forest model is improved. Further, the root mean square error or coefficient of determination is used to compare the test ground motion demand parameters DM with the ground motion demand parameters DM in the test set.
[0074] In this embodiment, the evaluation metrics include root mean square error (RMSE) and coefficient of determination (R²). 2 The two evaluation indicators can be defined as follows:
[0075]
[0076]
[0077] Where, N samples y represents the total number of samples in the training set. i and These represent the actual value and the predicted value for each sample, respectively. R represents the mean of the total sample. A lower RMSE score indicates a better predictive performance of the model. 2 The value range is [0,1]. The closer the index is to 1, the better the performance of the prediction model.
[0078] After training and validation of the model, a scatter plot of the model's predictions was plotted on the test set, as shown in the attached figure. Figure 6 As shown, the model's prediction performance is good; it can predict the DM index of the structure within 5% error for most data samples, and the accuracy meets the requirements.
[0079] S7: Input the undetermined ground motions into the trained optimized random forest model in sequence, sort the obtained ground motion demand parameters DM, and the ground motion corresponding to the maximum ground motion demand parameter DM is the target ground motion.
[0080] For a given structure, a trained random forest model is used to predict the structure's response in a large seismic motion dataset. The structure is then ranked according to the predicted seismic motion demand parameter DM. Seismic motions with larger DM indices have greater destructive potential, which means they are detrimental to the structure.
[0081] Based on the optimized random forest model, the DM (Displacement Degree) indices of structures under different seismic waves were predicted. In an exemplary embodiment, predictions were made for a seismically isolated continuous beam bridge structure and a six-story frame building structure. For the bridge structure, the maximum curvature of the piers and the maximum displacement of the supports were predicted; for the building structure, the maximum inter-story drift angle was predicted. In STEAD, the structural DM indices under all seismic waves were predicted. The distribution diagram of the predicted indices in the embodiment is shown below. Figure 7 As shown.
[0082] Seismic Failure Potential Ranking: Based on the predicted structural DM indices, seismic failure potentials are ranked. The ranking method involves comparing the results from step S32. The value, if a certain seismic wave is calculated The larger the value, the greater the destructive potential of the seismic wave on the structure. In the embodiment, the two structures are compared using all seismic waves from the STEAD database. The index names of the ten most destructive seismic waves, as calculated, are listed in Table 2.
[0083]
[0084] Table 2
[0085] A second aspect of this application provides a system for selecting the most unfavorable ground motion based on big data and random forest, the system comprising:
[0086] Analysis module: It is used to perform nonlinear time history analysis based on the finite element model of the structure under test, taking a preset number of ground motions as input, to obtain the ground motion demand parameter DM of the structure under test, and to obtain a training dataset by matching the ground motion intensity parameter IM of the ground motion with the ground motion demand parameter DM.
[0087] Normalization module: used to normalize the training dataset to obtain an initial dataset, and to divide the initial dataset into a training set and a test set;
[0088] Objective function acquisition module: used to divide the training set into multiple parts, take each part as the validation set in turn, and use the rest as the training set for cross-validation. The average value of the correlation coefficient between the predicted value DM and the true value DM obtained from each cross-validation is taken as the objective function of the Bayesian optimization algorithm.
[0089] Proxy model acquisition module: used to construct the acquisition function of the Bayesian optimization algorithm and construct the proxy model of the Bayesian optimization algorithm based on the objective function;
[0090] Iterative optimization module: used to preset multiple sets of random forest hyperparameters, the acquisition function collects the random forest hyperparameters as input to the surrogate model for iterative optimization, obtains the optimal random forest hyperparameters, and then obtains the optimized random forest model;
[0091] Training module: used to train the optimized random forest model using the training set respectively;
[0092] Target ground motion identification module: It is used to input the ground motions to be determined into the trained optimized random forest model in sequence, sort the obtained ground motion demand parameters DM, and the ground motion corresponding to the maximum ground motion demand parameter DM is the target ground motion.
[0093] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for selecting the most unfavorable seismic ground motion based on big data and random forest, characterized in that, The methods include: S1: Based on the finite element model of the structure under test, a preset number of ground motions are used as input for nonlinear time history analysis to obtain the ground motion demand parameter DM of the structure under test. The ground motion intensity parameter IM of the ground motion is matched one-to-one with the ground motion demand parameter DM to obtain the training dataset. S2: Normalize the training dataset to obtain an initial dataset, and divide the initial dataset into a training set and a test set; S3: Divide the training set into multiple parts, and use each part as the validation set in turn, while the rest are used as the training set for cross-validation. The maximum value of the average correlation coefficient between the predicted value DM and the true value DM obtained from each cross-validation is used as the objective function of the Bayesian optimization algorithm. S4: Construct the acquisition function of the Bayesian optimization algorithm and construct the surrogate model of the Bayesian optimization algorithm based on the objective function; S5: Multiple sets of random forest hyperparameters are preset. The acquisition function collects the random forest hyperparameters as input to the surrogate model for iterative optimization to obtain the optimal random forest hyperparameters, and then obtains the optimized random forest model. S6: Train the optimized random forest model using the training set; S7: Input the undetermined ground motions into the trained optimized random forest model in sequence, sort the obtained ground motion demand parameters DM, and the ground motion corresponding to the maximum ground motion demand parameter DM is the target ground motion.
2. The method for selecting the most unfavorable seismic motion according to claim 1, characterized in that, Step S1 also includes dimensionality reduction of the ground motion intensity parameter IM.
3. The method for selecting the most unfavorable seismic motion according to claim 1, characterized in that, Pearson correlation coefficient was used for correlation analysis to achieve dimensionality reduction of the seismic intensity parameter IM. The formula for calculating the Pearson correlation coefficient is as follows: Where, ρ x,y Let x be the correlation coefficient between two IM parameters x and y, and Cov(·) be the covariance. x Let σ be the mean of the IM parameter x in the dataset. y This represents the mean of the IM parameter y in the dataset.
4. The method for selecting the most unfavorable seismic motion according to any one of claims 1 to 3, characterized in that, Step S1 also includes normalizing the ground motion intensity parameter IM and the ground motion demand parameter DM.
5. The method for selecting the most unfavorable seismic motion according to claim 1, characterized in that, Step S6 further includes testing the trained optimized random forest model using a test set to obtain the test ground motion demand parameter DM, comparing the test ground motion demand parameter DM with the ground motion demand parameter DM in the test set, and improving the random forest model if the error is greater than a preset value.
6. The method for selecting the most unfavorable seismic motion according to claim 5, characterized in that, The root mean square error or coefficient of determination is used to compare the test ground motion demand parameter DM with the ground motion demand parameter DM in the test set.
7. A system for selecting the most unfavorable seismic ground motion based on big data and random forest, characterized in that, The system includes: Analysis module: It is used to perform nonlinear time history analysis based on the finite element model of the structure under test, taking a preset number of ground motions as input, to obtain the ground motion demand parameter DM of the structure under test, and to obtain a training dataset by matching the ground motion intensity parameter IM of the ground motion with the ground motion demand parameter DM. Normalization module: used to normalize the training dataset to obtain an initial dataset, and to divide the initial dataset into a training set and a test set; Objective function acquisition module: used to divide the training set into multiple parts, take each part as the validation set in turn, and use the rest as the training set for cross-validation. The average value of the correlation coefficient between the predicted value DM and the true value DM obtained from each cross-validation is taken as the objective function of the Bayesian optimization algorithm. Proxy model acquisition module: used to construct the acquisition function of the Bayesian optimization algorithm and construct the proxy model of the Bayesian optimization algorithm based on the objective function; Iterative optimization module: used to preset multiple sets of random forest hyperparameters, the acquisition function collects the random forest hyperparameters as input to the surrogate model for iterative optimization, obtains the optimal random forest hyperparameters, and then obtains the optimized random forest model; Training module: used to train the optimized random forest model using the training set respectively; Target ground motion identification module: It is used to input the ground motions to be determined into the trained optimized random forest model in sequence, sort the obtained ground motion demand parameters DM, and the ground motion corresponding to the maximum ground motion demand parameter DM is the target ground motion.
Citation Information
Patent Citations
Pile-soil-structure system peak earthquake response prediction method and system and medium
CN113128081A
Method for predicting earthquake vulnerability of nuclear power plant based on machine learning and MSA
CN116305434A