A method and device for finding a model of ground motion attenuation relationship based on symbol learning
Through the geoscillation attenuation relationship finding method based on symbol learning, the expression form of the earthquake prediction equation is simplified by using STRidge algorithm and sparse regression technology, the calculation efficiency and applicability are improved, the problems of complex expression and inefficiency in the existing technology are solved, and the concise characterization and physical relationship recognition of earthquake parameters are realized.
Patent Information
- Application Number
- CN202310256030.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-16
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-16
AI Technical Summary
In the prior art, the expression forms of earthquake prediction equations are complex, and the computational efficiency of machine learning methods is inefficient, especially when applied to new regions, the model parameters need to be retrained.
The geoscillation attenuation relationship finding method based on symbol learning is adopted, and the function expression P=Θξ is constructed, and the optimal coefficient is solved using the STRidge algorithm, combining sparse regression and regularization technology to simplify the expression method and improve the calculation efficiency.
The concise and interpretable characterization of complex earthquake systems is realized, the applicability and computational efficiency of the model in different regions is improved, the implicit physical relationships of earthquake variables can be identified, and a new regression paradigm that takes into account both physics and data-driven.
Smart Images

Figure CN116299703B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of earthquake engineering, and in particular to a method and device for finding a ground motion attenuation relationship based on symbolic learning. Background Art
[0002] Ground motion prediction equations (GMPEs) are an important part of seismic hazard analysis and are also an important basis for earthquake zoning and earthquake disaster prevention. Their core idea is to predict the peak ground acceleration (PGA) and peak ground velocity (PGV) based on actual ground motion record data or data generated by ground motion simulation by statistically analyzing terms such as magnitude, distance, hanging wall effect, fault type, linear and nonlinear site terms, basin effect, and regional differences.
[0003] The establishment of traditional GMPEs requires combining physical parameters given by seismology, geophysics, etc., predefining the form of the model equation, and obtaining GMPEs through statistical regression methods. However, both the data selection criteria and the selected model form will have a greater impact on GMPEs. The accuracy of this regression equation mainly depends on the regression method, data scale, and quality, and there is a certain degree of uncertainty. With the booming development of machine learning and deep learning methods, some scholars have used machine learning and deep learning methods (such as neural networks, genetic algorithms, decision trees) to establish GMPEs. The above methods have the ability of adaptive learning, can extract the hidden features of data in pattern recognition, and have demonstrated the prediction ability of GMPEs by dividing the test set.
[0004] However, the equation forms given by these methods are relatively complex. Especially when these models are reapplied in a new region, the model parameters need to be retrained, which is computationally inefficient for machine learning methods. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and device for finding a ground motion attenuation relationship based on symbolic learning, aiming to solve the problems of complex expression forms and low computational efficiency of machine learning methods in the prior art.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] On the one hand, a method for finding a ground motion attenuation relationship based on symbolic learning is provided, including a model construction step.
[0008] The functional expression of the model is shown in Equation (1):
[0009] P = Θξ (1)
[0010] In Equation (1), P represents the ground motion parameter matrix, which is defined as shown in Equation (2):
[0011]
[0012] Among them, p1, p2, p3, …, p n represent n groups of ground motion parameters, and n represents the number of groups of ground motion parameters; the ground motion parameters include peak ground acceleration PGA and peak ground velocity PGV; T represents matrix transpose; represents the set of real numbers;
[0013] In formula (1), Θ represents a candidate function library, which is defined as shown in formula (3):
[0014]
[0015] Among them, M w represents magnitude; R JB represents the closest distance (km) from the fault projection to the receiving station on the ground; V S30 represents the equivalent shear wave velocity at 30 m depth of the site (m / s); FM represents the source mechanism; Z 1.0 represents the formation depth corresponding to the shear wave velocity of 1.0 km / s from the ground; n represents the number of groups of candidate function library parameters; m represents the number of m candidate functions, including: constant term, first-order term, ln term, and square term;
[0016] In formula (1), ξ represents a sparse coefficient matrix, which is defined as shown in formula (4):
[0017]
[0018] Among them, ξ1, ξ2, ξ3, …, ξ m represent m groups of sparse coefficients; m represents the number of groups of sparse coefficients; T represents matrix transpose; represents the set of real numbers.
[0019] On the other hand, a ground motion attenuation relationship model finding device based on symbolic learning is provided, including at least one processor; and a memory, which stores data and instructions, and when the instructions are executed by at least one processor, the steps of the foregoing method are implemented.
[0020] The beneficial effects of the present invention are as follows: a novel symbolic learning method embedded with physical information is proposed, which realizes the automatic discovery of mathematical equation operators as symbols, obtains a concise and interpretable explicit representation of the complex ground motion system, greatly simplifies the expression method, and effectively improves the applicability of use; and it can be used to identify the implicit physical relationships of ground motion variables in different regions, give prediction equations, provide a new regression paradigm that combines physics and data-driven machine learning, and effectively improves the computational efficiency of machine learning methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic diagram of the device in the present invention;
[0022] Figure 2 is a schematic diagram of the platform site conditions in the present invention;
[0023] Figure 3 is a schematic diagram of the distribution of PGA and PGV with distance and magnitude in the present invention;
[0024] Figure 4 is a schematic diagram of the PISL - GMM framework in the present invention;
[0025] Figure 5 is a schematic diagram of the relationship between the number of items in the candidate function library and the initial threshold in the present invention;
[0026] Figure 6 is a schematic diagram of the comparison of prediction results in the present invention;
[0027] Figure 7 is the residual within an event and M in the present invention w , R JB and V S30 relationship distribution diagram;
[0028] Figure 8 is the distribution diagram of the residual between events relative to M w in the present invention;
[0029] Figure 9 is a comparison schematic diagram of the standard deviation τ of the residual between events and the standard deviation φ of the residual within an event of the three models of PISL - GMPE, BSSA14 and ANN in the present invention;
[0030] Figure 10 is a schematic diagram of the extrapolation ability of the test model in the present invention. Detailed implementation manners
[0031] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the drawings and implementation manners of the present invention.
[0032] Some implementation manners of the present invention relate to a device for realizing the type - finding of ground motion attenuation relationship based on symbolic learning, as Figure 1 shown, the type - finding device 1 includes at least one processor 2; and a memory 3, which stores data and instructions for implementing all steps in the following method implementation manners.
[0033] In some implementation manners of the method for finding the type of ground motion attenuation relationship based on symbolic learning, it includes a model construction step,
[0034] The functional expression of the model is as shown in Equation (1):
[0035] P = Θξ (1)
[0036] In Equation (1), P represents the ground motion parameter matrix, which is defined as shown in Equation (2):
[0037]
[0038] where p1, p2, p3, …, p n represent n groups of ground motion parameters, and n represents the number of groups of ground motion parameters; the ground motion parameters include peak ground acceleration PGA and peak ground velocity PGV; T represents matrix transpose; represents the set of real numbers;
[0039] In Equation (1), Θ represents the candidate function library, which is defined as shown in Equation (3):
[0040]
[0041] where M w represents the magnitude; R JB represents the closest distance (km) from the fault projection to the receiving station on the ground; V S30 represents the equivalent shear wave velocity at 30 m depth of the site (m / s); FM represents the source mechanism; Z 1.0 represents the formation depth corresponding to a shear wave velocity of 1.0 km / s from the ground; n represents the number of groups of candidate function library parameters; m represents the number of m candidate functions, including constant terms, linear terms, ln terms, and quadratic terms;
[0042] In Equation (1), ξ represents the sparse coefficient matrix, which is defined as shown in Equation (4):
[0043]
[0044] where ξ1, ξ2, ξ3, …, ξ m represent m groups of sparse coefficients; m represents the number of groups of sparse coefficients; T represents matrix transpose; represents the set of real numbers.
[0045] In some implementation manners of the ground motion attenuation relationship model finding method based on symbolic learning, it further includes the step of solving the optimal coefficients step,
[0046] Solving the optimal coefficients The algorithm is as shown in Equation (5):
[0047]
[0048] where Θ represents the candidate function library, ξ represents the sparse coefficient matrix; P represents the ground motion parameter matrix; λ represents the regularization coefficient; T represents matrix transpose; I represents the identity matrix.
[0049] In some implementation manners of the method for finding the attenuation relationship of ground motion based on symbol learning, it further includes preliminarily estimating the optimal solution ξ best and the residual error best step
[0050] Set the initial thresholds δ and l 0 , where l 0 represents the 0-order norm, set the norm regularization multiplication factor η, and preliminarily estimate the optimal solution ξ using Equation (6-1) best ;
[0051] Calculate the residual error of the validation set under the optimal solution ξ best , and its functional expression is as shown in Equation (6-2): best
[0052] ξ best =(Θ train ) -1 P train (6-1)
[0053]
[0054] where Θ train represents the training set of the candidate function library; P train represents the training set of the ground motion parameter matrix; Θ validation represents the validation set of the candidate function library; P validation represents the validation set of the ground motion parameter matrix; η represents the norm regularization multiplication factor; preferably, η = 10 -3 κ(Θ), where κ(Θ) is the condition number of the matrix Θ
[0055] In some implementation manners of the method for finding the attenuation relationship of ground motion based on symbol learning, it further includes an iterative optimization step
[0056] Set the number of iterations and the regularization coefficient λ;
[0057] Iterate the algorithm described in Equation (5);
[0058] During the iteration, increase the initial threshold δ to continuously reduce the residual error of the validation set best until the optimal solution is given
[0059] In some implementation manners of the method for finding the attenuation relationship of ground motion based on symbol learning, it further includes a data normalization step
[0060] The functional expression of the data normalization is as shown in Equation (7):
[0061]
[0062] Among them, x represents the candidate function library parameters and ground motion parameters respectively;
[0063] The candidate function library parameters include M w , R JB , V S30 , FM and Z 1.0 ; The ground motion parameters include PGA and PGV.
[0064] In some implementation manners of the ground motion attenuation relationship model finding method based on symbolic learning, it further includes a step of selecting seismic events.
[0065] Select seismic events from the seismic event database. The seismic events include ground motion parameters and their corresponding candidate function library parameters; the ground motion parameters construct the ground motion parameter matrix P, and the candidate function library parameters construct the candidate function library Θ.
[0066] In some implementation manners of the ground motion attenuation relationship model finding method based on symbolic learning, it further includes a step of processing seismic event data.
[0067] Delete the events that do not list any one of the candidate function library parameters. The candidate function library parameters include M w , FM, R JB , V S30 and Z 1.0 ;
[0068] Delete the events with less than 5 seismic records, and delete the events with focal depths less than 1 km and greater than 20 km.
[0069] In some implementation manners of the ground motion attenuation relationship model finding method based on symbolic learning, the candidate function library parameters of the candidate function library Θ further include the following parameters:
[0070] ln(R JB + 10), representing the distance saturation term, and setting the saturation factor to 10 km;
[0071] M w ln(R JB + 10), representing the magnitude saturation term, and setting the saturation factor to 10 km.
[0072] In some implementation manners of the ground motion attenuation relationship model finding method based on symbolic learning, it further includes a step of obtaining the ground motion parameter equation.
[0073] The functional expression of the ground motion parameter equation is shown in Equation (8):
[0074] P' = ξ ' best ·Θ' (8)
[0075] Among them, P' represents the obtained ground motion parameters, and the ground motion parameters include peak ground acceleration PGA and peak ground velocity PGV; ξ ' best represents the optimal solution of the obtained coefficient matrix; Θ' represents the obtained candidate function library.
[0076] Next, the technical solution of the present invention will be clearly and completely described by taking the data of the NGA-West2 project and combining with the drawings of the present invention.
[0077] First, select the data from the NGA-West2 project, and then delete those seismic events that do not list the magnitude (M w ), focal mechanism (FM), Joyner-Boore distance (R JB ), V S30 or the formation depth (Z 1.0 ) corresponding to the shear wave velocity of 1.0 km / s from the ground for any one of the parameters. Delete the seismic events with less than 5 seismic records. Delete the seismic events with the focal depth less than 1 km and greater than 20 km to ensure that the focal source is located in the shallow crust. Finally, 15,175 records corresponding to 282 shallow crust earthquakes recorded by 2,655 seismic stations including those in California, USA and the entire Japanese island are determined. These stations have different site conditions, such as Figure 2 shown. The specific parameter explanations and selection ranges are shown in Table 1. The selected parameters mainly include: M w , R JB , the equivalent shear wave velocity V S30 of the site at 30 m, FM, and the depth Z 1.0 corresponding to the shear wave velocity reaching 1 km / s. Figure 3 shows the distributions of PGA and PGV with distance and magnitude. Saturation phenomena related to magnitude and distance are shown in both PGA and PGV.
[0078] Table 1 Definitions and ranges of selected parameters
[0079]
[0080] For n groups of ground motion parameter definitions respectively represent the ground motion parameters PGA and PGV, are n groups of parameters M w , R JB , V S30 , FM, Z 1.0Construct a candidate function library Θ of m functions that include constant terms, linear terms, ln terms, and quadratic terms. The constant term is the residual of the model, representing the deviation of the model from the observed data. The linear terms, ln terms, and quadratic terms represent the non-linear attenuation of the model, and these terms are adopted in many GMPE models.
[0081] Specifically, to consider the distance saturation effect in the near field of large earthquakes, saturation terms related to magnitude and distance are added to the function library: ln(R JB +10) and M w ln(R JB +10), as shown in Figure 4 d. This function library is based on data fitting in the case of small and medium-sized earthquakes with rich data. In the case of sparse data in the near field of large earthquakes, prior physical knowledge is fully integrated to ensure the completeness and interpretability of the function library. In addition, since the influence of site amplification on ground motion cannot be ignored, the description of non-linear site amplification based on the BSSA14 model is referred to and simplified, and it is obtained that 1500 m / s is both the limiting velocity of V S30 , and also the limiting velocity of the reference bedrock site. Therefore, before normalizing the data, divide V S30 by 1500 m / s.
[0082] is the sparse coefficient matrix, which determines the active terms in Θ. Therefore, the function expression based on sparse regression is as follows:
[0083] P = Θξ (1)
[0084] All the function forms characterizing P are completely included in Θ; and the corresponding coefficient matrix ξ is sparse, indicating that P can be represented by some terms in Θ. Therefore, the algorithm for solving the optimal coefficient is shown in formula (5):
[0085]
[0086] In principle, STRidge (Sequential threshold ridge regression) is a combination of sequential threshold least squares regression and ridge regression, which can give a sparse STRidge provides clear interpretability in terms of managing equations and the original variable set. The present invention extends this method to the solution of ground motion models, forming the PISL-GMM algorithm, and its framework is schematically shown in Figure 4 .
[0087] Figure 4In the PISL-GMM framework shown, the arrows in the figure indicate the order of constructing the model and solving. (a) The data of NGA-West2 was collected. (b) A typical STRidge was constructed, and the short columns represent the parameters Ξ to be solved. (c) Using the STRidge algorithm, PGA and PGV were solved separately. (d) The calculated candidates and their coefficients, where ξ1 and ξ2 represent the coefficients corresponding to ln(PGA) and ln(PGV) respectively. (e) Some descriptions of this method, including the calculation of data normalization, the assumptions of the model, and a schematic diagram showing the selection of the initial threshold.
[0088] Taking PGA as an example, the specific calculation process will be introduced below.
[0089] 1) To avoid the inconsistent influence scale on PGA prediction caused by the parameters M w , FM, R JB , V S30 and Z 1.0 being in different magnitudes, the data was first normalized, as shown in Equation (7):
[0090]
[0091] where x represents M w , FM, R JB , V S30 , Z 1.0 and ln(PGA) respectively.
[0092] 2) PGA and its corresponding M w , FM, R JB , V S30 , Z 1.0 and other parameters were randomly divided into a training set, a validation set, and a test set, with their proportions being 60%, 20%, and 20% respectively:
[0093] 3) Set the initial threshold δ and the 0-order norm regularization multiplication factor η, where η is empirically set to 10 -3 κ(Θ), and κ(Θ) is the condition number of the matrix Θ.
[0094] 4) Use the training set to initially estimate ξ best and calculate the residual error best error best of the validation set under ξ
[0095] ξ best = (Θ train ) -1 P train (6-1)
[0096]
[0097] 5) Start the iterative algorithm 1 as shown in Table 2. Set the number of iterations to 500 and the regularization coefficient λ to 1e-7. During the iteration process, increase the initial threshold δ to continuously reduce the error of the validation set until the optimal ξ is given best is obtained best .
[0098] Table 2
[0099]
[0100] Since the value of the initial threshold δ of the coefficient matrix has a great influence on the result, during the model training process, after determining the regularization coefficient λ, different δ values are used to observe the accuracy of the model, and the relationship between the number of non-zero items in the function library and δ is plotted as Figure 5 shown Figure 5 The thresholds δ in the PGA and PGV prediction equations are given. It can be seen that as δ increases, the sparsity of the function library increases (the number of items decreases). The present invention selects the δ corresponding to the point where the number of items drops sharply as the optimal threshold. This not only ensures sparsity but also avoids the underfitting phenomenon of the model caused by too large δ
[0101] The PGA and PGV formulas predicted by PISL-GMPEs are shown in Formulas 9 and 10 respectively (see Figure 4 d). PISL-GMM not only provides a concise expression but also considers features such as near-field saturation and inelastic attenuation. The present invention compares PISL-GMM with the leading empirical GMM (BSSA14) and the data-driven ANN-based GMM (ANN). By evaluating the residuals and inference capabilities of the models, the feasibility of PISL-GMM is clarified
[0102]
[0103] Verification of the rationality of PISL-GMPE:
[0104] Compare the PGAs and PGVs predicted based on the NGA-West2 database with the prediction results of different M w , R JB and V S30 obtained by BSSA14 and ANN, as shown Figure 6 in
[0105] At V S30 = 560 and 760 m / s ( Figure 6 a), the PGAs predicted by the three models are similar in distance and amplitude. When V S30 = 200 m / s, at R JB > 30 km, the results of the three models are similar. However, at R JBWhen R < 30 km, the values predicted by PISL-GMPE are very high, which may be due to site nonlinearity.
[0106] At R JB < 30 km ( Figure 6 b), considering the prior knowledge (distance saturation), the results obtained by PISL-GMPE are consistent with those obtained by BSSA14; however, ANN does not reflect this (especially for the PGV predictions for M w = 4 and 5, see Figure 6 b). Due to the purely data-driven nature of ANN, they may be affected by fitting biases when there is no data or the data distribution is uneven. At R JB > 200 km, there is an obvious upward bias in the PGV predicted by PISL-GMPE, but the predictions of BSSA14 and ANN do not. This is similar to seismological observations. At R JB > 200 km, the energy of the Sb phase (upper crust diffracted S wave) is negligible, and the Sg phase (direct S wave) and Sn wave (Mohorovicic discontinuity diffracted S wave) in the 0.5 - 5 s period band become the main energy sources of seismic S waves, which leads to the upward curvature of the PGV attenuation curve.
[0107] Residual analysis of PISL-GMPE:
[0108] Residual analysis is important for evaluating the dispersion of the values predicted by a regression model. The present invention refers to the random effects model developed by Abrahamson and Youngs (1992) for residual analysis. The within-event and between-event residuals of the predicted ln(PGA) and ln(PGV) for the M w , R JB and V S30 distributions with respect to the measured values are calculated, as shown in Figure 7 and Figure 8 . Overall, the within-event residual distribution of ln(PGA) is between -4 and 4, and the residual values of all three variables are located near the residual mean value of 0. This means that the PGA and PGV estimated by PISL-GMM are unbiased with respect to M w , R JB and V S30 . Only when M w > 7, the residuals will decrease slightly ( Figure 7 f).
[0109] As shown in Figure 8 , the between-event residuals of PGA and PGV with respect to M w are shown. The residuals are mainly distributed between -1.5 and 1.5, and the mean value of the residuals is around zero. This indicates that the model does not show a systematic bias. Similarly, when M w> 7 o'clock, the residual increases ( Figure 8 ), indicating that the values predicted by PISL-GMPE do not have sufficient correlation with the measured values. This may be due to the scarcity of data for large earthquakes. The obvious magnitude correlation observed in the inter-event residuals may not represent the actual probability change at large magnitudes.
[0110] To evaluate the differences among PISL-GMPE, BBSA14, and ANN, the present invention calculated the standard deviations of the inter-event (τ) and intra-event (φ) residuals of the three models according to the method of Boore8.
[0111] To calculate the τ and φ values related to M w , the expressions are shown in Equations 11 and 12.
[0112]
[0113] where the values of τ1, τ2, φ1, and φ2 are shown in Table 3:
[0114] Table 3 τ1, τ2, φ1, and φ2 in PISL-GMMs
[0115]
[0116] Using Equations 11 and 12, the relationships between the standard deviation of the inter-event residual τ and the standard deviation of the intra-event residual φ and M w were calculated, and the three models, PISL-GMPE, BSSA14, and ANN, were compared, as Figure 9 shown. Figure 9 a shows the variation of τ with magnitude. Compared with the BSSA14 model, the PISL-GMPE model has a lower τ when M w < 4.5, but increases when M w > 5.5. Compared with the ANN model, the τ of PISL-GMPE is lower both when M w < 4.5 and when M w > 5.5. Figure 9 The comparison of the φ values in b shows that the values of the three models are similar, except for the increase in φ shown by ANN at PGV. PISL-GMPE shows a standard deviation of residuals of the same order of magnitude as the other models. This model captures the variations corresponding to magnitude, distance, and site, and is therefore suitable for predicting ground motion intensity. In addition, the above three models fit the data well from three different perspectives: physical statistics (BSSA14), data-driven with prior physical knowledge (PISL), and data-driven (ANN), which results in different variations among these models.
[0117] Analysis of the extrapolation ability of PISL-GMPE:
[0118] When the model is extrapolated beyond the specified range of the training data, the error becomes large. Seismic hazard analysis of major infrastructure such as nuclear power plants and gas pipelines needs to consider extreme cases. Although these cases may not be within the range of historical seismic data, they should not be ignored. In particular, data-driven GMPEs have no physical constraints, and their extrapolation capabilities must be tested.
[0119] To test the extrapolation capabilities of the PISL and ANN models, the present invention re-trained them only with data where R JB ≥ 30 km. Subsequently, the present invention observed the prediction behavior of PISL and ANN at different magnitudes for R JB < 30 km, as Figure 10 shown. According to the results, the PISL and ANN models had similar prediction behavior at M w = 6 and 7, but the ANN had an obvious linear trend at M w = 4 and 5. The ANN model is completely data-driven. The distribution of seismic data shows that the proportion of small earthquake data is large (about 71% of the total data when M w < 5), and the proportion of large earthquake data is small (about 29% of the total data when M w ≥ 5). Therefore, the inference of the ANN will result in an unrealistic data distribution. In contrast, PISL performs very prominently in the near field because of the addition of seismic magnitude and distance saturation terms in the function library.
[0120] Embodiments and functional operations of the subject matter described in this specification can be implemented in: digital electronic circuits, tangible computer software or firmware, computer hardware, including the structures disclosed in this specification and structural equivalents thereof, or a combination of one or more of the foregoing. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on one or more tangible non-transitory program carriers for execution by, or to control the operation of, a data processing apparatus.
[0121] As an alternative or addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to an appropriate receiver device for execution by a data processing apparatus. A computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of the foregoing devices.
[0122] The term "data processing device" encompasses all kinds of devices, apparatuses, and machines for processing data, including, by way of example, programmable processors, computers, or multiprocessors or multi-computers. The device may include dedicated logic circuits, such as, for example, FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits). In addition to hardware, the device may also include code that creates an execution environment for relevant computer programs, such as code that constitutes processor firmware, protocol stacks, database management systems, operating systems, or a combination of one or more of them.
[0123] A computer program (which may also be referred to as or described as a program, software, software application, module, software module, script, or code) can be written in any form of programming language, including compiled language or interpreted language or declarative language or procedural language, and the computer program can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The computer program may or may not correspond to a file in a file system. The program may be stored in a part of a file that holds other programs or data, for example, stored in one or more of the following scripts: in a markup language document; in a single file dedicated to the relevant program; or in multiple cooperating files, such as files that store one or more modules, subroutines, or portions of code. The computer program can be deployed to execute on one computer or multiple computers, which are located in one place, or distributed to multiple locations and interconnected via a communication network.
Claims
1. A method for finding a model of ground motion attenuation relationship based on symbolic learning, characterized in that, The method includes a model construction step, The functional expression of the model is shown in Equation (1): P = Θξ (1) In Equation (1), P represents the ground motion parameter matrix, which is defined as shown in Equation (2): Among them, p1, p2, p3, …, p n represent n groups of ground motion parameters, where n represents the number of groups of ground motion parameters; the ground motion parameters include peak ground acceleration PGA and peak ground velocity PGV; T represents matrix transpose; represents the set of real numbers; In Equation (1), Θ represents the candidate function library, which is defined as shown in Equation (3): Among them, M w represents the magnitude; R JB represents the shortest distance (km) from the fault projection to the receiving station on the ground; V S30 represents the equivalent shear wave velocity at 30 m depth of the site (m / s); FM represents the source mechanism; Z 1.0 represents the formation depth corresponding to the shear wave velocity of 1.0 km / s from the ground; n represents the number of groups of candidate function library parameters; m represents the number of m candidate functions, including constant terms, first-order terms, ln terms, and quadratic terms; In Equation (1), ξ represents the sparse coefficient matrix, which is defined as shown in Equation (4): Among them, ξ1, ξ2, ξ3, …, ξ m represent m groups of sparse coefficients; m represents the number of groups of sparse coefficients; T represents matrix transpose; represents the set of real numbers.
2. The method for determining the attenuation relationship model of ground motion based on symbol learning according to claim 1, wherein The method further includes solving for the optimal coefficients step Algorithm for solving the optimal coefficients is shown in Equation (5) as follows: Wherein, Θ represents the candidate function library, ξ represents the sparse coefficient matrix; P represents the ground motion parameter matrix; λ represents the regularization coefficient; T represents the matrix transpose; I represents the identity matrix.
3. A method for finding a model of ground motion attenuation relationship based on symbol learning according to claim 2, characterized in that The method further includes a step of preliminarily estimating an optimal solution ξ best and a residual error best step Set the initial thresholds δ and l 0 , where l 0 represents the 0th norm, set the norm regularization multiplication factor η, and initially estimate the optimal solution ξ using Equation (6-1) best ; Calculate the residual error best of the validation set at the optimal solution ξ best , and its functional expression is shown in Equation (6-2): ξ best = (Θ train ) -1 P train (6 - 1) Among them, Θ train represents the candidate function library training set; P train represents the ground motion parameter matrix training set; Θ validation represents the candidate function library validation set; P validation represents the ground motion parameter matrix validation set; η represents the norm regularization multiplication factor; η = 10 -3 κ(Θ), where κ(Θ) is the condition number of the matrix Θ.
4. A method for finding a ground motion attenuation relationship model based on symbolic learning according to claim 3, characterized in that, The method further includes an iterative optimization step, Set the number of iterations and the regularization coefficient λ; Iterate the algorithm described in Equation (5); During the iteration process, increase the initial threshold δ so that the residual error of the validation set best continually decreases until the optimal solution is given 5. A method for finding a ground motion attenuation relationship model based on symbolic learning according to any one of claims 1-4, characterized in that, The method further includes a data normalization step, The functional expression of the data normalization is shown in Equation (7): Wherein, x represents the candidate function library parameters and the ground motion parameters respectively; The candidate function library parameters include M w , R JB , V S30 , FM, and Z 1.0 ; the ground motion parameters include PGA and PGV.
6. A method for finding a model of ground motion attenuation relationship based on symbol learning according to any one of claims 1-4, characterized in that The method further includes a seismic event selection step, Select seismic events from the seismic event database. The seismic events include ground motion parameters and their corresponding candidate function library parameters; the ground motion parameters construct the ground motion parameter matrix P, and the candidate function library parameters construct the candidate function library Θ.
7. A method for finding a ground motion attenuation relationship model based on symbol learning according to any one of claims 1-4, characterized in that The method further includes a seismic event data processing step, Delete the event that does not list any of the candidate library parameters, where the candidate library parameters include M w , FM, R JB , V S30 and Z 1.0 ; Delete events with less than 5 seismic records, and delete events with focal depths less than 1 km and greater than 20 km.
8. A method for finding a ground motion attenuation relationship model based on symbolic learning according to any one of claims 1-4, characterized in that The candidate function library parameters of the candidate function library Θ further include the following parameters: ln(R JB + 10), which represents the distance saturation term, and the saturation factor is set to 10 km; M w ln(R JB + 10), representing the magnitude saturation term, with the saturation factor set to 10 km.
9. A method for finding a ground motion attenuation relationship model based on symbol learning according to any one of claims 1-4, characterized in that, The method further includes a ground motion parameter equation obtaining step, The functional expression of the ground motion parameter equation is shown in Equation (8): P' = ξ ' best ·Θ' (8) Among them, P' represents the obtained ground motion parameters, and the ground motion parameters include peak ground acceleration PGA and peak ground velocity PGV; ξ' best represents the optimal solution of the obtained coefficient matrix; Θ' represents the obtained candidate function library.
10. A device for finding a model of ground motion attenuation relationship based on symbol learning, characterized in that, The device includes at least one processor; and a memory that stores data and instructions. When the instructions are executed by at least one processor, the steps of the method according to any one of claims 1-9 are implemented.
Citation Information
Patent Citations
Method for building seismic oscillation attenuation relation based on two-stage residual analysis
CN103926621A
Performance-level seismic motion hazard analysis method based on three-layer dataset neural network
US20210302603A1