Deep learning seismic source mechanism inversion system and method based on physical driving
By combining the three-dimensional finite element method and deep learning, physical constraint optimizer is introduced to solve the complexity and accuracy of the source mechanism calculation, and efficient and accurate source mechanism inversion is achieved, which is suitable for complex earthquake environments.
Patent Information
- Application Number
- CN202510504060.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art has problems such as high computational complexity, strong dependence on high-quality data, poor adaptability and insufficient real-time performance in the calculation of source mechanisms, especially in complex geological conditions.
The deep learning source mechanism inversion system based on physics-driven is adopted, combined with the three-dimensional finite element method to simulate seismic wave propagation, and the moment tensor symmetry, Coulomb rupture criterion and stress conservation constraints are introduced, and the source mechanism prediction model is optimized through multi-level feature extraction networks and hybrid training strategies.
It realizes efficient and accurate inversion of the source mechanism, improves the credibility and physical consistency of predictions, reduces the dependence of computing resources, has good robustness and generalization capabilities, and is suitable for complex earthquake environments.
Smart Images

Figure CN120409118A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - technical field of seismology, artificial intelligence, and computer science, and specifically relates to a physical - driven deep - learning focal mechanism inversion system and method. Background Technique
[0002] The focal mechanism solution is an important part of seismological research, which describes the direction of earthquake fault rupture, the sliding mode, and the characteristics of the source stress field. The calculation of the focal mechanism is of great significance for earthquake activity research, seismic tectonic analysis, geodynamic research, and earthquake disaster prediction.
[0003] Currently, the calculation of the focal mechanism mainly relies on traditional methods such as the P - wave first - motion method, far - field waveform inversion method, generalized moment tensor inversion method, and finite - element - based numerical simulation method. The P - wave first - motion method infers the focal mechanism solution by analyzing the polarity distribution of the P - wave first - motion recorded by seismic stations. This method is relatively simple to calculate, but has high requirements for the distribution of stations and is easily affected by noise. The far - field waveform inversion method uses long - period far - field seismic waveform data for focal mechanism inversion, which is suitable for the analysis of large earthquakes. However, for small and medium - sized earthquakes, due to the weak far - field signals, the inversion accuracy is limited. The generalized moment tensor inversion method combines various types of waveform data and uses full - waveform inversion technology to solve the seismic moment tensor to obtain a more accurate focal mechanism solution. However, this method is computationally complex, has high requirements for the accuracy of the velocity model, and has a large computational cost. The finite - element - based numerical simulation method simulates the fault rupture process by establishing a finite - element model to infer the focal mechanism. However, the computational amount is huge and it is difficult to meet the requirements of rapid earthquake emergency.
[0004] The above - mentioned traditional methods all have certain limitations, such as high computational complexity, strong dependence on high - quality data, poor adaptability to inhomogeneous media, and insufficient real - time performance. In addition, these methods usually need to rely on a large amount of seismic station data and have high requirements for the accuracy of phase identification, resulting in limited solution accuracy under complex geological conditions or uneven seismic station coverage.
[0005] In recent years, with the development of Deep Learning technology, data-driven methods have shown great application potential in seismic data processing, seismic phase identification, focal mechanism calculation, etc. For example, deep learning technologies such as Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), and Long Short-Term Memory Network (LSTM) have been widely used in seismic data analysis and achieved good results in aspects such as automatic seismic phase identification and seismic event classification. However, the pure data-driven deep learning method lacks physical constraints and is easily affected by training data, resulting in insufficient generalization ability and difficulty in adapting to complex seismic tectonic environments. In addition, deep learning models usually require a large amount of labeled data for training, and the acquisition cost of high-quality data for focal mechanism solution is relatively high, which limits the practical application of data-driven methods. To effectively solve the deficiencies of the existing technology, there is an urgent need to provide a physically-driven deep learning focal mechanism inversion method. Summary of the Invention
[0006] Aiming at the problems existing in the above-mentioned prior art, the present invention provides a physically-driven deep learning focal mechanism inversion system and method. The system has a simple structure and high intelligence, and can perform efficient and accurate focal mechanism inversion on unknown seismic events. The method has a low implementation cost and high intelligence, can solve the problems of low efficiency, poor accuracy, and insufficient physical interpretation ability existing in the inversion of focal mechanism parameters in the prior art, and can achieve efficient and accurate prediction of focal parameters.
[0007] To achieve the above object, the present invention provides a physically-driven deep learning focal mechanism inversion system, including a seismic wave acquisition module, a seismic wave simulation and feature extraction module, a data preprocessing module, a data feature extractor, a physical constraint optimizer, a model training module, and a focal mechanism prediction module;
[0008] The seismic wave acquisition module is used to obtain a large amount of measured seismic wave waveforms through long-term on-site acquisition operations;
[0009] The seismic wave simulation and feature extraction module is used to use the three-dimensional finite element method to perform high-precision simulation on the seismic wave propagation process under different combinations of focal mechanism parameters, generate a large amount of synthetic seismic wave waveforms, and at the same time, extract spectral features from the large amount of synthetic seismic wave waveforms to obtain a multi-scale feature sequence;
[0010] The data preprocessing module is used to preprocess the large amount of measured seismic waves and synthetic seismic waves, and send the preprocessed data to the data feature extractor;
[0011] The data feature extractor is used to extract seismic waveform data features from a large number of measured seismic waves and synthetic seismic waves through a multi-level feature extraction network combined with dilated convolution, global context multi-scale pooling, and spatial attention mechanism, and obtain an enhanced feature sequence;
[0012] The physical constraint optimizer is used to introduce moment tensor symmetry, Coulomb failure criterion, and stress conservation constraints into the loss function to form a comprehensive loss term function;
[0013] The model training module is used to train the focal mechanism prediction model by adopting a dynamically adjusted learning rate and a hybrid training strategy, and combine the multi-scale feature sequence, enhanced feature sequence, and comprehensive loss term function to optimize the parameters of the focal mechanism prediction model;
[0014] The focal mechanism prediction module is used to realize the rapid inversion of unknown seismic events based on the built-in focal mechanism prediction model and obtain the inversion result.
[0015] In the present invention, the seismic wave simulation and feature extraction module uses the three-dimensional finite element method to generate synthetic waveforms and extract spectral features, which can provide high-quality input data for the prediction process of the focal mechanism prediction module. The data feature extractor extracts multi-scale time-frequency features from the waveforms through a multi-level feature extraction network combined with dilated convolution, global context multi-scale pooling, and spatial attention mechanism, and can establish a mapping relationship with the source parameters. The physical constraint optimizer introduces moment tensor symmetry, Coulomb failure criterion, and stress conservation constraints, ensuring that the prediction results can effectively conform to physical laws and is conducive to improving the prediction accuracy. The model training module uses a hybrid training of synthetic and measured data, dynamically adjusts the learning rate and constraint weights, and can effectively optimize the performance of the focal mechanism prediction model. The focal mechanism prediction module performs rapid focal mechanism inversion on unknown seismic events based on the built-in focal mechanism prediction model, can realize near-real-time focal mechanism analysis and earthquake response, and at the same time provides support for structural analysis, source spectrum estimation, and regional stress field inversion.
[0016] The system has a simple structure and a high degree of intelligence, and can perform efficient and accurate focal mechanism inversion on unknown seismic events.
[0017] The present invention also provides a physical-driven deep learning focal mechanism inversion method, which adopts a physical-driven deep learning focal mechanism inversion system, and includes the following steps:
[0018] Step 1: Collection of the mixed original dataset; Use the seismic wave acquisition module to obtain a large amount of measured seismic wave waveforms through long-term field acquisition operations; At the same time, use the seismic wave simulation and feature extraction module to forward simulate the seismic wave propagation process using the three-dimensional finite element method, and generate a variety of synthetic seismic wave waveforms by adjusting the source mechanism parameters, and then extract multi-scale spectral features that can reflect the source mechanism characteristics from the synthetic seismic wave waveforms;
[0019] Step 2: Preprocess the mixed data; Preprocess the measured seismic wave waveforms and synthetic seismic wave waveforms, including denoising, normalization, and time truncation processing of the data, and then input the preprocessed data into the feature extractor;
[0020] Step 3: Construct a data feature extractor, and use the data feature extractor to extract multi-level enhanced features from the measured seismic wave waveforms and synthetic seismic wave waveforms;
[0021] Step 4: Integrate the multi-scale spectral features and multi-scale spectral features to form a mixed training dataset, and divide the sample dataset into a training set, a validation set, and a test set according to a set ratio;
[0022] Step 5: Construct a source mechanism prediction model;
[0023] S51: Construct a physical constraint optimizer, and use the physical constraint optimizer to introduce fault physical mechanisms, stress tensor conservation, and source tensor constraints into the loss function to form a comprehensive loss term;
[0024] S52: Adopt a neural network structure based on the long short-term memory network to establish a mapping relationship between the spectral features and the source mechanism parameters, and construct an initial source mechanism prediction model; Initialize the initial source mechanism prediction model, and use the training set to train the initial source mechanism prediction model. During the training process, adopt a dynamic learning rate adjustment and mixed training strategy, and optimize the model parameters by minimizing the comprehensive loss term. At the same time, adopt a cosine annealing learning rate schedule and an Adam optimizer, combined with a step-by-step weight adjustment method, to make the model steadily transition from data-driven to physical-driven, significantly improving the robustness and adaptability of the model under complex source conditions. At the same time, use the root mean square error to evaluate the performance of the model, and complete the training process after reaching the set number of iterations or meeting the convergence conditions to obtain an optimized source mechanism prediction model;
[0025] Step 6: Embed the data feature extractor, the physical constraint optimizer, and the source mechanism prediction model into the seismic monitoring system for source mechanism inversion of actual seismic events to achieve fast and accurate parameter prediction;
[0026] S61: waveform data input and preprocessing: receiving raw waveform data from a seismic station. The raw waveform data is a three-component time series, which includes vibration records in the east-west, north-south, and vertical directions, and then performing data preprocessing on the raw waveform data;
[0027] S62: Feature extraction: inputting the preprocessed waveform data into a data feature extractor to generate a multi-scale feature sequence;
[0028] S63: Model reasoning and parameter prediction; input the multi-scale feature sequence as input data into the focal mechanism prediction model, use the focal mechanism prediction model to predict the focal mechanism parameters, and output the focal parameter prediction results.
[0029] As a preferred embodiment, in step 1, the process of the seismic wave simulation and feature extraction module generating a synthetic seismic wave waveform and performing spectrum feature extraction is as follows:
[0030] S11: The earth medium is first modeled as a three-dimensional space, whose physical properties include density ρ, shear modulus μ and Lame constant λ. The earth medium is then discretized into multiple tetrahedral units, and the propagation of seismic waves is obtained according to formula (1) to follow the governing equation of elastic wave dynamics.
[0031]
[0032] Where u is the displacement vector, f is the body force, σ is the stress tensor, Where,∈ is the strain tensor,
[0033] S12: The earthquake source is introduced as a point source, and its mechanical characteristics are represented by the moment tensor M, as shown in formula (2);
[0034]
[0035] Where M is a symmetric matrix that satisfies xy =M yx , M xz =M zx , M yz =M zy ;
[0036] S13: Generate a variety of synthetic seismic waveforms by adjusting the focal mechanism parameters. At the same time, set up virtual receiving points in the simulation area and record the three-component displacement time series u x (t),u y (t),u z (t);
[0037] S14: First, denoise and normalize the synthetic seismic wave waveform, and then perform a fast Fourier transform according to formula (3) to obtain the frequency f k The corresponding complex spectrum U(k). Then, calculate the amplitude spectrum A(f k ) from the complex spectrum according to formula (4) as the spectral feature;
[0038]
[0039] where u(n) is the discrete sampling point of the time-domain waveform, N is the total number of sampling points, k is the frequency index, Δt is the sampling interval, and U(k) is the complex spectrum corresponding to the frequency f k The corresponding complex spectrum;
[0040]
[0041] where Re(U(k)) and Im(U(k)) are the real and imaginary parts of U(k), respectively;
[0042] S15: Extract the values in a specific frequency band from the amplitude spectrum A(f k ) and arrange them in the order of frequency to form a feature vector. By calculating the amplitude spectrum A(f x (t), u y (t), u z (t) for the three-component waveforms respectively, and merge the feature vectors to form a multi-dimensional feature vector. k As an optimization, in step two, the data feature extractor is constructed through the following process:
[0043] As an optimization, in step two, the data feature extractor is constructed through the following process:
[0044] S21: Use 4 convolutional modules as the backbone network of the multi-level feature extraction network. Among them, each convolutional module consists of several one-dimensional convolutional layers and a max-pooling operation;
[0045] S22: In the 3rd and 4th layers of the multi-level feature extraction network, add a dilated convolution design for expanding the receptive field as shown in formula (5); at the same time, in the 2nd, 3rd, and 4th convolutional modules of the multi-level feature extraction network, set the dilation rate to r = 2, 4, 8 respectively corresponding to the expansion of the receptive field from local to global;
[0046]
[0047] where y(i) is the value of the convolutional output feature at the time position i, w(k) is the weight of the convolutional kernel, x(i + r·k) is the value of the input waveform at the time position i + r·k, K is the size of the convolutional kernel, and r is the dilation rate, representing the interval between each sampling point in the convolutional kernel;
[0048] S23: Introduce multi-scale pooling that enhances global context after the 4th layer of the multi-level feature extraction network, as shown in formulas (6), (7), and (8);
[0049] P k = Pool k (F) (6);
[0050]
[0051] where k = 1, 2, 4, 8; F is the input waveform feature, and Pool k represents the pooling operation with scale k, Upsample is linear interpolation upsampling, and Concat is the feature concatenation operation;
[0052] S24: Add a spatial attention design that highlights key features after the multi-scale pooling that enhances global context; First, apply global average pooling and global max pooling to the input waveform feature F respectively to generate two single-channel feature sequences F avg and F max ; Then, concatenate the pooled feature sequences along the channel dimension and generate an attention weight sequence through a convolutional layer with a one-dimensional size of 1, as shown in formula (11); Finally, multiply the weight sequence S with the input waveform feature F point by point according to formula (12) to generate the enhanced feature F att ;
[0053]
[0054] where C is the number of channels of the waveform feature, and F c is the feature sequence of the c-th channel;
[0055] S = σ′(W · Concat(F avg , F max ) + b) (11);
[0056] where σ′ is the sigmoid activation function, W and b are the weights and biases of the convolutional kernel respectively, and the value of S is between 0 and 1;
[0057] F att = S ⊙ F (12);
[0058] where ⊙ represents element-wise multiplication.
[0059] As a preference, in step three, the physical constraint optimizer is constructed through the following process:
[0060] S31: Introduce the symmetry constraint of the moment tensor and enforce the constraint by penalizing the deviation of the asymmetric components, where the symmetry loss term L is obtained through formula (13) sym;
[0061] L sym =(M xy -M yx ) 2 +(M xz -M zx ) 2 +(M yz -M zy ) 2 (13);
[0062] S32: Introduce the Coulomb failure criterion and obtain the Coulomb loss term L through Equation (14) coulomb ;
[0063] L coulomb =max(0,σ n tanφ + c - τ) 2 (14);
[0064] Wherein, σ n is the normal stress, σ n =n·(M·n); τ is the shear stress, τ = n·(M·d); n and d are calculated by geometric transformation from the predicted strike angle φ, dip angle δ and slip angle λ;
[0065] S33: Introduce the stress tensor conservation relation and obtain the stress conservation loss term L through Equation (15) stress ;
[0066]
[0067] Wherein, ▽·σ is calculated by finite difference approximation based on the spatial derivative of the moment tensor M at the hypocenter;
[0068] S34: Measure the difference between the predicted value and the true value output by the data feature extractor through data loss. Specifically, the mean square error is used for measurement, and the data loss term L is obtained according to Equation (16) data ;
[0069]
[0070] Wherein, N is the number of samples, y i,j is the true hypocenter parameter, is the predicted value;
[0071] S35: Perform weighted combination on the symmetry loss term L sym , Coulomb loss term L coulomb , stress conservation loss term L stress and data loss term L data to obtain the comprehensive loss function L according to Equation (17);
[0072] L = L data + λ1L sym + λ2L coulomb + λ3L stress (17);
[0073] Wherein, λ1, λ2, and λ3 are the symmetry loss term L sym , the Coulomb loss term L coulomb and the stress conservation loss term L stress weight coefficients respectively.
[0074] As an optimization, in step four, the sample data set is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0075] As an optimization, in S52 of step five, during the initialization process of the initial focal mechanism prediction model, the weights of the convolutional layer are initialized using the Xavier initialization method. When setting the optimizer, the bias initial value is set to zero, and the output layer of the moment tensor in the physical constraint optimizer is initially set to a random minimum value.
[0076] As an optimization, in S52 of step five, the root mean square error RMSE of each focal parameter is calculated according to formula (19);
[0077]
[0078] As an optimization, in S62 of step six, the data preprocessing of the original waveform data is as follows:
[0079] First, remove environmental noise and instrument drift through band-pass filtering, and retain the main frequency band of the seismic signal; then, normalize the waveform amplitude to a unified range; next, intercept a fixed time window, centered on the arrival time of the first motion of the focal point, to ensure that sufficient seismic phase information is included. If the waveform length is insufficient, pad it with zeros. If it is too long, extract the key segment; at the same time, perform anomaly detection on the original waveform data. If there are significant missing or distorted parts in the waveform, mark it as unavailable and skip the subsequent processing steps.
[0080] As an optimization, in S61 of step six, the original waveform data comes from multiple seismic stations. In S62 of step six, the generated multi-scale feature sequences are multiple groups and correspond to multiple seismic stations respectively. In S63 of step six, the multiple groups of multi-scale feature sequences are respectively input into the focal mechanism prediction model, and multiple focal parameter prediction results are output. Then, the final focal parameters are obtained by fusing the multiple focal parameter prediction results through weighted averaging.
[0081] The present invention provides a physical-driven deep learning seismic source mechanism inversion method to achieve high-precision and high-efficiency calculation of seismic source mechanisms, thereby improving the scientificity and reliability of earthquake disaster early warning and engineering seismic design. First, different from traditional single-scale feature extraction methods, the data feature extractor in the present invention adopts a design architecture that combines a multi-level feature extraction network and an adaptive attention mechanism. The multi-level feature extraction network constructs a multi-level receptive field structure by stacking dilated convolutional layers, capturing both local detail features and long-range spatio-temporal correlations simultaneously; combined with a global context multi-scale pooling module, it realizes the collaborative representation of different-scale temporal patterns in waveform data; and introduces a spatial attention dynamic weighting mechanism, enabling the seismic source mechanism prediction model to focus on the key phase features of seismic signals, effectively simplifying the structure of the seismic source mechanism prediction model. At the same time, it reduces the computational amount during the prediction process and greatly improves the prediction efficiency. Through the collaborative optimization of the above technologies, this feature extractor can generate highly discriminative high-order feature representations, providing a robust and physically meaningful data representation basis for subsequent seismic source parameter inversion. In addition, the physical constraint optimizer in the present invention introduces fault physical mechanisms, including seismic source tensor symmetry, Coulomb failure criterion, and stress tensor conservation relationship, during the neural network training process, constructs a physical consistency loss function, and then combines it with the traditional prediction error to form a mixed loss term. This mechanism effectively suppresses the model's tendency to fit non-physical solutions, improving the credibility and physical interpretability of the inversion results. Furthermore, during the training process of the seismic source mechanism prediction model in the present invention, a mixed training strategy of synthetic data and measured data is adopted, and perturbation testing, cross-validation, and regularization mechanisms are introduced to enhance the generalization ability of the seismic source mechanism prediction model. At the same time, during the training process, a cosine annealing learning rate scheduler and an Adam optimizer are used, combined with a step-by-step weight adjustment method, enabling the model to robustly transition from a data-driven stage to a physical-driven stage, significantly enhancing the robustness and adaptability of the seismic source mechanism prediction model under complex seismic source conditions. In addition, the trained seismic source mechanism prediction model in the present invention can be deployed to existing seismic monitoring systems, and then rapid seismic source mechanism inversion of unknown seismic events can be carried out, realizing near-real-time seismic source analysis and earthquake response, and at the same time providing support for tectonic analysis, seismic source spectrum estimation, and regional stress field inversion.
[0082] Compared with the prior art, the present invention has the following advantages:
[0083] 1) Integrating physical laws and deep learning, improving the credibility and physical consistency of predictions: By introducing seismic source tensor constraints and fault rupture criteria, the present invention overcomes the "black box" problem of traditional neural networks, making the seismic source mechanism inversion results more consistent with seismic physical laws.
[0084] 2) Greatly improved the inversion efficiency and reduced the dependence on computing resources: By using neural networks to replace the traditional Green's function matching calculation process, the iteration process and operation time are significantly reduced, which is applicable to scenarios of rapid earthquake response and massive event processing.
[0085] 3) The obtained model has good scalability and generalization ability: Through joint training of synthetic-measured data and embedding of physical priors, the model can maintain stable prediction accuracy under different geological backgrounds, data loss, or strong noise interference, and has good robustness and transfer adaptability.
[0086] In summary, the physical-driven deep learning focal mechanism inversion method provided by the present invention forms an end-to-end, scalable, and highly robust source parameter inversion scheme through the organic combination of physical constraints, spectral feature extraction, finite element simulation, and deep neural network inference in a multi-module structure. This method not only solves the problems of low efficiency, poor accuracy, and insufficient physical interpretation ability in the inversion of focal mechanism parameters in the prior art, but also lays a solid technical foundation for the future construction of an automated analysis system for practical earthquake monitoring applications, with significant engineering practical value and broad academic research prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 is the architecture flowchart of the present invention;
[0088] Figure 2 is the structural diagram of the multi-level feature extraction network in the present invention;
[0089] Figure 3 is the structural diagram of the physical constraint optimizer in the present invention;
[0090] Figure 4 is the training flowchart of the model in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0091] The object of the present invention is to provide a focal mechanism inversion method based on a deep learning model with physical constraints for the technical problems existing in the prior art in focal mechanism inversion, such as low computational efficiency, insufficient accuracy, and neglect of physical laws. As Figure 1As shown in the figure, first, through the seismic wave simulation and feature extraction module, the three-dimensional finite element method is used to accurately simulate the propagation of seismic waves in the earth's medium, and the spectral features are extracted from the synthetic waveforms to provide high-quality input data for the subsequent deep learning model. Next, the focal mechanism prediction model adopts a neural network structure based on long short-term memory networks to establish the mapping relationship between the spectral features and the focal mechanism parameters. Through the physical constraint optimizer, physical laws such as fault physical mechanisms, stress tensor conservation, and focal tensor symmetry are introduced into the loss function to ensure the physical consistency of the prediction results. The model training module adopts a hybrid training strategy of synthetic and measured data and optimizes the generalization ability of the model through dynamic weight adjustment. Finally, the focal mechanism prediction module realizes the rapid inversion of unknown seismic events based on the focal mechanism prediction model to meet the dual needs of real-time monitoring and scientific research.
[0092] To more clearly clarify the objectives, technical solutions, and advantages of the present invention, the present invention will be further described in detail below in combination with specific implementation steps.
[0093] The present invention provides a physically-driven deep learning focal mechanism inversion system, including a seismic wave acquisition module, a seismic wave simulation and feature extraction module, a data preprocessing module, a data feature extractor, a physical constraint optimizer, a model training module, and a focal mechanism prediction module.
[0094] The seismic wave acquisition module is used to obtain a large amount of measured seismic wave waveforms through long-term on-site acquisition operations.
[0095] The seismic wave simulation and feature extraction module is used to accurately simulate the seismic wave propagation process under different combinations of focal mechanism parameters using the three-dimensional finite element method and generate a large amount of synthetic seismic wave waveforms. At the same time, it is used to extract spectral features from the large amount of synthetic seismic wave waveforms to obtain a multi-scale feature sequence to construct the input feature expression of the focal mechanism. This processing process not only retains the physical characteristics of the focal response but also enhances the sensitivity of the focal mechanism prediction model to focal changes, providing a high-quality data basis for subsequent deep learning modeling.
[0096] As an optimization, the seismic wave simulation and feature extraction module includes three units, namely, a seismic wave propagation simulation unit, a synthetic seismic waveform generation unit, and a spectral feature extraction unit. Among them, the seismic wave propagation simulation unit uses the three-dimensional finite element method to accurately simulate the propagation process of seismic waves in the earth's medium.
[0097] The data preprocessing module is used to preprocess the large amount of measured seismic waves and the large amount of synthetic seismic waves and send the preprocessed data to the data feature extractor.
[0098] The feature extraction module is used to extract seismic waveform data features from a large number of measured seismic waves and a large number of synthetic seismic waves through a multi-level feature extraction network combined with dilated convolution, global context multi-scale pooling, and spatial attention mechanism, and obtain an enhanced feature sequence;
[0099] The physical constraint optimizer is used to introduce moment tensor symmetry, Coulomb failure criterion, and stress conservation constraints into the loss function to form a comprehensive loss term function, so as to enhance the physical consistency of the prediction results;
[0100] The model training module is used to train the focal mechanism prediction model by dynamically adjusting the learning rate and hybrid training strategy, and combine the multi-scale feature sequence, enhanced feature sequence, and comprehensive loss term function to optimize the parameters of the focal mechanism prediction model, so that it can accurately predict the focal mechanism parameters while maintaining physical consistency and generalization ability; as Figure 4 shown, the model training module adopts a hybrid training method to gradually improve the generalization performance of the model by dynamically adjusting the learning rate, regularization strategy (comprehensive loss term), and physical constraint weight.
[0101] The focal mechanism prediction module is used to complete the deployment of the focal mechanism prediction model based on the built-in focal mechanism prediction model and realize the rapid inversion of unknown seismic events to obtain the inversion results.
[0102] In the present invention, the seismic wave simulation and feature extraction module uses the three-dimensional finite element method to generate synthetic waveforms and extract spectral features, which can provide high-quality input data for the prediction process of the focal mechanism prediction module. The data feature extractor extracts multi-scale time-frequency features from the waveforms through a multi-level feature extraction network combined with dilated convolution, global context multi-scale pooling, and spatial attention mechanism, and can establish a mapping relationship with the source parameters. The physical constraint optimizer introduces moment tensor symmetry, Coulomb failure criterion, and stress conservation constraints, ensuring that the prediction results can effectively conform to physical laws and is beneficial to improving the prediction accuracy. The model training module uses a hybrid training of synthetic and measured data, dynamically adjusts the learning rate and constraint weights, and can effectively optimize the performance of the focal mechanism prediction model. The focal mechanism prediction module performs a rapid focal mechanism inversion of unknown seismic events based on the built-in focal mechanism prediction model, can realize near-real-time focal mechanism analysis and seismic response, and at the same time provide support for structural analysis, source spectrum estimation, and regional stress field inversion.
[0103] This system has a simple structure and high intelligence, and can perform efficient and accurate focal mechanism inversion on unknown seismic events.
[0104] The present invention also provides a physically-driven deep learning focal mechanism inversion method, which adopts a physically-driven deep learning focal mechanism inversion system, including the following steps:
[0105] Step 1: Collection of the original dataset;
[0106] Using the seismic wave acquisition module, a large number of measured seismic wave waveforms are obtained through long-term field acquisition operations (acquisition during real seismic events); as an alternative, measured seismic wave waveforms can be obtained from the public seismic dataset library USGS. Approximately 10,000 groups of labeled waveform data can be collected to ensure that the model can adapt to real scenarios;
[0107] Meanwhile, using the seismic wave simulation and feature extraction module, the three-dimensional finite element method is used to forward simulate the seismic wave propagation process, and by adjusting the source mechanism parameters, a variety of synthetic seismic wave waveforms are generated. As an optimization, approximately 100,000 groups of sample data can be generated to effectively cover a variety of geological backgrounds and source characteristics. Then, multi-scale spectral features that can reflect the source mechanism characteristics are extracted from the synthetic seismic wave waveforms to provide high-quality input data for the deep learning model;
[0108] As an optimization, the process of the seismic wave simulation and feature extraction module generating synthetic seismic wave waveforms and performing spectral feature extraction is as follows:
[0109] S11: First, the earth medium is modeled as a three-dimensional space, whose physical properties include density ρ, shear modulus μ, and Lame constant λ. These parameters are set according to geological data or standard earth models. For numerical simulation, the earth medium is then discretized into multiple tetrahedral elements, and according to Equation (1), the propagation of seismic waves follows the governing equation of elastic wave dynamics;
[0110]
[0111] where u is the displacement vector, f is the body force, σ is the stress tensor, and the stress tensor σ is related to the strain tensor ∈ through Hooke's law, so there is: σ = λ(▽·u)I + 2με,
[0112] S12: The source is introduced in the form of a point source, and its mechanical properties are represented by the moment tensor M, as shown in Equation (2);
[0113]
[0114] In the formula, M is a symmetric matrix. The moment tensor M is a 3×3 symmetric matrix, which satisfies M xy = M yx M xz = M zx M yz = M zy ;
[0115] S13: The source force is realized by applying equivalent body forces at the source location. To reduce the influence of boundary reflections on the waveform, absorbing boundary conditions are adopted. By adjusting the source mechanism parameters (such as the strike angle, dip angle, slip angle, and source moment tensor), various synthetic seismic wave waveforms are generated. At the same time, virtual receiving points are set within the simulation area to record the three-component displacement time series u x (t), u y (t), u z (t);
[0116] S14: To reduce the data dimension and extract key information, spectral features of the synthetic waveforms are extracted. First, the synthetic seismic wave waveforms are denoised and normalized, and then the fast Fourier transform is performed according to formula (3) to obtain the complex spectrum U(k) corresponding to the frequency f k . Then, according to formula (4), the amplitude spectrum A(f k ) is calculated from the complex spectrum as the spectral feature; the amplitude spectrum A(f k ) reflects the energy distribution of the waveform at different frequencies and is closely related to the source mechanism parameters.
[0117]
[0118] where u(n) is the discrete sampling point of the waveform in the time domain, N is the total number of sampling points, k is the frequency index corresponding to the frequency f k = k / (NΔt), and Δt is the sampling interval;
[0119]
[0120] where Re(U(k)) and Im(U(k)) are the real and imaginary parts of U(k), respectively;
[0121] S15: To reduce the feature dimension, values in a specific frequency band (such as the low-frequency to mid-frequency range) are extracted from the amplitude spectrum A(f k ) and arranged in frequency order to form a feature vector. By calculating the amplitude spectra A(f x ) for the three-component waveforms u y (t), u z (t), u k )(t) respectively and merging the feature vectors, a multi-dimensional feature vector is formed as the input for the subsequent deep learning model;
[0122] Step 2: Preprocess the mixed data; To effectively unify the data format, the synthetic and measured waveform data need to be preprocessed;
[0123] Preprocess the measured seismic wave waveform and the synthetic seismic wave waveform, including denoising, normalizing, and time truncating the data; as an optimization, during the denoising process, band-pass filtering is preferably used, and the frequency band of 2 - 20 Hz is retained. During the normalization process, the amplitude is scaled to the range of [-1, 1]. During the time truncating process, a fixed length of 60 seconds is truncated; then the preprocessed measured seismic wave waveform and synthetic seismic wave waveform are input into the feature extractor;
[0124] As an optimization, the data feature extractor is constructed through the following process:
[0125] The data feature extractor is the core module of the entire segmentation framework, responsible for extracting general multi-level features from the synthetic seismic waveform and providing high-quality input data for subsequent deep learning inversion tasks. This module uses a convolutional neural network as the backbone network and combines the dilated convolution design for expanding the receptive field, the multi-scale pooling for enhancing the global context, and the spatial attention design for highlighting key features to enhance the feature expression ability. Its design aims to capture local details, global structures, and key frequency band features closely related to the source mechanism parameters in the seismic waveform. The specific construction process is as follows:
[0126] Constructing a multi-level feature extraction network is the core part of constructing the data feature extractor, as Figure 2 shown, mainly used to extract multi-level features from the synthetic seismic waveform, including low-level time-domain fluctuations, high-level frequency-domain information, and global structures. These features can help the model capture key information related to the source mechanism parameters (such as the strike angle, dip angle, and slip angle), providing comprehensive feature support for subsequent inversion tasks.
[0127] Four convolutional modules are used as the backbone network of the multi-level feature extraction network. Each convolutional module consists of several one-dimensional convolutional layers and a max-pooling operation. By gradually compressing the time dimension of the waveform and extracting deeper frequency-domain features simultaneously, the model can have a stronger understanding ability for complex waveform patterns. At the same time, in the third and fourth layers of the multi-level feature extraction network, a dilated convolution design for expanding the receptive field is added. Its role is to capture a wider range of time-frequency context information by expanding the receptive field while keeping the time resolution unchanged. This design has a significant effect on detecting short-time high-frequency fluctuations and long-time low-frequency trends in seismic waveforms. In addition, multi-scale pooling for enhancing global context is introduced after the fourth layer. Global frequency-domain information is extracted through pooling operations at different scales, thereby enhancing the model's understanding ability of the overall structure of the waveform. Finally, a spatial attention design for highlighting key features is added after the multi-scale pooling for enhancing global context. Its role is to highlight the key time-frequency features closely related to the source mechanism while suppressing background noise. This attention mechanism can improve the sensitivity and robustness of the model to details.
[0128] The dilated convolution design for expanding the receptive field is an efficient method for expanding the receptive field, which can capture a larger range of time-frequency context information without increasing the number of convolutional kernel parameters, and is particularly effective for the extraction of multi-scale features in seismic waveforms. By introducing the dilated convolution design for expanding the receptive field, the model can better capture the short-time high-frequency components and long-time low-frequency components in the waveform while maintaining the global perception of the overall structure. The core idea of the dilated convolution design for expanding the receptive field is to introduce holes in the sampling interval of the one-dimensional convolutional kernel to expand the time range covered by the convolutional kernel while not changing the time resolution of the input waveform, as shown in formula (5);
[0129]
[0130] Where y(i) is the value of the convolutional output feature at time position i, w(k) is the weight of the convolutional kernel, x(i+r·k) is the value of the input waveform at time position i+r·k, K is the size of the convolutional kernel, and r is the dilation rate, representing the interval between each sampling point in the convolutional kernel; in the 2nd, 3rd, and 4th convolutional modules of the multi-level feature extraction network, the dilation rates are set to r = 2, 4, and 8 respectively, corresponding to the expansion of the receptive field from local to global; this design of multi-scale dilated convolution can capture waveform features within different time scales. For example, short-time high-frequency components may be related to fault slip characteristics, while long-time low-frequency components are related to the focal depth. The dilated convolution design that expands the receptive field retains the time resolution of the waveform while expanding the receptive field, enabling the model to perceive a wider time-frequency context information. This plays a key role in the extraction of local details and high-frequency features in seismic waveforms. Combining the features of the dilated convolution design that expands the receptive field, the model can capture rich information at multiple scales, providing high-quality input features for subsequent inversion tasks.
[0131] Multi-scale pooling that enhances global context is a key module for enhancing global frequency domain information and is particularly suitable for handling scale differences in time-frequency features of seismic waveforms. The features of seismic waveforms have different physical meanings in different time and frequency ranges. For example, long-time low-frequency components may be related to the overall characteristics of the earthquake source, while short-time high-frequency components reflect local details. Multi-scale pooling that enhances global context enhances the model's comprehensive perception ability of global and local information by extracting pooling features at different scales. Specifically, it extracts global and local features of the input waveform features by setting multiple pooling windows of different scales, such as 1, 2, 4, and 8. Each pooling operation generates a feature of a fixed size, which is then restored to the original feature length through linear interpolation. Finally, all pooling features are concatenated into a unified multi-scale feature representation, as shown in formulas (6), (7), and (8);
[0132] P k =Pool k (F) (6);
[0133]
[0134] Where k = 1, 2, 4, 8; F is the input waveform feature, Pool kPooling operation with scale k, Upsample is linear interpolation upsampling, and Concat is feature concatenation operation; the role of multi-scale pooling for enhancing global context is to capture the global frequency domain information of seismic waveforms. For example, the branch with a larger pooling window can capture the overall trend of long-term low-frequency components, while the branch with a smaller pooling window focuses on the detailed features of short-term high-frequency components. The fusion of such multi-scale features enables the model to take into account both the details and the overall structural features of the waveform, thereby improving the robustness and accuracy of the inversion task. The multi-scale pooling for enhancing global context is constructed after the output of the 4th convolutional module in the multi-level feature extraction network, providing global enhanced features for the subsequent spatial attention design for highlighting key features and the inversion task.
[0135] The construction of the spatial attention design for highlighting key features aims to enhance the extraction of key time-frequency features by focusing on the time distribution of waveform features, while suppressing irrelevant background noise. This mechanism is of great significance for the recognition of time-frequency features closely related to source mechanism parameters in seismic waveforms. The spatial attention design for highlighting key features generates a weight sequence S by calculating the importance distribution of waveform features at each time position, and multiplies it pointwise with the input waveform features to highlight the feature expression in the key region.
[0136] The specific construction steps are as follows: First, apply global average pooling and global max pooling to the input waveform feature F respectively to generate two single-channel feature sequences F avg and F max ;
[0137]
[0138] where C is the number of channels of the waveform feature, and F c is the feature sequence of the c-th channel;
[0139] Secondly, concatenate the pooled feature sequences along the channel dimension and generate an attention weight sequence through a convolutional layer with a one-dimensional size of 1, as shown in formula (11);
[0140] S = σ′(W·Concat(F avg , F max ) + b) (11);
[0141] where σ′ is the sigmoid activation function, W and b are the weights and biases of the convolutional kernel respectively, and the value of S is between 0 and 1;
[0142] Finally, multiply the weight sequence S with the input waveform feature F pointwise according to formula (12) to generate the enhanced feature F att ;
[0143] F att= S⊙F (12);
[0144] In the formula, ⊙ represents element-wise multiplication; a spatial attention design that highlights key features is added after multi-scale pooling that enhances global context, for further refining key features. Its role is to highlight the time-frequency features that contribute to the inversion task, while reducing the interference of irrelevant time periods. The features combined with the spatial attention design that highlights key features not only retain important details, but also strengthen the model's ability to identify key time-frequency features, providing more accurate input features for deep learning inversion.
[0145] Through the combined effects of constructing a multi-level feature extraction network, dilated convolution design that expands the receptive field, multi-scale pooling that enhances global context, and spatial attention design that highlights key features, the data feature extractor finally outputs a high-quality multi-scale feature sequence. This feature sequence combines global frequency domain information and local time domain detail features, providing strong support for subsequent deep learning inversion tasks. The design of integrating multi-module feature outputs realizes the comprehensiveness and accuracy of feature extraction. Among them, the dilated convolution design that expands the receptive field enhances the perception of multi-scale waveform features, the multi-scale pooling that enhances global context strengthens the global frequency domain expression, and the spatial attention design that highlights key features highlights the key features closely related to the focal mechanism. Through the multi-module fusion of the data feature extractor, the model can simultaneously process complex waveform morphologies and feature scale differences, thus providing a basic guarantee for high-precision focal mechanism inversion.
[0146] Through the above Step 1 and Step 2, high-quality and diverse training samples can be provided for the model.
[0147] Step 3: Construct a data feature extractor, and use the data feature extractor to extract multi-level enhanced features from the measured seismic wave waveforms and synthetic seismic wave waveforms;
[0148] As an optimization, the physical constraint optimizer is constructed through the following process:
[0149] S31: Introducing the moment tensor symmetry constraint is the primary task of the physical constraint optimizer, which is used to ensure that the moment tensor predicted by the model meets the basic requirements of seismic physics. However, the moment tensor components directly output by the data feature extractor may violate this condition due to the free fitting of the network. To solve this problem, a symmetry loss term is designed to enforce the constraint by penalizing the deviation of the asymmetric components. Among them, the symmetry loss term L is obtained through formula (13) sym ; this loss term measures the degree of violation of the symmetry condition. If the predicted value completely satisfies the symmetry, then L sym = 0. This constraint ensures the physical consistency of the moment tensor and avoids non-physical results caused by data noise or model overfitting.
[0150] L sym = (M xy - M yx ) 2 + (M xz - M zx ) 2 + (M yz - M zy ) 2 (13);
[0151] S32: Implementing the Coulomb failure criterion constraint is another important step for the physical constraint optimizer, aiming to ensure that the predicted focal mechanism parameters conform to the mechanical conditions of fault slip. According to the Coulomb failure criterion, the condition for fault slip to occur is that the shear stress τ must exceed the combination of the normal stress σ n and the frictional force, as shown by τ ≥ σ n tanφ + c, where φ is the internal friction angle and c is the cohesion, and both are the inherent properties of the medium. To incorporate this criterion into the model, first calculate the shear stress τ and the normal stress σ n of the fault plane based on the predicted moment tensor M, the fault plane normal vector n, and the slip direction vector d. Among them, τ = n·(M·d), σ n = n·(M·n); n and d are obtained through geometric transformation from the predicted strike angle φ, dip angle δ, and slip angle λ; if the prediction result does not satisfy the Coulomb criterion (i.e., τ < σ n tanφ + c), it indicates that the fault slip condition is not achieved and the prediction result deviates from the physical reality. Therefore, the Coulomb loss term L coulomb is obtained through formula (14); this loss term only incurs a penalty when the Coulomb criterion is violated. If τ ≥ σ n tanφ + c, then L coulomb is 0. By introducing this constraint, the model can avoid predicting moment tensors or combinations of fault parameters that cannot trigger fault slip, thereby enhancing the mechanical rationality of the results.
[0152] L coulomb = max(0, σ n tanφ + c - τ) 2 (14);
[0153] S33: Applying the stress conservation relation constraint aims to ensure that the predicted moment tensor is consistent with the physical conservation relation of the surrounding stress field. According to the principle of conservation of momentum, in the absence of additional body forces, the divergence of the stress tensor σ should be zero, as shown by ▽·σ = 0; in earthquake source mechanism inversion, σ can be approximately represented by the stress field of the moment tensor M near the earthquake source. Since directly calculating the divergence of the three-dimensional stress field requires complex medium models and boundary conditions, a simplified method is adopted in this module to construct the constraint by evaluating the divergence of the stress field induced by the moment tensor. Thus, after the stress tensor conservation relation, the stress conservation loss term L is obtained according to formula (15). stress ;
[0154]
[0155] Where ▽·σ is approximately calculated by finite difference based on the spatial derivative of the moment tensor M at the earthquake source point; if the medium model is known, it can be further accurately evaluated in combination with the stress field obtained by simulation. This loss term encourages the model to generate a moment tensor consistent with stress balance and avoid stress field anomalies caused by data-driven factors.
[0156] S34: Designing the comprehensive loss function is a key step in combining the above physical constraints with the data-driven loss, which is used to optimize both the prediction accuracy and physical consistency during training. The data loss measures the difference between the predicted value and the true value output by the data feature extractor. Specifically, the mean square error is used for measurement, and the data loss term L is obtained according to formula (16). data ;
[0157]
[0158] Where N is the number of samples, y i,j is the true earthquake source parameter, is the predicted value;
[0159] S35: The symmetry loss term L sym , Coulomb loss term L coulomb , stress conservation loss term L stress and data loss term L data are weighted and combined, and the comprehensive loss function L is obtained according to formula (17);
[0160] L = L data + λ1L sym + λ2L coulomb + λ3L stress (17);
[0161] Where λ1, λ2, λ3 are the symmetry loss term L sym , Coulomb loss term L coulomb and stress conservation loss term Lstress The weight coefficients; these weights are set to (0.1, 0.05, 0.02) at the beginning of training and gradually increase to 1.0 as the training progresses to ensure that the model first learns the data pattern and then gradually satisfies the physical constraints. This weighting strategy can effectively prevent the model from falling into local extrema and at the same time ensure the physical rationality of the prediction results.
[0162] The implementation details of optimizing the physical constraints involve the specific operations of integrating the above loss function into the deep learning framework. During the training process, the comprehensive loss function optimizes the model parameters through the backpropagation algorithm and uses the Adam optimizer. To further improve the constraint effect, the moment tensor is symmetrized after each iteration (such as taking M xy =(M xy +M yx ) / 2), and it is verified whether the Coulomb criterion is satisfied. If not, the weight of L coulomb is increased. In addition, the implementation of the stress conservation constraint can combine the medium parameters of the seismic wave simulation module to further refine the accuracy of the divergence calculation.
[0163] Step 4: Use the multi-scale spectral features in Step 1 and the multi-level enhanced features in Step 3 as the sample data set, and divide the sample data set into a training set, a validation set, and a test set according to a set ratio. Among them, the training set is used to optimize the parameters of the focal mechanism prediction model, the validation set is used to adjust the hyperparameters of the focal mechanism prediction model, and the test set is used to evaluate the final performance of the focal mechanism prediction model; preferably, the sample data set is divided into a training set, a validation set, and a test set according to a ratio of 8:1:1.
[0164] Step 5: Construct a focal mechanism prediction model;
[0165] S51: Construct a physical constraint optimizer, and use the physical constraint optimizer to introduce the fault physical mechanism, stress tensor conservation, and focal tensor constraints into the loss function to form a comprehensive loss term;
[0166] As Figure 3 shown, the core task of the physical constraint optimizer is to integrate the seismic physical laws into the deep learning model to ensure that the prediction results output by the data feature extractor not only depend on data driving but also conform to the physical consistency of the focal mechanism, thereby effectively improving the credibility and scientificity of the inversion results.
[0167] This module designs a series of regularization terms based on physical laws to constrain the focal mechanism parameters output by the model (such as the strike angle φ, dip angle δ, slip angle λ, and six independent components M xx 、M xy 、M xz 、M yy 、Myz 、M zz ), preventing the model from generating results that deviate from physical reality. The design of the physically constrained optimizer combines the moment tensor symmetry, the Coulomb failure criterion, and the stress conservation relationship to form a comprehensive loss function that works together with the data loss to optimize the model.
[0168] S52: Using a neural network structure based on long short-term memory network, a mapping relationship between spectrum characteristics and focal mechanism parameters is established to construct an initial focal mechanism prediction model; the initial focal mechanism prediction model is initialized and trained using a training set, such as Figure 4 As shown in the figure, during the training process, dynamic adjustment of learning rate and hybrid training strategy are adopted to optimize the model parameters by minimizing the comprehensive loss term. At the same time, cosine annealing learning rate scheduling and Adam optimizer are used, combined with a stepwise weight adjustment method, to enable the model to robustly transition from the data-driven to the physical-driven stage, significantly improving the model's robustness and adaptability under complex source conditions. At the same time, the root mean square error is used to evaluate the model's performance. The training process is completed after reaching the set number of iterations or meeting the convergence condition, and the optimized focal mechanism prediction model is obtained.
[0169] As a further preference, in order to enhance the robustness of the model to noise and data missing, a data enhancement strategy is added to the training set, specifically, randomly adding Gaussian noise or truncating part of the time period.
[0170] As a further optimization, in order to ensure the convergence speed and stability of the focal mechanism prediction model, during the initialization of the initial focal mechanism prediction model, the convolutional layer weights are initialized using the Xavier method to ensure that the variances of the input and output are consistent and to avoid the problem of gradient vanishing or exploding. When setting the optimizer, the initial bias value is set to zero to reduce the impact of the initial deviation on the training, and the torque tensor output layer in the physical constraint optimizer is initially set to a random minimum value to ensure that the initial training does not deviate too far from the physical constraints.
[0171] Dynamically adjusting the learning rate and training strategy is a core component of the training process, balancing the model's convergence speed and accuracy. Training is divided into multiple phases, totaling 100 epochs, with the number of samples per batch set to 640 to balance computational efficiency and gradient stability. The learning rate is scheduled using a cosine annealing strategy, as shown in Equation (18); this strategy gradually decreases the learning rate from its initial value, ensuring rapid learning of data patterns in the early stages and refined parameter adjustments in the later stages.
[0172]
[0173] Where η max =0.001,η min= 0.0001, T = 100, where T is the cycle length;
[0174] To enhance the model's adaptability to physical laws, the physical constraint weights λ1, λ2, and λ3 adopt a dynamic adjustment strategy: keep the initial values for the first 20 epochs and focus on optimizing L data to fit the data; from the 20th to the 60th epoch, increase the weights by 20% every 10 epochs to gradually strengthen the physical constraints; after the 60th epoch, fix them at 1.0 to ensure that physical consistency dominates. In addition, add an L2 regularization term with a coefficient of 0.0001 to prevent the model from overfitting complex waveform data. During training, evaluate the loss on the validation set every 5 epochs. If the validation loss does not decrease for 10 consecutive epochs, trigger the early stopping mechanism to avoid overtraining.
[0175] The integration of hybrid training and physical constraints is an innovation in the present invention. By jointly training with synthetic and measured data and gradually incorporating physical constraints, the generalization ability of the model is improved. Synthetic data provides diverse coverage of source parameters to help the model learn a wide range of waveform patterns; measured data introduces the noise and complexity of real earthquakes to enhance the model's adaptability to actual scenarios. In each epoch, the training batches are randomly mixed with synthetic and measured data in a ratio of 8:2 to ensure that the model benefits from the advantages of both types of data. The integration of physical constraints is achieved through a comprehensive loss function. The data loss L data ensures the agreement between the predicted value and the true value, while the symmetry loss L sym , Coulomb loss L coulomb and stress conservation loss L stress restrict the physical rationality of the prediction results. To verify the constraint effect, output the symmetry deviation of the moment tensor and the satisfaction rate of the Coulomb criterion every 10 epochs. If the deviation is too large or the satisfaction rate is lower than 90%, temporarily increase the corresponding λ to 50% until the expected standard is reached. This hybrid training strategy enables the model to gradually transition from data-driven to physics-driven, ensuring that the prediction results are both accurate and interpretable.
[0176] The performance evaluation and model verification process is the final step of the training framework, used to test the prediction accuracy and robustness of the model. After training is completed, evaluate the model performance on the test set and calculate the root mean square error RMSE of each source parameter according to formula (19);
[0177]
[0178] At the same time, check the satisfaction degree of physical constraints, including the moment tensor symmetry error (L sym < 0.01), Coulomb criterion satisfaction rate (> 95%) and stress conservation loss (L stress(< 0.05). To test the generalization ability of the model, additional verification is carried out on the measured data with noise and the synthetic data without known source parameters. If the RMSE is higher than the preset threshold (such as 5° or 10% source moment error) or the physical constraints are not satisfied, then the hyperparameters are adjusted (such as increasing the number of training epochs or adjusting the value of λ), and retraining is performed. The final model needs to achieve high accuracy on both synthetic and measured data while maintaining physical consistency before entering the deployment stage.
[0179] Through the data collection and preprocessing process, initialization of model parameters and optimizer settings, dynamic adjustment of the learning rate and training strategy, integration of hybrid training and physical constraints, and performance evaluation and model verification process, the complete process from data preparation to model optimization is completed. The trained focal mechanism prediction model can effectively utilize synthetic and measured waveform data, and combine physical constraints to predict focal mechanism parameters, laying a solid foundation for the subsequent deployment of the focal mechanism prediction system.
[0180] By means of a reasonable training strategy, the synthetic seismic wave waveforms and multi-scale spectral features generated by the seismic wave simulation and feature extraction module, the multi-level enhanced features extracted by the data feature extractor, and the comprehensive loss function constructed by the physical constraint optimizer are effectively combined to optimize the parameters of the focal mechanism prediction model, enabling it to accurately predict focal mechanism parameters while maintaining physical consistency and generalization ability.
[0181] Step 6: Embed the data feature extractor, physical constraint optimizer, and focal mechanism prediction model into the seismic monitoring system for focal mechanism inversion of actual seismic events, and achieve fast and accurate parameter prediction by providing efficient automated analysis capabilities, meeting the dual requirements of real-time monitoring and scientific research. To seamlessly embed the focal mechanism prediction model into the seismic monitoring system and optimize its operation efficiency, the trained focal mechanism prediction model is saved in a lightweight format, deployed on a high-performance server or edge computing device, and docked with the data stream of the seismic network through a standard interface to facilitate real-time reception of waveform data collected by seismic stations and complete the full process. Preferably, the seismic monitoring system can adopt a modular design, divided into a data reception and preprocessing module, a model inference module (focal mechanism prediction model), and a result postprocessing module, and each module runs in parallel through multi-threading to improve processing efficiency. To meet the real-time requirements, the seismic monitoring system can adopt hardware acceleration technology to shorten the single prediction time to the millisecond level, and support distributed deployment so that multiple servers can cooperate to process data of large-scale seismic events and improve data throughput. The operation log records the input waveform, output parameters, and consistency evaluation of each prediction for subsequent analysis and improvement. In addition, the seismic monitoring system can provide a user interaction interface to support offline analysis of historical waveforms or adjustment of operation parameters to adapt to different application scenarios.
[0182] S61: Waveform data input and preprocessing; Waveform data input and preprocessing is the first step in deploying the focal mechanism prediction system, aiming to convert the measured waveform data from the seismic monitoring network into an input format recognizable by the model.
[0183] Receive the original waveform data from seismic stations. The original waveform data is a three-component time series, which includes vibration records in the east-west, north-south, and vertical directions. The sampling rate and duration vary depending on the station configuration.
[0184] To ensure consistency with the training data, all input waveforms need to undergo standardized preprocessing. The specific process is as follows: First, remove environmental noise and instrument drift through band-pass filtering, and retain the main frequency band of the seismic signal; then, normalize the waveform amplitude to a unified range to avoid affecting feature extraction due to amplitude differences; then intercept a fixed time window, centered on the arrival time of the first motion of the earthquake source, to ensure sufficient seismic phase information is included. If the waveform length is insufficient, it can be padded with zeros; if it is too long, the key segment is intercepted. As an option, to improve the robustness of the system, the preprocessing process also includes anomaly detection. If there are significant missing or distorted parts in the waveform, it is marked as unavailable and the subsequent processing steps are skipped.
[0185] S62: Directly input the preprocessed waveform data into the data feature extractor, and extract high-quality time-frequency features through a multi-level feature extraction network, dilated convolution design, multi-scale pooling, and spatial attention design to generate a multi-scale feature sequence, where the multi-scale feature sequence is consistent with the input format in the training stage.
[0186] S63: Model inference and parameter prediction; Model inference and parameter prediction is the core part of the system, and use the trained focal mechanism prediction model to quickly predict the focal mechanism parameters from the multi-scale feature sequence.
[0187] Take the multi-scale feature sequence as the input data and input it into the focal mechanism prediction model, use the focal mechanism prediction model to predict the focal mechanism parameters, and output the prediction results of the focal parameters; the focal parameters include the strike angle, dip angle, slip angle, and six independent components of the moment tensor. In the focal mechanism prediction model, the inference process uses the forward propagation algorithm, and the calculation time is extremely short, much lower than the time consumption of traditional methods, meeting the requirements of real-time applications.
[0188] To ensure prediction stability, joint inference of waveforms from multiple stations can be supported: obtain the corresponding raw waveform data from multiple seismic stations, then generate multi-scale feature sequences for multiple seismic stations through a data feature extractor. Then, input them into the focal mechanism prediction model respectively to generate multiple predicted focal parameter results. Finally, obtain the final focal parameters by fusing multiple predicted focal parameter results through weighted averaging, where the weights are based on the distance between seismic stations and data quality. This multi-station fusion strategy can effectively reduce the impact of single waveform noise or anomalies on the results and improve the robustness of the prediction.
[0189] To effectively verify and correct the focal parameters output by the model and ensure their compliance with physical laws and practical applications, physical consistency verification and result post-processing can be carried out. The inferred moment tensor is first verified for symmetry. If the off-diagonal components are found to be asymmetric, symmetrization processing is performed. Subsequently, based on the predicted strike angle, dip angle, slip angle, and moment tensor, the shear stress and normal stress on the fault plane are calculated to verify whether the mechanical conditions for fault slip are met. If the conditions are not met, the result is marked and the confidence level is reduced. In addition, the balance of the stress field is also checked. If the deviation is too large, it is prompted that more station data may be required. In the post-processing stage, the moment tensor is decomposed into a scalar seismic moment and fault parameters, generating a user-friendly output format, such as a double couple solution or a fault plane solution, for direct use by seismologists. All results are accompanied by physical consistency evaluation indicators to ensure that users understand the credibility of the prediction.
[0190] The output result evaluation and feedback mechanism is the final step of the seismic monitoring system, which is used to evaluate the reliability and practicality of the prediction results and provide feedback to users. Each set of prediction results is accompanied by a confidence score, taking into account both data fitting and physical consistency. If the confidence level is lower than expected, the seismic monitoring system will prompt the user that more station data or manual verification may be required. The predicted parameters are output in a standard format, compatible with seismological analysis tools. To continuously improve the model, the seismic monitoring system collects the waveforms and results of each prediction and feeds them back to the training dataset. If systematic deviations are found, the model parameters can be updated regularly. In addition, the system supports users to upload historical waveforms for analysis or adjust the configuration to meet specific needs, ensuring its wide applicability in real-time monitoring and scientific research.
[0191] Through waveform data input and preprocessing, model inference and parameter prediction, physical consistency verification and result post-processing, and output result evaluation and feedback mechanism, the seismic monitoring system deploying the focal prediction mechanism can achieve efficient inversion from measured waveforms to focal parameters, and significantly improve the inversion speed. Specifically, an inversion speed of seconds can be achieved. At the same time, the scientific nature of the results is ensured through physical constraints, making it applicable to seismic monitoring, disaster assessment, and geodynamic research, providing a solid guarantee for the practical application of the present invention.
[0192] The present invention provides a physical-driven deep learning seismic source mechanism inversion method to achieve high-precision and high-efficiency calculation of seismic source mechanisms, thereby improving the scientific nature and reliability of earthquake disaster early warning and engineering seismic design. First, different from traditional single-scale feature extraction methods, the data feature extractor in the present invention adopts a design architecture that combines a multi-level feature extraction network and an adaptive attention mechanism. The multi-level feature extraction network constructs a multi-level receptive field structure by stacking dilated convolutional layers, capturing local detail features and long-range spatio-temporal correlations simultaneously; combined with a global context multi-scale pooling module, it realizes the collaborative representation of different-scale temporal patterns in waveform data; and introduces a spatial attention dynamic weighting mechanism, enabling the seismic source mechanism prediction model to focus on the key phase features of seismic signals, effectively simplifying the structure of the seismic source mechanism prediction model. At the same time, it reduces the computational amount during the prediction process and greatly improves the prediction efficiency. Through the collaborative optimization of the above technologies, this feature extractor can generate high-order feature representations with strong discriminability, providing a robust and physically meaningful data representation basis for subsequent seismic source parameter inversion. In addition, in the present invention, the physical constraint optimizer introduces fault physical mechanisms, including seismic source tensor symmetry, Coulomb failure criterion, and stress tensor conservation relationship, during the neural network training process, constructs a physical consistency loss function, and then combines it with the traditional prediction error to form a mixed loss term. This mechanism effectively suppresses the fitting tendency of the model to non-physical solutions and improves the credibility and physical interpretability of the inversion results. Furthermore, during the training process of the seismic source mechanism prediction model in the present invention, a mixed training strategy of synthetic data and measured data is adopted, and perturbation testing, cross-validation, and regularization mechanisms are introduced to enhance the generalization ability of the seismic source mechanism prediction model. At the same time, during the training process, a cosine annealing learning rate scheduling and an Adam optimizer are used, combined with a step-by-step weight adjustment method, enabling the model to robustly transition from a data-driven stage to a physical-driven stage, significantly enhancing the robustness and adaptability of the seismic source mechanism prediction model under complex seismic source conditions. In addition, the trained seismic source mechanism prediction model in the present invention can be deployed to existing seismic monitoring systems, and then rapid seismic source mechanism inversion of unknown earthquake events can be carried out, realizing near-real-time seismic source analysis and earthquake response, and at the same time providing support for tectonic analysis, seismic source spectrum estimation, and regional stress field inversion.
[0193] In summary, the present invention proposes a physical-constrained deep learning model seismic source mechanism inversion method, realizing a complete process from seismic wave simulation to real-time prediction. Compared with traditional methods, the present invention significantly improves the calculation efficiency and prediction accuracy, and at the same time enhances the scientific nature of the results through physical constraints. The seismic monitoring system embedded with the seismic source mechanism prediction model can support multi-station joint inference and real-time operation, and is applicable to seismic monitoring, disaster assessment, and scientific research, demonstrating the advantages of the combination of data-driven and physical laws, and providing an efficient and practical solution for the field of seismology.
Claims
1. A physical-driven deep learning focal mechanism inversion system, characterized in that It includes a seismic wave acquisition module, a seismic wave simulation and feature extraction module, a data preprocessing module, a data feature extractor, a physical constraint optimizer, a model training module, and a focal mechanism prediction module; The seismic wave acquisition module is used to obtain a large amount of measured seismic wave waveforms through long-term on-site acquisition operations; The seismic wave simulation and feature extraction module is used to accurately simulate the seismic wave propagation process under different combinations of focal mechanism parameters using the three-dimensional finite element method, and generate a large amount of synthetic seismic wave waveforms. At the same time, it is used to extract spectral features from the large amount of synthetic seismic wave waveforms to obtain a multi-scale feature sequence; The data preprocessing module is used to preprocess the large amount of measured seismic waves and synthetic seismic waves, and send the preprocessed data to the data feature extractor; The data feature extractor is used to extract seismic waveform data features from the large amount of measured seismic waves and synthetic seismic waves through a multi-level feature extraction network combined with dilated convolution, global context multi-scale pooling, and spatial attention mechanism to obtain an enhanced feature sequence; The physical constraint optimizer is used to introduce moment tensor symmetry, Coulomb failure criterion, and stress conservation constraints into the loss function to form a comprehensive loss term function; The model training module is used to train the focal mechanism prediction model using a dynamically adjusted learning rate and a hybrid training strategy, and combine the multi-scale feature sequence, the enhanced feature sequence, and the comprehensive loss term function to optimize the parameters of the focal mechanism prediction model; The focal mechanism prediction module is used to achieve rapid inversion of unknown seismic events based on the built-in focal mechanism prediction model and obtain the inversion result.
2. A physically-driven deep learning focal mechanism inversion method, which uses a physically-driven deep learning focal mechanism inversion system as described in claim 1, characterized in that It includes the following steps: Step 1: Collection of the mixed original dataset; Use the seismic wave acquisition module to obtain a large amount of measured seismic wave waveforms through long-term on-site acquisition operations; At the same time, use the seismic wave simulation and feature extraction module to forward simulate the seismic wave propagation process using the three-dimensional finite element method, and generate a variety of synthetic seismic wave waveforms by adjusting the focal mechanism parameters, and then extract the multi-scale spectral features that can reflect the characteristics of the focal mechanism from the synthetic seismic wave waveforms; Step 2: Preprocess the mixed data; Preprocess the measured seismic wave waveforms and synthetic seismic wave waveforms, including denoising, normalization, and time truncation processing of the data, and then input the preprocessed data into the feature extractor; Step 3: Construct a data feature extractor, and use the data feature extractor to extract multi-level enhanced features from the measured seismic wave waveforms and synthetic seismic wave waveforms; Step 4: Integrate the multi-scale spectral features and multi-scale spectral features to form a mixed training dataset, and divide the sample dataset into a training set, a validation set, and a test set according to a set ratio; Step 5: Construct a focal mechanism prediction model; S51: Construct a physical constraint optimizer, and use the physical constraint optimizer to introduce fault physical mechanisms, stress tensor conservation, and focal tensor constraints into the loss function to form a comprehensive loss term; S52: Adopt a neural network structure based on long short-term memory networks to establish the mapping relationship between spectral features and focal mechanism parameters, and construct an initial focal mechanism prediction model; initialize the initial focal mechanism prediction model, and use the training set to train the initial focal mechanism prediction model. During the training process, adopt a dynamic learning rate adjustment and hybrid training strategy, and optimize the model parameters by minimizing the comprehensive loss term. At the same time, adopt a cosine annealing learning rate scheduler and an Adam optimizer, and cooperate with the step-by-step weight adjustment method to enable the model to robustly transition from data-driven to physics-driven stages, significantly improving the robustness and adaptability of the model under complex source conditions. At the same time, use the root mean square error to evaluate the performance of the model, and complete the training process after reaching the set number of iterations or meeting the convergence conditions to obtain an optimized focal mechanism prediction model; Step Six: Embed the data feature extractor, physical constraint optimizer, and focal mechanism prediction model into the seismic monitoring system for focal mechanism inversion of actual seismic events to achieve fast and accurate parameter prediction; S61: Waveform data input and preprocessing; Receive the original waveform data from the seismic station. The original waveform data is a three-component time series, which includes vibration records in the east-west, north-south, and vertical directions, and then perform data preprocessing on the original waveform data; S62: Feature extraction; Input the preprocessed waveform data into the data feature extractor to generate a multi-scale feature sequence; S63: Model inference and parameter prediction; Use the multi-scale feature sequence as the input data and input it into the focal mechanism prediction model, and use the focal mechanism prediction model to predict the focal mechanism parameters and output the focal parameter prediction result.
3. A physical-driven deep learning seismic source mechanism inversion method according to claim 2, characterized in that In Step One, the process of generating synthetic seismic wave waveforms and performing spectral feature extraction by the seismic wave simulation and feature extraction module is as follows: S11: First, model the earth medium as a three-dimensional space, whose physical properties include density ρ, shear modulus μ, and Lame constant λ. Then, discretize the earth medium into multiple tetrahedral elements, and obtain the control equation for seismic wave propagation following elastic wave dynamics according to Equation (1); Among them, u is the displacement vector, f is the body force, σ is the stress tensor, σ = λ(▽·u)I + 2μ∈, where ∈ is the strain tensor. S12: Introduce the source in the form of a point source, and its mechanical characteristics are represented by the moment tensor M, as shown in Equation (2); where M is a symmetric matrix satisfying M xy = M yx , M xz = M zx , M yz = M zy ; S13: Generate a variety of synthetic seismic wave waveforms by adjusting the source mechanism parameters. At the same time, set virtual receiving points in the simulation area to record the three-component displacement time series \(u_{x}(t)\), \(u_{y}(t)\), \(u_{z}(t)\); x (t), \(u_{x}\) y (t), \(u_{y}\) z (t); S14: First, perform denoising and normalization on the synthetic seismic wave waveform, and then perform a fast Fourier transform according to formula (3) to obtain the frequency f k and the corresponding complex spectrum U(k). Then, calculate the amplitude spectrum A(f k ) from the complex spectrum according to formula (4) as the spectral feature; where \(u(n)\) is the discrete sampling point of the time-domain waveform, \(N\) is the total number of sampling points, \(k\) is the frequency index, \(\Delta t\) is the sampling interval, and \(U(k)\) is the complex spectrum corresponding to the frequency \(f\). k The corresponding complex spectrum; In the formula, Re(U(k)) and Im(U(k)) are the real part and imaginary part of U(k), respectively; S15: Extract the values of a specific frequency band from the amplitude spectrum A(f k ), arrange them in the order of frequency, and form a feature vector. By calculating the amplitude spectra A(f x ) for the three-component waveforms u y (t), u z (t), u k (t) respectively, and merge the feature vectors to form a multi-dimensional feature vector.
4. A physical-driven deep learning seismic source mechanism inversion method according to claim 2, characterized in that In Step Two, the data feature extractor is constructed through the following process: S21: Use 4 convolutional modules as the backbone network of the multi-level feature extraction network. Among them, each convolutional module consists of several one-dimensional convolutional layers and a max-pooling operation; S22: In the 3rd and 4th layers of constructing the multi-level feature extraction network, add a dilated convolution design to expand the receptive field, as shown in Equation (5); at the same time, in the 2nd, 3rd, and 4th convolutional modules of constructing the multi-level feature extraction network, set the dilation rate r = 2, 4, 8 corresponding to the expansion of the receptive field from local to global; Where y(i) is the value of the convolution output feature at time position i, w(k) is the weight of the convolution kernel, x(i + r·k) is the value of the input waveform at time position i + r·k, K is the size of the convolution kernel, and r is the dilation rate, representing the interval between each sampling point in the convolution kernel; S23: Introduce multi-scale pooling that enhances global context after the 4th layer of the multi-level feature extraction network, as shown in formulas (6), (7), and (8); P k = Pool k (F) (6); where k = 1, 2, 4, 8; F is the input waveform feature, Pool k represents the pooling operation with scale k, Upsample is the linear interpolation upsampling, and Concat is the feature concatenation operation; S24: Add a spatial attention design that highlights key features after multi-scale pooling that enhances the global context; first, apply global average pooling and global max pooling to the input waveform feature F respectively to generate two single-channel feature sequences F avg and F max ; then, concatenate the pooled feature sequences along the channel dimension, and generate an attention weight sequence through a convolutional layer with a one-dimensional size of 1, as shown in formula (11); finally, multiply the weight sequence S and the input waveform feature F point by point according to formula (12) to generate the enhanced feature F att ; Where C is the number of channels of the waveform feature, and F c is the feature sequence of the c-th channel; S = σ′(W·Concat(F avg , F max ) + b) (11); Where σ′ is the sigmoid activation function, W and b are the weight and bias of the convolution kernel respectively, and the value of S is between 0 and 1; F att = S⊙F (12); Where ⊙ represents element-wise multiplication.
5. A physical-driven deep learning seismic source mechanism inversion method according to claim 2, characterized in that, In step three, the physical constraint optimizer is constructed through the following process: S31: Introduce the symmetry constraint of the moment tensor and enforce the constraint by penalizing the deviation of the asymmetric components, where the symmetry loss term L is obtained by formula (13). sym ; L sym = (M xy - M yx ) 2 + (M xz - M zx ) 2 + (M yz - M zy ) 2 (13); S32: Introduce the Coulomb failure criterion and obtain the Coulomb loss term L through formula (14) coulomb ; L coulomb = max(0, σ n tanφ + c - τ) 2 (14); where, σ n is the normal stress, σ n = n·(M·n); τ is the shear stress, τ = n·(M·d); n and d are calculated by geometric transformation from the predicted strike angle φ, dip angle δ, and slip angle λ S33: Introduce the stress tensor conservation relation and obtain the stress conservation loss term L through formula (15) stress ; Where ▽·σ is calculated by finite difference approximation based on the spatial derivative of the moment tensor M at the source point; S34: Measure the difference between the predicted value and the true value output by the data feature extractor through data loss. Specifically, use the mean square error for measurement and obtain the data loss term L according to formula (16). data ; where N is the number of samples, and y i,j is the true source parameter, and is the predicted value; S35: Combine the symmetry loss term \(L\) sym , the Coulomb loss term \(L\) coulomb , the stress conservation loss term \(L\) stress and the data loss term \(L\) data with weights, and obtain the comprehensive loss function \(L\) according to Equation (17); L = L data + λ1L sym + λ2L coulomb + λ3L stress (17); where λ1, λ2, and λ3 are the weight coefficients of the symmetry loss term L sym , the Coulomb loss term L coulomb , and the stress conservation loss term L stress respectively.
6. The inversion method for the seismic source mechanism based on physics-driven deep learning according to claim 2, wherein In step four, the sample data set is divided into a training set, a validation set, and a test set in the ratio of 8:1:
1.
7. A physical-driven deep learning seismic source mechanism inversion method according to claim 2, characterized in that In S52 of step five, during the initialization process of the initial source mechanism prediction model, the weights of the convolutional layer are initialized using the Xavier initialization method. When setting the optimizer, the initial value of the bias is set to zero, and the output layer of the moment tensor in the physical constraint optimizer is initially set to a random minimum.
8. A physical-driven deep learning seismic source mechanism inversion method according to claim 2, characterized in that In S52 of step five, calculate the root mean square error RMSE of each source parameter according to formula (19); 9. A physical-driven deep learning seismic source mechanism inversion method according to claim 2, characterized in that In S62 of step six, the data preprocessing of the original waveform data is as follows: First, remove environmental noise and instrument drift through band-pass filtering, and retain the main frequency band of the seismic signal; then, normalize the waveform amplitude to a unified range; next, intercept a fixed time window, centered on the arrival time of the source initial motion, to ensure that sufficient seismic phase information is included. If the waveform length is insufficient, pad it with zeros. If it is too long, extract the key segment; at the same time, perform anomaly detection on the original waveform data. If there are significant missing or distorted parts in the waveform, mark it as unavailable and skip the subsequent processing steps.
10. A physical-driven deep learning seismic source mechanism inversion method according to claim 2, characterized in that In S61 of step six, the original waveform data comes from multiple seismic stations. In S62 of step six, the generated multi-scale feature sequences are in multiple groups and correspond to multiple seismic stations respectively. In S63 of step six, the multiple groups of multi-scale feature sequences are respectively input into the source mechanism prediction model, and multiple source parameter prediction results are output. Then, the final source parameters are obtained by fusing the multiple source parameter prediction results through weighted averaging.
Citation Information
Cited By
Farmland drainage ditch purification effect evaluation method and system based on Internet of Things
CN120875275A
Earthquake fault detection method and device, storage medium and electronic equipment
CN121657132A
Seismic data noise suppression method and system based on deep learning
CN121806117A
Neural network seismic source mechanism inversion optimization method based on brightness field representation
CN122283886A
Deep learning full waveform inversion method and system based on physical law constraint
CN122488219A