Method and device for automatically picking up seismic phase based on deep learning
Through the deep learning-based PSNet neural network, feature extraction and prediction of seismic waveform data is solved, and efficient and accurate seismic phase pickup is achieved.
Patent Information
- Application Number
- CN202510078478.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, the automatic picking of seismic signals is inefficient and the error is large, making it difficult to cope with the problem of rapid growth in the number of earthquakes.
Using a deep learning-based method, by constructing a PSNet neural network, the three-component seismic waveform data are preprocessed and feature extracted to generate the predicted probability of P and S wave arrival time.
It realizes efficient and accurate automatic seismic phase pickup, improves processing efficiency, reduces errors, and can adapt to different signal-to-noise ratios and instrument data robustness.
Smart Images

Figure CN119960026A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning automatic target signal picking, and specifically relates to a method and device for automatic seismic phase picking based on deep learning. Background Art
[0002] Accurately and efficiently extracting earthquake signals from massive amounts of data is the basis for earthquake data analysis. Accurate waveform arrival time picking helps to more finely analyze the source location, source depth, source type, and formation parameters. At the same time, these data can also be used for aftershock analysis and extraction, providing basic data for establishing aftershock sequences, analyzing earthquake occurrence patterns, and subsequent earthquake prevention and disaster reduction work.
[0003] Picking up earthquake phases is an important foundation for conducting seismology-related research. With the continuous growth in the number of seismic stations and the density of the network, the number of earthquakes that can be monitored has increased exponentially, making the work of picking up seismic phases more and more difficult. The frequency composition of seismic waves recorded by seismic stations is particularly complex, and is often accompanied by intricate noise interference. The traditional manual selection and processing method is time-consuming, labor-intensive, inefficient, and has large errors. It is difficult to cope with the current situation of the rapid increase in the number of earthquakes, and it is even more unable to meet the needs of real-time processing such as earthquake early warning and rapid emergency output reports. Therefore, people have put forward higher requirements and challenges for the automatic picking technology of earthquake phases. The development of modern seismology has an increasingly strong demand for efficient automatic earthquake detection methods.
[0004] At present, due to the complexity of the earthquake source and the medium, seismic signals show complex characteristics in the time and frequency domains. The essence of arrival time extraction is to analyze and process the waveform features. The complex seismic wave morphology makes it difficult for traditional feature analysis methods to cope with it. The fixed features extracted by traditional time domain and frequency domain methods are difficult to apply to complex seismic wave morphology. They rely on manual feature definition and analysis during the arrival time analysis process. This makes the features lack in completeness and universality, so traditional waveform picking methods encounter difficulties in extracting arrival times.
[0005] In early work, waveform extraction mainly relied on manual identification, which was feasible when processing a small amount of data. With the improvement of observation instruments and the increase in seismic stations, the amount of seismic data has increased significantly. At the same time, research on different directions such as landslides and secondary disasters has gradually increased, which has also increased the accuracy and speed requirements for waveform picking. Early manual picking was difficult to meet the requirements in terms of speed. Therefore, the automatic waveform picking method has been developed and gradually become the mainstream in existing technologies. The automatic picking method is to extract and quantify the waveform features, and then determine the signal selection by setting a threshold. The features used in this process include time domain features, frequency domain features, and comprehensive features. Common ones include: methods based on amplitude and energy ratio, such as the STA / LTA (Short-Term Average / Long-TermAverage) algorithm, which detects earthquakes by calculating whether the average energy ratio in the long and short windows of a fixed length reaches the manually set threshold. When the earthquake signal arrives, the average energy ratio in the long and short windows will change suddenly. When this ratio is greater than the manually set threshold, it is automatically determined as an earthquake signal. This method is affected by factors such as the length of long and short time windows, trigger threshold setting, and feature function selection, and has low detection sensitivity for events with low signal-to-noise ratio. Based on waveform correlation, there is a "one-to-many" template matching method based on cross-correlation, which detects earthquakes by comparing the template earthquake event waveform with the station record cross-correlation. Due to the limitation of template selection, the applicability of template matching is low, and with the increase of template database, the calculation amount of detection also increases rapidly, and its calculation efficiency and memory efficiency are low. There is also a "many-to-many" matched filtering method based on autocorrelation, which has higher waveform cross-correlation sensitivity and can also detect unknown sources with similar waveforms, but because of its large calculation amount, it is still very time-consuming to detect earthquakes in large continuous data sets. Despite the above efforts, the accuracy of the automatic phase selection algorithm still lags behind the results of experienced manual picking. This is because the seismic waveform is highly complex due to multiple effects, including focal mechanism, stress drop, scattering, site effect, phase conversion, and interference from multiple noise sources. Traditional automatic phase selection algorithms use manually defined features and require careful data processing, such as bandpass filtering and setting activation thresholds. The threshold was set manually in the early days, but in subsequent work it gradually relied on the learning of machine learning methods. Summary of the invention
[0006] The present invention aims to solve the problems that manual seismic phase picking is time-consuming, inefficient and has large errors, and that the automatic picking algorithm depends on the signal-to-noise ratio of the data and has a large amount of calculation, making it difficult to cope with the rapid increase in the number of earthquakes. The present invention provides a method and device for automatic seismic phase picking based on deep learning, which can efficiently and accurately pick up rapid seismic phases in the reservoir area, effectively assist in the identification and analysis of reservoir-induced earthquakes, and further evaluate the earthquake risks that may be caused by reservoirs built in complex geological areas in the southwest. The present invention not only provides a scientific basis for the safe operation of reservoirs, but also enriches the database of earthquake research in the region, and provides an important reference for earthquake early warning and response in the Baihetan Reservoir and surrounding areas.
[0007] The above-mentioned object of the present invention is achieved by the following technical solutions:
[0008] A method for automatically picking up seismic phases based on deep learning, comprising the following steps:
[0009] Step 1: Collect three-component seismic waveform data including P-wave arrival time and S-wave arrival time;
[0010] Step 2: Preprocess the three-component seismic waveform data as samples to construct a training set, a validation set, and a test set;
[0011] Step 3: Convert the P-wave arrival time and S-wave arrival time corresponding to each sample into a Gaussian distribution curve of the P-wave arrival time and a Gaussian distribution curve of the S-wave arrival time with the same number of time points as the sample, and construct a noise distribution curve;
[0012] Step 4: Construct PSNet neural network and loss function;
[0013] Step 5: Use the sample set and the validation set to train the PSNet neural network, and obtain the optimal weight and bias parameters of the PSNet neural network based on minimizing the loss function value;
[0014] Step 6: Preprocess the three-component seismic waveform data to be processed and input them into the trained PSNet neural network to generate the predicted probability of each category at each time point.
[0015] The preprocessing of the three-component seismic waveform data in step 2 as described above includes the following steps:
[0016] In the three-component seismic waveform data, the three-component seismic waveform window data including the P wave arrival time and the S wave arrival time are selected by setting the window length, and the position of the window intercepting the three-component seismic waveform data is adjusted to obtain different three-component seismic waveform window data;
[0017] Adjust the sampling frequency of each three-component seismic waveform window data to be consistent;
[0018] The data of each time point of the three-component window data of each seismic waveform are normalized by removing the mean and dividing by the standard deviation to obtain the corresponding samples.
[0019] As described above, in step 3, the amplitude ranges of the P-wave arrival time Gaussian distribution curve and the S-wave arrival time Gaussian distribution curve are both 0 to 1, and the value corresponding to the time point on the noise distribution curve = 1-the value of the corresponding time point of the P-wave arrival time Gaussian distribution curve-the value of the corresponding time point of the S-wave arrival time Gaussian distribution curve.
[0020] As mentioned above, the PSNet neural network includes four downsampling stages, four upsampling stages, and a softmax function.
[0021] As mentioned above, the loss function Cross_entropy of the PSNet neural network is based on the following formula:
[0022]
[0023] Among them, i is the category number, x is the time point, and p i (x) is the actual probability that the value of the sample at time point x is the i-th category, q i (x) is the predicted probability of the i-th category of the value at time point x.
[0024] As mentioned above, the predicted probability q i (x) is based on the following formula:
[0025]
[0026] Where i is 1, 2 and 3 respectively represent noise, P wave and S wave, z i (x) is the value of the i-th category at the time point x output by the last upsampling stage, z k (x) is the value of the kth category at the time point x output by the last upsampling stage, i and k are both category numbers,
[0027] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned method for automatically picking up seismic phases when executing the computer program.
[0028] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for automatically picking up seismic phases.
[0029] A computer program product comprises a computer program, which implements the steps of the above-mentioned method for automatically picking up seismic phases when executed by a processor.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] (1) Due to the combined deep learning for automatic seismic phase picking, the PSNet neural network based on U-Net proposed in the present invention can efficiently extract features from three-component seismic waveform data, extract multi-level features of input data, fuse low-level and high-level features through jump connections between multiple downsampling stages and upsampling stages, and gradually restore the spatial resolution of the feature map through the upsampling stage, and finally generate a precise prediction probability, which solves the problem that manual seismic phase picking is time-consuming, labor-intensive, inefficient and has large errors.
[0032] (2) Since the seismic phases of different three-component seismic waveform data are picked up by combining deep learning and search algorithms, the present invention can reasonably allocate resources, thereby improving the evaluation speed of hyperparameter combinations and finding the relatively optimal solution more quickly.
[0033] (3) Since the seismic phase picking of seismic waveform data with different signal-to-noise ratios is performed by joint deep learning, the present invention can make the differential three-component seismic waveform data of different instruments be used for the training of the PSNet neural network, which can effectively improve the robustness of the seismic phase picking of the three-component seismic waveform data with different signal-to-noise ratios on different instruments. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flow chart of the present invention;
[0035] Figure 2 The Gaussian distribution diagram of the PSNet neural network predicting the seismic phase P and S waves. Specifically, given a continuous earthquake waveform of about 30 seconds containing three components of P and S waves at the RJOB station in the BW network at 20:00 on August 24, 2009, the PSNet neural network is used to pick up the P and S waves. DETAILED DESCRIPTION
[0036] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below in conjunction with embodiments. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0037] A method for automatically picking up seismic phases based on deep learning, such as Figure 1 As shown, the following steps are included:
[0038] Step 1: Collect three-component seismic waveform data including P-wave arrival time and S-wave arrival time;
[0039] Step 1 specifically includes the following steps:
[0040] Collect and download three-component seismic waveform data with P-wave arrival time and S-wave arrival time. The three-component seismic waveform data contains a large number of manually picked P-wave arrival times and S-wave arrival times, which represent extremely rich labeled training data and are very suitable for deep learning applications. In this embodiment, three-component seismic waveform data with P-wave arrival time and S-wave arrival time are obtained through site stratified sampling.
[0041] Step 2: Preprocess the three-component seismic waveform data into a format that can be used for deep neural network training. The preprocessed three-component seismic waveform data is used as a sample, and a training set, a validation set, and a test set are constructed based on the sample. The training set and the validation set are used for training and validation of the PSNet neural network, while the test set is specifically used to evaluate the final performance and results of the PSNet neural network.
[0042] The preprocessing process in step 2 includes the following steps:
[0043] Step 2.1: First, in the three-component seismic waveform data, the three-component seismic waveform window data including the P-wave arrival time and the S-wave arrival time are selected by setting a window with a set time length. In this embodiment, the time length is set to 30 seconds. In order to ensure that the algorithm is not just learning a specific window scheme, the position of the window intercepting the three-component seismic waveform data is adjusted according to different seismic event data, and the corresponding P-wave arrival time and S-wave arrival time in the three-component seismic waveform window data will also be adjusted accordingly.
[0044] Step 2.2: Then, the sampling frequencies of the three-component seismic waveform window data are adjusted to be consistent. In this embodiment, the sampling frequencies of the three-component seismic waveform window data are 100 Hz, which is also the most common sampling rate in the three-component seismic waveform data. Therefore, each three-component seismic waveform window data with a window length of 30 seconds contains data at 3001 time points.
[0045] Step 2.3: For each time point of the three-component window data of each seismic waveform, the data is normalized by removing the mean and dividing by the standard deviation to obtain the corresponding samples.
[0046] Through the above preprocessing, the different three-component seismic waveform data of different instruments can be used for the training of the PSNet neural network, which can effectively improve the robustness of seismic phase picking for three-component seismic waveform data with different signal-to-noise ratios on different instruments.
[0047] Step 3: Convert the P-wave arrival time and S-wave arrival time corresponding to each sample into the Gaussian distribution curve of the P-wave arrival time and the Gaussian distribution curve of the S-wave arrival time with the same number of time points as the sample, construct the noise distribution curve, and further obtain the actual probability p of the i-th category of the value of the sample at the time point x.i (x);
[0048] The P-wave arrival time and S-wave arrival time manually selected in the sample may have certain uncertainties and may not be the true arrival time. For this reason, it is assumed that the manually selected P-wave arrival time and S-wave arrival time conform to the Gaussian distribution. In this distribution, the manually picked P-wave arrival time and S-wave arrival time have the highest probability, that is, the maximum value of the Gaussian distribution curve, and the value (probability) of the time point near the maximum value of the Gaussian distribution curve gradually decreases. In this embodiment, the standard deviation of the Gaussian distribution curve of the P-wave arrival time and the Gaussian distribution curve of the S-wave arrival time are both set to 0.1 seconds.
[0049] The P-wave arrival time and S-wave arrival time corresponding to each sample are respectively converted into a P-wave arrival time Gaussian distribution curve and an S-wave arrival time Gaussian distribution curve with the same number of time points as the sample, the amplitude ranges of the P-wave arrival time Gaussian distribution curve and the S-wave arrival time Gaussian distribution curve are both 0-1, the value corresponding to each time point on the P-wave arrival time Gaussian distribution curve is the P-wave arrival time probability, and the value corresponding to each time point on the S-wave arrival time Gaussian distribution curve P is the S-wave arrival time probability, and a noise distribution curve is constructed, the number of time points of the noise distribution curve is the same as that of the sample, the value corresponding to the time point on the noise distribution curve = 1-the value of the corresponding time point of the P-wave arrival time Gaussian distribution curve-the value of the corresponding time point of the S-wave arrival time Gaussian distribution curve, and the value corresponding to each time point on the noise distribution curve is the noise probability;
[0050] Step 4: Construct PSNet neural network and loss function;
[0051] The above step 4 includes the following steps:
[0052] The structure of the PSNet neural network is based on U-Net generation, which is a deep neural network method for biomedical image processing, which aims to extract attributes from images. In the setting of this method, the PSNet neural network is used to extract three types of attributes in the time series: P wave, S wave and noise. The input is a sample, and the output is the probability distribution of P wave, S wave and noise at each time point of the sample. In this embodiment, the input sample of the PSNet neural network contains data of 3001 time points (30 seconds in length, sampled at 100Hz), and the data of each time point of the PSNet neural network output sample belongs to the probability of P wave, S wave and noise. The input sample will be processed by four downsampling stages and four upsampling stages of the PSNet neural network. Jump connections are made between the corresponding downsampling stages and upsampling stages, and one-dimensional convolution and rectified linear unit (ReLU) activation functions are applied in each downsampling stage and each upsampling stage. The downsampling process aims to extract useful information from the original seismic data and reduce it to a smaller number of neurons, and the upsampling process expands and converts this information into the probability values of P wave, S wave and noise at each time point of the sample. The last layer is connected to a softmax function to generate the predicted probability of each category at each time point on the sample:
[0053]
[0054] Where i is 1, 2 and 3 respectively represent noise, P wave and S wave, q i (x) is the predicted probability that the value of the sample at time point x is the i-th category, z i (x) is the value of the i-th category at the time point x output by the last upsampling stage, z k (x) is the value of the kth category at time point x output by the last upsampling stage, and i and k are both category numbers.
[0055] The loss function Cross_entropy is based on the predicted probability q of the sample at time point x being the i-th category. i (x) and the value of the sample at time point x is the actual probability p of the i-th category i The cross entropy between (x):
[0056]
[0057] Step 5: Use the sample set and the validation set to train the PSNet neural network, and obtain the optimal weight and bias parameters of the PSNet neural network based on minimizing the loss function value.
[0058] In deep learning, hyperparameters are different from the weights and biases of the model. They are set before training the model, rather than learned through the training process. The choice of hyperparameters plays a key role in the performance and behavior of the model, so choosing the right hyperparameters is crucial to obtaining good model performance. These parameters usually include learning rate, batch size, number of iterations, etc., which need to be manually adjusted or optimized to determine the optimal value to improve the training effect and generalization ability of the model. In this embodiment, keras-tuner is used in combination with the Hyperband search algorithm to optimize the hyperparameters of the PSNet neural network during model training. The Hyperband algorithm is a technique for hyperparameter optimization, which combines stratified sampling and early stopping principles, and aims to accelerate the experimental process through dynamic resource allocation, so as to find a near-optimal hyperparameter combination under limited resources, and obtain the PSNet neural network with relatively optimal hyperparameters by traversing different hyperparameter combinations and training, named PSNet.
[0059] Each sample in the test set is input into the trained PSNet neural network, and the predicted probability of the category corresponding to the data at each time point output by the PSNet neural network is analyzed. The accuracy of its P-wave and S-wave arrival time is evaluated using certain accuracy indicators.
[0060] For the predicted probability of the category corresponding to the data at each time point output by the PSNet neural network, this method uses precision, recall and F1 score to evaluate the effect of PSNet seismic phase picking P and S waves, where:
[0061]
[0062] T P 、F P , T N 、F N They are the abbreviations of True Positive, False Positive, True Negative, and False Negative. Specifically: T P : refers to the number of samples that are actually positive and correctly predicted to be positive; F P : refers to the number of samples that are actually negative but are mistakenly predicted to be positive; T N : refers to the number of samples that are actually negative and correctly predicted to be negative; F N : Refers to the number of samples that are actually positive but are mistakenly predicted as negative.
[0063] In this embodiment, if the difference between the predicted value and the actual value is within 0.1 second, it is considered to be TP . The F1 score is a balance between precision and recall. For example, if an extremely high threshold is applied to the positive picks, only the best picks will be classified as positive, which will account for a very small fraction of the total true picks. In this case, we may get extremely high precision but very low recall. Conversely, setting the threshold very low will result in lower precision but higher recall. In both cases, it will result in a lower F1 score, so using F1 is a more accurate way to evaluate the performance of the algorithm.
[0064] Step 6: Preprocess the three-component seismic waveform data to be processed and input them into the trained PSNet neural network to generate the predicted probability of each category at each time point.
[0065] This application provides a method for automatically picking up earthquake phases based on deep learning. It uses deep learning methods to quickly pick up the P-wave arrival time and S-wave arrival time of earthquakes in the Baihetan Reservoir area, which effectively helps to identify and analyze reservoir-induced earthquakes, and further evaluates the earthquake risks that may be caused by reservoirs built in complex geological areas in the southwest. This technology not only provides a scientific basis for the safe operation of reservoirs, but also enriches the database of earthquake research in the region, and provides an important reference for earthquake early warning and response in the Baihetan Reservoir and surrounding areas.
[0066] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.
[0067] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0068] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0069] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0070] The above are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should be regarded as the protection scope of the present invention.
Claims
1. A method for automatically picking up seismic phases based on deep learning, characterized in that: The following steps are involved: Step 1: Collect three-component seismic waveform data including P-wave arrival time and S-wave arrival time; Step 2: Preprocess the three-component seismic waveform data as samples to construct a training set, a validation set, and a test set; Step 3: Convert the P-wave arrival time and S-wave arrival time corresponding to each sample into a Gaussian distribution curve of the P-wave arrival time and a Gaussian distribution curve of the S-wave arrival time with the same number of time points as the sample, and construct a noise distribution curve; Step 4: Construct PSNet neural network and loss function; Step 5: Use the sample set and the validation set to train the PSNet neural network, and obtain the optimal weight and bias parameters of the PSNet neural network based on minimizing the loss function value; Step 6: Preprocess the three-component seismic waveform data to be processed and input them into the trained PSNet neural network to generate the predicted probability of each category at each time point.
2. A method for automatically picking up seismic phases based on deep learning according to claim 1, characterized in that: The preprocessing of the three-component seismic waveform data in step 2 comprises the following steps: In the three-component seismic waveform data, the three-component seismic waveform window data including the P wave arrival time and the S wave arrival time are selected by setting the window length, and the position of the window intercepting the three-component seismic waveform data is adjusted to obtain different three-component seismic waveform window data; Adjust the sampling frequency of each three-component seismic waveform window data to be consistent; The data of each time point of the three-component window data of each seismic waveform are normalized by removing the mean and dividing by the standard deviation to obtain the corresponding samples.
3. The method for automatically picking up seismic phases based on deep learning according to claim 1, characterized in that: In step 3, the amplitude ranges of the P wave arrival time Gaussian distribution curve and the S wave arrival time Gaussian distribution curve are both 0 to 1, and the value corresponding to the time point on the noise distribution curve = 1-the value of the corresponding time point of the P wave arrival time Gaussian distribution curve-the value of the corresponding time point of the S wave arrival time Gaussian distribution curve.
4. The method for automatically picking up seismic phases based on deep learning according to claim 1, characterized in that: The PSNet neural network includes four downsampling stages, four upsampling stages, and a softmax function.
5. The method for automatically picking up seismic phases based on deep learning according to claim 1, characterized in that: The loss function Cross_entropy of the PSNet neural network is based on the following formula: Among them, i is the category number, x is the time point, and p i (x) is the actual probability that the value of the sample at time point x is the i-th category, q i (x) is the predicted probability of the i-th category of the value at time point x.
6. The method for automatically picking up seismic phases based on deep learning according to claim 5, characterized in that: The predicted probability q i (x) is based on the following formula: Where i is 1, 2 and 3 respectively represent noise, P wave and S wave, z i (x) is the value of the i-th category at the time point x output by the last upsampling stage, z k (x) is the value of the kth category at time point x output by the last upsampling stage, and i and k are both category numbers.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method for automatically picking up seismic phases according to any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for automatically picking up seismic phases according to any one of claims 1 to 6 are implemented.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for automatically picking up seismic phases described in any one of claims 1 to 6 are implemented.