Earthquake nucleation phase identification and model training method, system, terminal and medium

By using machine learning models trained with seismic simulation data of different roughness, seismic data features are processed to identify seismic phase moments, the problem of inaccurate identification of nuclear seismic phases caused by the small amount of natural seismic data in the prior art is solved, and higher recognition accuracy and efficiency are achieved.

CN119535594BActive Publication Date: 2025-05-06SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510108560.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The prior art has a small amount of natural seismic data with labels, so the trained prediction model cannot accurately identify the nuclear seismic phase.

Method used

By obtaining the seismic data to be measured, performing feature extraction, and using a machine learning model trained on seismic simulation data of different roughness to process data, determine the seismic phase time, thereby realizing the identification of seismic nuclear seismic phase.

Benefits of technology

The diversity and data volume of training samples are expanded, the generalization ability of the model is improved, and the nuclear seismic phase can be more accurately identified and improved the accuracy and efficiency of seismic data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119535594B_ABST
    Figure CN119535594B_ABST
Patent Text Reader

Abstract

The earthquake nucleation phase identification and model training method, system, terminal and medium provided by the present invention specifically relate to the field of data monitoring technology, and the scheme includes: obtaining the earthquake data to be measured; extracting features of the earthquake data to be measured to obtain several types of data features; processing the data features using a trained machine learning model to determine several phase moments, and the machine learning model is trained based on earthquake simulation data with different roughness; based on the phase moment, the earthquake nucleation phase identification result is obtained. The scheme uses earthquake simulation data with different roughness as training samples to train the model, which expands the diversity and data volume of the training samples, and is conducive to improving the generalization ability of the model; at the same time, because the machine learning model has the ability to process complex time series data and multi-target prediction, the trained model has good prediction ability, can process complex time series data and is conducive to improving the accuracy of nucleation phase identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data monitoring technology, and in particular to an earthquake nucleation phase identification and model training method, system, terminal and medium. Background Art

[0002] The nucleation phase is the initial stage of an earthquake, which refers to the tiny shock waves caused by stress accumulation in the earthquake source area. These shock waves usually appear as weaker vibrations than the subsequent rupture phase. Therefore, accurately identifying and extracting the nucleation phase is crucial for seismological research, earthquake early warning and post-earthquake analysis.

[0003] At present, although machine learning methods have been applied to the field of seismic data analysis, due to the relatively small reserves of existing earthquake research data, only a small amount of labeled data can be obtained. Therefore, the prediction model trained based on machine learning algorithms using a small amount of labeled natural earthquake data as a data set cannot accurately identify the nucleation phase. Summary of the invention

[0004] In view of the above-mentioned deficiencies in the prior art, the object of the present invention is to provide a method, system, terminal and medium for earthquake nucleation phase identification and model training, aiming to solve the problem in the prior art that the trained prediction model cannot accurately identify the nucleation phase due to the small amount of labeled natural earthquake data.

[0005] In order to achieve the above-mentioned object, the present invention provides a first aspect of a method for identifying earthquake nucleation phases, comprising obtaining earthquake data to be measured;

[0006] Extracting features from the seismic data to be measured to obtain several types of data features;

[0007] Processing the data features using a trained machine learning model to determine a number of seismic phase moments, wherein the trained machine learning model is trained based on earthquake simulation data with different roughness;

[0008] Based on the seismic phase moment, an earthquake nucleation phase identification result is obtained.

[0009] Optionally, the feature extraction is performed on the seismic data to be measured to obtain several types of data features, including:

[0010] Downsampling the seismic data to be measured to obtain key seismic data;

[0011] Based on a preset time window, the key seismic data is segmented to obtain key seismic data segments;

[0012] Feature extraction is performed on key seismic data segments corresponding to each time window to obtain several types of data features, which at least include statistical features and envelope features.

[0013] Optionally, the data features are processed using a trained machine learning model to determine a number of seismic phase moments, including:

[0014] Based on a preset sampling factor, a preset step size and a preset time window, the data features are processed using a trained machine learning model to determine a number of seismic phase moments, which include at least the start time of the nucleation phase, the end time of the nucleation phase and the time when the earthquake occurs.

[0015] A second aspect of the present invention provides a machine learning model training method, which is applied to the steps of implementing any of the above-mentioned earthquake nucleation phase identification methods, including:

[0016] Collecting earthquake simulation data of different roughness, wherein the earthquake simulation data includes a time data set and an amplitude data set;

[0017] Performing time synchronization and downsampling on the seismic simulation data to obtain seismic data samples;

[0018] Segmenting and extracting features from the seismic data samples according to a preset time window to obtain data feature samples within each time window;

[0019] Generate label data for each time point based on a fault point label file and the time data set, wherein the fault point label file is generated based on a file containing earthquake event data in the earthquake simulation data;

[0020] Construct a training set based on the data feature samples and corresponding label data in each time window;

[0021] Based on the machine learning algorithm, build the initial machine learning model;

[0022] The initial machine learning model is trained using the training set until a preset training termination condition is reached to obtain a trained machine learning model.

[0023] Optionally, segmenting and extracting features from the seismic data samples according to preset time windows to obtain data feature samples in each time window includes:

[0024] Segmenting the seismic data sample according to a preset time window to obtain a plurality of seismic data sample segments;

[0025] Feature extraction is performed on the amplitude data in the amplitude data set corresponding to each segment of the seismic data sample to obtain data feature samples in each time window.

[0026] Optionally, generating label data for each time point based on the fault point label file and the time data set includes:

[0027] Based on the fault point label file, determining a number of fault points corresponding to the earthquake event;

[0028] The time corresponding to each of the fault points is compared with the time corresponding to the time data in the time data set to generate label data for each time point.

[0029] Optionally, after generating the label data at each time point, the method further includes:

[0030] Downsampling the label data at each time point to obtain optimized label data;

[0031] The data feature samples corresponding to each time window are aligned with the optimized label data to update the label data.

[0032] A third aspect of the present invention provides an earthquake nucleation phase identification system, the system comprising:

[0033] A data acquisition module, used to obtain the seismic data to be measured;

[0034] A feature extraction module is used to extract features from the seismic data to be measured to obtain several types of data features;

[0035] A seismic phase moment identification module, used to process the data features using a trained machine learning model to determine a number of seismic phase moments, wherein the trained machine learning model is trained based on earthquake simulation data of different roughnesses;

[0036] The identification result prediction module is used to obtain the earthquake nucleation phase identification result based on the seismic phase moment.

[0037] A fourth aspect of the present invention provides a terminal, which includes a memory for storing executable instructions; a processor for calling and running the executable instructions in the memory to execute any step of the above-mentioned earthquake nucleation phase identification method, and / or execute any step of the above-mentioned machine learning model training method.

[0038] A fifth aspect of the present invention provides a computer-readable storage medium, which stores program instructions. When the program instructions are executed by a processor, the steps of any one of the above-mentioned earthquake nucleation phase identification methods are implemented, and / or the steps of any one of the above-mentioned machine learning model training methods are executed.

[0039] Compared with the prior art, the beneficial effects of this solution are as follows:

[0040] The present invention uses earthquake simulation data of different roughness as training samples to train the model, which not only expands the diversity and data volume of training samples, but also ensures that the model can learn the common characteristics and differences of the two types of data, which is beneficial to improving the generalization ability of the model; at the same time, since the machine learning model has the ability to process complex time series data and multi-target prediction, the trained model has good prediction ability, can process complex time series data and is beneficial to improving the accuracy of nucleation phase identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0042] Figure 1 A brief flow chart of the earthquake nucleation phase identification method of the present invention;

[0043] Figure 2 It is a flowchart of the overall process of the earthquake nucleation phase identification method of the present invention;

[0044] Figure 3 It is a comparison curve diagram of the actual value and the predicted value of the normalized nucleation start time of the present invention;

[0045] Figure 4 It is a comparison curve diagram of the actual value and the predicted value of the normalized nucleation end time of the present invention;

[0046] Figure 5 A comparison curve diagram of the actual value and the predicted value at the time of earthquake occurrence of the present invention;

[0047] Figure 6 A curve diagram showing the change of friction coefficient with time before an earthquake occurs according to the present invention;

[0048] Figure 7 It is a schematic diagram of the change of acoustic emission signals monitored by multiple acoustic emission sensors at the same time over time;

[0049] Figure 8 It is a schematic diagram of the modules of the earthquake nucleation phase identification system of the present invention;

[0050] Fig. 9 It is a schematic diagram of the terminal structure of the present invention. DETAILED DESCRIPTION

[0051] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present invention.

[0052] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0053] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0054] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0055] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0056] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0058] The present invention faces the problem that the trained prediction model cannot accurately identify the nucleation phase due to the small amount of labeled natural earthquake data in the prior art, and proposes a method for identifying the nucleation phase of earthquake, which mainly includes obtaining laboratory earthquake data, especially acoustic emission data, extracting multiple features after processing; using machine learning algorithms, especially random forest regression models, to train and identify the start time, end time and earthquake occurrence time of the earthquake nucleation phase; further migrating to natural earthquakes for testing, and generating results for identifying and predicting the nucleation phase of natural earthquakes. The present invention can break through the limitations of traditional methods in the identification of earthquake nucleation phases, especially under complex environments and noise interference conditions, and can improve the accuracy and efficiency of identification by comprehensively analyzing laboratory data and natural earthquake data, thereby more comprehensively characterizing and depicting the earthquake nucleation process, and providing accurate technical support for earthquake early warning.

[0059] The embodiment of the present invention provides a method for identifying earthquake nucleation phases, which is deployed on electronic devices such as computers and servers and is applied to scenarios for identifying earthquake nucleation phases. Specifically, Figure 1 and Figure 2 As shown, the steps of the method in this embodiment include:

[0060] Step S100: Acquire the earthquake data to be measured.

[0061] Specifically, natural earthquake data to be tested are collected from real earthquake events. If the format of the natural earthquake data is not a preset data format, the natural earthquake data to be tested is converted into a preset data format through format conversion, and the data after format conversion are spliced ​​in chronological order to obtain the earthquake data to be tested. For example, if the natural earthquake data is stored in a standard data format SAC (Seismic Analysis Code, a time series signal processing software, which is one of the most widely used data analysis software packages in the field of natural seismology), and the preset data format is the standard format MAT for data storage in Matlab software or the file format HDF5 (Hierarchical Data Format 5, hierarchical data format version 5) for storing and managing large amounts of data, the natural earthquake data to be tested is format converted, such as using earthquake moment analysis and drawing software to convert the data in SAC format into MAT format or HDF5 format. Since natural earthquake data has high noise and uncontrollable factors, white noise is used to fill the missing parts during splicing, and the filling standard deviation is the average value of the background noise before and after the splicing data.

[0062] Step S200: extracting features from the seismic data to be measured to obtain several types of data features.

[0063] Specifically, feature extraction is performed on the seismic signal in the seismic data to be measured to obtain data features including one or more of the following types, specifically:

[0064] Extracting time domain features from the earthquake data to be tested, such as extracting the maximum and minimum amplitude values ​​in the earthquake waveform used to reflect the magnitude of the earthquake, that is, the maximum and minimum values ​​of the amplitude peak; extracting the amplitude used to measure the overall energy of the earthquake signal; extracting the duration of the earthquake signal; extracting the arrival time of seismic waves such as longitudinal waves (i.e. P waves, Primary Wave, earthquake longitudinal waves) and transverse waves (i.e. S waves, Secondary Wave, earthquake transverse waves), which helps to locate the earthquake source; extracting the envelope characteristics of the signal through Hilbert transform can effectively capture the instantaneous amplitude changes of the seismic phase, enhance the model's perception of complex signals, and thus help improve the robustness of the model.

[0065] Extract frequency domain features from the seismic data to be tested, such as extracting the main frequency components in the seismic data to be tested that are used to reflect the physical characteristics of seismic waves; extracting the distribution of seismic signals in the frequency domain, which helps to distinguish different types of seismic events; extracting the energy distribution in different frequency bands, that is, spectral energy, to provide frequency characteristics of seismic signals; extracting the frequency centroid of seismic signals to reflect the main frequency components of seismic waves.

[0066] Extract statistical features from the seismic data to be tested, such as calculating the arithmetic mean and root mean square of the seismic signal; calculating the variance used to measure the degree of fluctuation of the seismic signal; calculating the skewness used to describe the symmetry of the distribution shape of the seismic signal; calculating the kurtosis used to describe the sharpness of the distribution shape of the seismic signal.

[0067] For data with large high-frequency noise, wavelet analysis features are extracted from the seismic data to be tested, such as wavelet coefficients calculated through wavelet transform to reflect the local characteristics of seismic signals at different scales; extracting the distribution of wavelet energy at different scales is helpful to identify the detailed characteristics of seismic signals. Among them, wavelet transform can capture the local characteristics of the signal in both the time and frequency domains and reduce the impact of high-frequency noise on seismic phase identification.

[0068] Extract high-order statistical features from the seismic data to be measured, such as calculating the high-order moments of the seismic data to be measured, such as the third-order moment (i.e. skewness) and the fourth-order moment (i.e. kurtosis), to provide non-Gaussian information of the seismic signal; calculate the nonlinear characteristics of the seismic signal that are used to suppress Gaussian noise, i.e. high-order cumulants.

[0069] Extract phase features from the seismic data to be tested, such as extracting instantaneous phase information from the seismic data to be tested, reflecting the propagation direction and speed changes of seismic waves; calculating the phase differences between different seismic waves, which helps to locate the earthquake source and determine the type of earthquake; calculating the coherence between different seismic signals to determine the correlation of seismic events; determining the fractal dimension of seismic signals and other features to reflect the complexity and irregularity of seismic waveforms.

[0070] In practical applications, the method of feature extraction may vary depending on the type, quality and analysis purpose of seismic data. Therefore, when selecting a feature extraction method, it is necessary to comprehensively consider the characteristics of seismic data and the analysis requirements. At the same time, machine learning algorithms can also be combined for feature selection and dimensionality reduction to improve the accuracy and efficiency of seismic data analysis.

[0071] All data features are constructed into a feature matrix and saved as a binary format file for storing array data in the NumPy library in Python, i.e., the .npy file format, for subsequent analysis. The various data features extracted in this embodiment can fully reflect the changes in seismic signals and can help the model to perform more accurate nucleation phase identification.

[0072] Step S300: Use a trained machine learning model to process the data features to determine a number of seismic phase moments, and the trained machine learning model is trained based on fusion data of seismic simulation data with different roughness.

[0073] Specifically, taking into account factors such as the propagation path of seismic waves, medium characteristics, the location of observation points, and the performance of seismographs, one or more types of data features extracted above are used as inputs of the trained machine learning model to output several seismic phase moments. Among them, the seismic phase moment refers to the precise time when seismic waves (such as P waves, S waves, and surface waves, etc.) arrive at a specific observation point (such as a seismic station). These seismic waves include different types of seismic phases, and each seismic phase has its own unique propagation characteristics and arrival time characteristics. The trained machine learning model is obtained by training the initial machine learning model using seismic simulation data of different roughness to ensure that the model can learn the common features and differences of laboratory seismic data of different roughness. Among them, the initial machine learning model can be a random forest regression model, a support vector machine (SVM), a deep neural network (DNN), a transfer learning model, and all other machine learning models that can process data features.

[0074] Step S400: Based on the seismic phase moment, an earthquake nucleation phase identification result is obtained.

[0075] Specifically, by combining the phase moment and waveform characteristics, the duration, amplitude, frequency and other parameters of the nucleation phase are determined, thereby determining the start and end times of the nucleation phase, and obtaining the earthquake nucleation phase identification results to understand the energy release process of the nucleation phase and the earthquake preparation mechanism.

[0076] In this embodiment, this embodiment uses earthquake simulation data of different roughness as training samples to train the model, which not only expands the diversity and data volume of training samples, but also ensures that the model can learn the common characteristics and differences of the two types of data, which is beneficial to improving the generalization ability of the model; at the same time, since the machine learning model has the ability to process complex time series data and multi-target prediction, the trained machine learning model has good prediction ability, can process complex time series data and is beneficial to improving the accuracy of nucleation phase identification.

[0077] In one embodiment, the feature extraction of the seismic data to be measured in step S200 is performed to obtain several types of data features, including:

[0078] Step S210: down-sampling the seismic data to be measured to obtain key seismic data;

[0079] Step S220: segmenting the key seismic data based on a preset time window to obtain key seismic data segments;

[0080] Step S230: extracting features from the key seismic data segments corresponding to each time window to obtain several types of data features, wherein the data features include at least statistical features and envelope features.

[0081] Specifically, firstly, the time data in the seismic data to be measured is subjected to time synchronization processing to ensure the continuity of the time series, specifically including: if the seismic data to be measured has multiple time series or multi-dimensional time series (for example, data recorded by multiple seismic stations, or seismic waveform data in different directions recorded by the same seismic station), it means that the dimension of the time data in the seismic data to be measured is greater than 1, therefore, this embodiment selects the corresponding flattening strategy according to the analysis target. For example, if the goal is to identify the common seismic phase characteristics of all stations, the waveform data of all stations are spliced ​​into a long time series; if the goal is to analyze the waveform differences in different directions, the data in each direction are processed separately. Then, the seismic data to be measured is flattened according to the selected strategy to reorganize the data, such as converting a multi-dimensional array into a two-dimensional array (wherein a row represents a sampling point of a time series and a column represents a time point in the time series), or merging the data of multiple time series into a time series. If the seismic data to be measured comes from different stations or different directions, and the timestamps are not completely consistent, time alignment is performed by interpolation or resampling to ensure that all data points are at the same time point. During the flattening process, if missing data are encountered (for example, data from a station is not available for a certain period of time), the earthquake data to be measured are smoothed by interpolation or filling with zero values.

[0082] Furthermore, in order to improve the availability of the seismic data to be tested, the present embodiment uses an algorithm for calculating the discrete difference of elements in an array to calculate the time difference between adjacent time points in a time series, and determines the tolerance range of the time difference by calculating the average value and standard deviation of each time difference. If the correspondence between the time data and the amplitude data in the seismic data to be tested is unclear, and any time difference is detected to exceed the preset maximum allowable difference, a warning is issued to remind the user to check the seismic data to be tested.

[0083] Furthermore, in order to reduce the amount of data, improve computing efficiency, and optimize computing results, this embodiment downsamples the time data and amplitude data according to the scale of the data and the requirements of specific tasks to obtain key seismic data. Through downsampling, this embodiment can not only retain sufficient seismic phase information, but also reduce redundant information and the computational burden of subsequent processing, meet the demand for rapid response, and improve the accuracy of prediction to a certain extent, achieving a balance between real-time and accuracy, and providing the possibility for real-time processing of large-scale seismic data.

[0084] Then, based on a preset time window (for example, including 1000 to 10000 samples), the key seismic data is segmented to obtain key seismic data segments, ensuring that each time window contains sufficient information to avoid excessive occupation of computing resources by long time series data. Finally, in order to ensure the accurate calculation of local features, feature extraction is performed on the key seismic data segments corresponding to each time window to obtain several types of data features, which include at least statistical features and envelope features. In other preferred embodiments, a dynamic time window can be used according to the characteristics of different seismic events. The dynamic time window can automatically adjust the size of the window according to the complexity and change rate of the seismic signal. For areas with drastic changes in seismic phases, a smaller time window is used; for areas with gentle changes, a larger window is used to refine the feature extraction process and reduce redundant data.

[0085] As an illustrative example, the adjustment principles for the time window, step size and downsampling factor for different seismic data are as follows:

[0086] The principle of selecting the window size is related to the time interval of laboratory earthquake events in the meter-level rock friction experiment, which is generally 2 / 5 to 4 / 5 of the time interval between two adjacent laboratory earthquake events; the step size is 1 / 20 to 1 / 10 of the time window size; the sampling factor is 10 / 100 / 1000, etc. Adjustment principle: After each training, the size of the determination coefficient is adjusted according to the size of the determination coefficient, and the value is increased or decreased within the value range, and the time window, step size, and sampling factor value with a determination coefficient value closer to 1 are selected.

[0087] In one embodiment, the step S300 uses a trained machine learning model to process the data features to determine a number of seismic phase moments, including:

[0088] Based on a preset sampling factor, a preset step size and a preset time window, the data features are processed using a trained machine learning model to determine a number of seismic phase moments, which include at least the start time of the nucleation phase, the end time of the nucleation phase and the time when the earthquake occurs.

[0089] Specifically, first, according to the type of data features and the size of the data volume, the sampling factor is preset to adjust the data sampling rate, so as to reduce the data volume and computational complexity while maintaining the key information of the data. According to the scale of the data and the requirements of the specific task, the time window and step size are preset to divide the continuous seismic waveform data into multiple time periods, ensuring that each time period contains enough information to identify the seismic phase. Then, the data features in each time window are used as the input of the trained machine learning model, and the seismic phase and non-seismic phase are distinguished by setting the seismic phase threshold, and all time windows are traversed. When the output of the model exceeds the threshold, the seismic phase is considered to be detected. When the characteristics of the nucleation phase are detected for the first time, it is recorded as the start time of the nucleation phase; when the characteristics of the nucleation phase gradually weaken to below the threshold, it is recorded as the end time of the nucleation phase; the time of earthquake occurrence is determined according to a certain offset of the start time of the nucleation phase, or according to the arrival time of other seismic phases (such as P waves and S waves). It can be seen that the size of the time window, the step size and the downsampling factor in this embodiment can be flexibly adjusted according to factors such as the type of data features and the amount of data, so that the system can optimize the processing of seismic data of different scales and characteristics, and by adjusting these parameters, the model can adapt to different seismic phase events and data noise environments, which is conducive to maintaining good performance of the model in different data environments.

[0090] In one embodiment, the training process of the machine learning model in step S300 includes:

[0091] Step M100: collecting earthquake simulation data of different roughnesses, wherein the earthquake simulation data includes a time data set and an amplitude data set.

[0092] Specifically, earthquake simulation data of different roughnesses are collected, such as acoustic emission data of different roughnesses obtained from the laboratory, for simulating earthquake events, including a time data set and an amplitude data set. Among them, the unit of the time data set is seconds, and the amplitude data set is the corresponding earthquake signal intensity. These laboratory data are suitable as basic data for model training because they have less noise and stronger controllability. The storage format of the earthquake simulation data used in this embodiment is MAT or HDF5 format, and it can also be selected according to the requirements of data volume, storage duration, etc., to store in a column storage format (such as Parquet format or Apache Arrow), which is conducive to improving the efficiency of data storage and reading, especially in real-time processing applications of large-scale earthquake data.

[0093] Step M200: time-synchronize and downsample the seismic simulation data to obtain seismic data samples.

[0094] Specifically, firstly, the time data in the earthquake simulation data is synchronized to align the data from different sources on the same time axis to ensure the continuity of the time series, including: if the dimension of the time data in the earthquake simulation data is greater than 1, the corresponding flattening strategy is selected according to the analysis objective, and the earthquake simulation data is flattened according to the selected strategy to obtain earthquake data samples. During the flattening process, if missing data is encountered (for example, the data of a station is not available in a certain time period), the earthquake simulation data is smoothed by interpolation or filling with zero values.

[0095] Furthermore, in order to improve the availability of data, this embodiment uses an algorithm for calculating the discrete difference of elements in an array to calculate the time difference between adjacent time points in the time series, and determines the tolerance range of the time difference by calculating the average and standard deviation of each time difference. If the correspondence between the time data and the amplitude data is unclear, and any time difference is detected to exceed the preset maximum allowable difference, a warning is issued to remind the user to check the data.

[0096] Furthermore, in order to reduce the amount of data, improve computing efficiency, and optimize computing results, this embodiment downsamples the time data and amplitude data according to the scale of the data and the requirements of the specific task to obtain key seismic data. Through downsampling, this embodiment can not only retain sufficient seismic phase information, but also reduce redundant information and the computational burden of subsequent processing, ensure the processing speed of the data, and improve the accuracy of the prediction to a certain extent, which makes it possible to process large-scale seismic data in real time. Furthermore, the sampling rate is dynamically adjusted according to the characteristics of the data. For example, a lower sampling rate is used in relatively stable areas of the seismic signal, and a higher sampling rate is used for processing in areas where the signal changes dramatically. As other implementations, the data set can also be optimized by preprocessing steps other than time synchronization and downsampling.

[0097] Step M300: Segment and extract features from the seismic data samples according to a preset time window to obtain data feature samples in each time window.

[0098] Specifically, the seismic data samples are segmented according to a preset time window to obtain a number of seismic data sample segments; the amplitude data in the amplitude data set corresponding to each seismic data sample segment are respectively feature extracted to obtain data feature samples in each time window, wherein the data feature samples include but are not limited to the various features obtained in step S200, and by dividing the sliding window into two parts in a preset ratio, calculating the average value of the preset percentiles in the two parts, and taking each average value as a percentile feature value, thereby obtaining a number of percentile features. For example, for each window of a certain seismic data sample, the average value corresponding to each percentile from the 1st percentile to the 9th percentile, and the average value corresponding to each percentile from the 91st percentile to the 99th percentile are respectively calculated, and finally 18 average values ​​are obtained, that is, 18 percentile feature values ​​are obtained.

[0099] Step M400: Generate label data for each time point based on the fault point label file and the time data set, wherein the fault point label file is generated based on a file containing earthquake event data in the earthquake simulation data.

[0100] Specifically, first, the fault points are read from multiple label files containing earthquake simulation data of earthquake events (such as label files of the start time of the nucleation phase, label files of the end time of the nucleation phase, and label files of the earthquake arrival time), and the time corresponding to the fault point is subtracted from the time data corresponding to different time in the time data set to generate the corresponding fault point label file, wherein each label file contains two columns, the first column is the number of the earthquake event, and the second column is the time of the fault point corresponding to the event. Then, based on the fault point label file, several fault points corresponding to the earthquake event are determined. For each label file, a zero array with the same size as the time data is created to store the time difference between each fault point and the time data in the time data set of the earthquake data sample. For each set of label files, a function for converting the data in the text format label file into the NumPy array format is used to read the data and extract the time value of the fault point. Finally, the time corresponding to each fault point is compared with the time corresponding to the time data in the time data set to generate the label data of each time point, which is used to supervise the training of the model and help the model learn the time characteristics of the earthquake phase. For example, by calculating the time corresponding to each fault point minus the time data that is less than its own time, and generating corresponding label data based on these differences, the calculated data features and label data are saved as NumPy array files for subsequent training.

[0101] For each fault point, the label of the time point is obtained by calculating the time difference between the time point and the current fault point. If it is the first fault point, all time differences earlier than the fault point in the time column are marked; for subsequent fault points, the time difference between the previous fault point and the current fault point is calculated, and all time points in the time column within this time period are marked.

[0102] For example, suppose there is a time series in the time data set of earthquake data samples, and the earthquake time contained in the label file has some fault point data. The process of constructing the label array is as follows:

[0103] Time series (time_data): [0,1,2,3,4,5,6,7,8] (unit: seconds);

[0104] Fault point data (fault_points): [(1,3),(2,7)] (format is (event number, fault point time)), indicating that the first fault point occurs at 3 seconds and the second fault point occurs at 7 seconds.

[0105] Step 1: First, initialize a label array for the time data. Assume that the size of this label array is the same as the time data and the initial values ​​are all zero.

[0106] Step 2: Calculate the label of the first fault point. Since the first fault point occurs at 3 seconds, calculate the time difference between all time points earlier than 3 seconds and 3 seconds. The calculation process is as follows:

[0107] For the time 0 second in time_data, the label value is 3-0=3 seconds;

[0108] For the time of 1 second in time_data, the label value is 3-1=2 seconds;

[0109] For the time 2 seconds in time_data, the label value is 3-2=1 second;

[0110] For the time 3 seconds in time_data, the label value is 3-3=0 seconds.

[0111] The label array constructed by obtaining the first set of label data based on the first fault point is: [3,2,1,0].

[0112] Calculate the label of the second fault point. Since the second fault point occurs at 7 seconds, calculate the time difference between all time points earlier than 7 seconds and 7 seconds. The calculation process is as follows:

[0113] For the time 3 seconds in time_data, the label value is 7-3=4 seconds;

[0114] For the time 4 seconds in time_data, the label value is 7-4=3 seconds;

[0115] For the time 5 seconds in time_data, the label value is 7-5=2 seconds;

[0116] For the time 6 seconds in time_data, the label value is 7-6=1 second;

[0117] For the time 7 seconds in time_data, the label value is 7-7=0 seconds.

[0118] The second set of label data is obtained according to the second fault point to construct a label array: [4,3,2,1,0]. Finally, all the label data are used to construct a complete label array: [3,2,1,0,4,3,2,1,0]. The label data includes the start time, end time and earthquake occurrence time of the seismic phase. Finally, all the generated label data are saved as NumPy array files for subsequent use.

[0119] In one implementation, after generating the label data at each time point in step M400, step M800 is included, specifically:

[0120] Step M810: down-sampling the label data at each time point to obtain optimized label data;

[0121] Step M820: aligning the data feature samples corresponding to each time window with the optimized label data, and updating the label data.

[0122] Specifically, in order to align the label with the feature data, this embodiment performs downsampling processing on the label data at each time point to obtain optimized label data. By setting the window size and step size, the data feature samples corresponding to each time window are aligned with the optimized label data, such as retaining the last label value of each window to update the label array, ensuring the synchronization of data features and label data, which is convenient for subsequent model training.

[0123] Step M500: construct a training set based on the data feature samples and corresponding label data in each time window.

[0124] Specifically, the data feature samples in each time window after time alignment and downsampling are paired with the label data at the corresponding position in the label array to construct a training set.

[0125] Step M600: Build an initial machine learning model based on the machine learning algorithm.

[0126] Specifically, based on machine learning algorithms such as random forest regression model, support vector machine, deep neural network, and transfer learning model, an initial machine learning model is constructed to enable the model to process complex time series data and improve the recognition accuracy of nucleation phases.

[0127] It is easy to understand that this embodiment builds a model based on a machine learning algorithm. As another preferred implementation mode, the model can also be built based on a deep learning model (such as a convolutional neural network or a recurrent neural network, etc.). The deep learning model is particularly suitable for processing high-dimensional data, especially time series data. It can automatically learn complex feature expressions from the data and can show good performance in seismic phase identification.

[0128] Step M700: Use the training set to train the initial machine learning model until a preset training termination condition is reached to obtain a trained machine learning model.

[0129] Specifically, the training is performed by inputting the training set into the initial machine learning model. For example, for the initial machine learning model built based on the random forest regression model, the model builds multiple decision trees through multiple trainings, and randomly extracts multiple sample subsets (with replacement sampling) from the original data set to provide different training data for each tree. Training is performed on different data subsets, and the tree depth, splitting criteria, minimum number of sample splits, etc. are optimized. In the splitting process of each node, a part of the features is randomly selected instead of considering all the features to increase the differences between the trees. A decision tree is constructed on each subset until the preset training termination conditions (such as the depth of the tree, the number of samples of the leaf node, etc.) are reached, and finally a trained machine learning model that can accurately identify earthquake phases is obtained.

[0130] After training, the performance of the model is evaluated through the cross-validation method of time series data to ensure its generalization ability to the test data and avoid overfitting. At the same time, the evaluation indicators include mean square error and coefficient of determination to measure the accuracy and predictive ability of the model. The trained machine learning model can be used to identify new earthquake data in real time. By extracting and analyzing the characteristics of the nucleation phase, the normalized start time, end time and earthquake occurrence time of the nucleation phase are output to achieve the identification and extraction of the nucleation phase, thereby obtaining the earthquake prediction trend.

[0131] Then, the trained model is tested using natural earthquake data, including:

[0132] Natural earthquake data stored in SAC format are collected from real earthquake events, and the format of natural earthquake data is converted and spliced ​​in chronological order to obtain natural earthquake data, which also includes time data set and amplitude data set. Since natural earthquake data has high noise and uncontrollable factors, white noise is used to fill the missing parts during splicing, and the filling standard deviation is the average value of the background noise before and after the splicing data.

[0133] According to the processing method of step M200, the time data in the natural earthquake data is time-synchronized to align data from different sources on the same time axis to ensure the continuity of the time series. Since natural earthquake data often only contains one earthquake event, the label data can be directly constructed using the earthquake time, and the time difference between each time point and the earthquake time can be calculated, which can simplify the label data generation process and is particularly suitable for rapid testing of natural earthquake data. Then, according to the processing method of step M800, the label data of each time point in the natural earthquake data is downsampled to obtain optimized label data, and the data feature samples corresponding to each time window are aligned with the optimized label data to update the label data.

[0134] During the model testing process, the test window, step size, and downsampling factor are set using labeled data generated based on natural earthquake data, including: Test window size: The recommended range is 5000 to 20000, and multiples of 100 are selected for machine learning. The specific adjustment can be made according to the scale of the test data and the proportion of the window when training the model; Test step size: It is recommended to be a multiple of 100, and the value range is between one-fifth and one-twentieth of the training window size; Downsampling factor: Select according to different data volumes, the sampling factor is 10 / 100 / 1000, etc.; The adjustment principle is the same as the adjustment of training parameters.

[0135] Furthermore, in order to improve the predictive performance of the model, transfer learning or multi-task learning can be introduced so that the trained model can be applied to different types of earthquake events, or jointly trained for multiple earthquake events, which is conducive to further improving the diversity of training data and the generalization ability of the model.

[0136] In this embodiment, the model is jointly trained by utilizing high-quality complex features of laboratory data of different roughness, and the model is migrated to the application of natural earthquake data through transfer learning methods, thereby greatly improving the accuracy and adaptability of the model.

[0137] In summary, the beneficial effects of the method of the present invention mainly include:

[0138] The present invention constructs a training set by fusing earthquake simulation data of different roughness, and constructs an initial machine learning model based on a machine learning algorithm, which can accurately identify the start time, end time and earthquake occurrence time of the nucleation phase, and improve the accuracy and generalization ability of phase identification. Since the earthquake simulation data generated in the laboratory has a large amount of data, low noise and strong controllability, it is suitable for extracting clear phase characteristics; natural earthquake data is obtained through an earthquake monitoring network, has high noise and complex phases, but a small amount of data. Combining earthquake simulation data of different roughness can fully reflect the characteristics of the nucleation phase in multiple dimensions, which is conducive to improving the recognition ability of the model. The random forest regression model is used to automatically learn the potential laws in the data. Compared with the traditional rule-based method, it can more accurately identify complex phase patterns, especially in the case of noise interference, and can effectively improve the recognition accuracy.

[0139] The present invention can process complex time series data, especially when the data noise is large, it can ensure the accuracy of seismic phase identification, breaking through the limitations of traditional methods. By calculating a variety of signal features, it can comprehensively reflect the various changing trends of seismic signals. For example, by converting the signal into a complex form through Hilbert transform, the envelope is obtained, and then the envelope characteristics of the signal are extracted, which can effectively capture the instantaneous amplitude changes of the seismic phase and reflect the fluctuation of the signal. In particular, it has obvious advantages in the identification of weak signals such as nucleation phases. The envelope characteristics can also enhance the model's perception of complex signals and further improve the robustness of the model.

[0140] The present invention divides the seismic signal into time windows, and the amount of data in each time window is moderate, thereby avoiding excessive occupation of computing resources by long time series data. At the same time, the feature extraction method in the time window ensures the accurate calculation of local features, thereby realizing efficient data processing and model training. By downsampling the data, the redundant information in the original data is reduced and the processing efficiency is improved. The downsampled data not only retains sufficient seismic phase information, but also reduces the amount of calculation required for model training, making it possible to process large-scale seismic data in real time.

[0141] The present invention adopts a strategy of flexibly adjusting the time window size and sampling factor, so that the model can automatically adjust the time scale for different data types (such as long-period earthquakes, short-period vibrations, etc.) and different noise environments, and select the most appropriate parameters for phase identification, so that the system can optimize the processing of seismic data of different scales and characteristics. By adjusting these parameters, the model can adapt to different seismic phase events and data noise environments. Since the present invention makes full use of the advantages of laboratory seismic data of different roughness, which is less noise and rich and diverse, it can adapt to various types of seismic data to ensure the best recognition accuracy and calculation efficiency, and avoid overfitting or underfitting of the model in certain data situations.

[0142] The present invention introduces a technical method for multi-label generation, which not only generates labels for the start time of the nucleation phase, but also generates labels for the end time and the time when the earthquake occurs. In addition, after the label is generated, the label is downsampled and aligned to ensure accurate alignment between the label and the feature data, so that the model can learn multiple phase moments (such as the start time of the phase, the end time of the phase, and the time when the phase occurs) at the same time during the training process, effectively improving the model's ability to recognize the entire process of the phase.

[0143] The present invention can detect the nucleation phase in the early stage of the earthquake source by accurately identifying and extracting the nucleation phase, which helps to improve the response speed and accuracy of the earthquake early warning system.

[0144] It should be stated that the present invention is not only applicable to laboratory data and natural earthquake data, but can also be extended to other fields, such as acoustic emission data analysis, artificial earthquake induced event monitoring, etc. The technical framework of the present invention can be extended to other types of time series data analysis tasks. For example, acoustic emission data analysis can be used to monitor the rupture process of underground structures, or used in fields such as equipment monitoring in industrial production. The core technical features of the present invention, such as time windows, feature extraction methods, label generation, and machine learning model training, all have strong cross-domain applicability. By adjusting relevant parameters and algorithms, the concept of the method of the present invention can be applied to different types of time series data and a variety of monitoring scenarios. Therefore, all model training and prediction functions implemented based on the concept of the present invention are within the scope of protection of the present invention.

[0145] In order to verify the effectiveness of the method of the present invention, the following training of seismic phase identification is performed based on the random forest regression model, and the key parameters are set as follows: the number of trees (n_estimators) is set to 776; the maximum depth (max_depth) is set to 250 to limit the maximum depth of each tree; the minimum number of sample splits (min_samples_split) is set to 54, which is the minimum number of samples required when the decision tree splits. If the number of samples at a node is less than this value, the tree will not split further; the minimum number of sample leaves (min_samples_leaf) is set to 23, which indicates the minimum number of samples required for each leaf node to control the growth of the tree and avoid leaf nodes being too sparse; the bootstrap method (bootstrap) is set to True, indicating that the bootstrap method (with replacement sampling) is used when constructing each tree to improve the stability and accuracy of the model.

[0146] Based on the above parameter settings, we get Figure 3 The comparison curve of the true value and the predicted value of the normalized nucleation start time is shown in the figure. It can be seen that the gradual decrease of the true value before the start of the nucleation phase represents the gradual approach to the specific time of the nucleation start. The predicted value of the remaining time of the nucleation start calculated by substituting the amplitude data of the test data into the trained random forest model shows a downward trend, indicating that the normalized remaining time of the nucleation start gradually decreases with the passage of time. In the initial stage (approximately 0 to 40,000 seconds), the gap between the predicted value and the true value is large, indicating that the model cannot fit the real data well in this time period. As time goes by (especially between 40,000 seconds and 80,000 seconds), the gap between the predicted value and the true value gradually decreases, indicating that the prediction performance of the model in the later period is improved. This shows that the model can capture some potential dynamic changes or influencing factors near the nucleation, making its prediction ability stronger. Overall, the model can provide reasonable predictions in certain time periods, but the prediction accuracy may be insufficient in a longer time range.

[0147] Get as Figure 4The comparison curve of the true value and predicted value of the normalized nucleation end time is shown in the figure. It can be seen that the gradual decrease of the true value before the end of the nucleation phase represents the specific time of gradually approaching the end of nucleation. The predicted value of the remaining time of nucleation end calculated by substituting the amplitude data of the test data into the trained random forest model shows a downward trend, indicating that the normalized remaining time of nucleation end gradually decreases with the passage of time. In the initial stage (approximately 0 to 50,000 seconds), the gap between the predicted value and the true value is large, indicating that the model cannot fit the real data well in this time period. As time goes by (especially between 50,000 seconds and 70,000 seconds), the gap between the predicted value and the true value gradually decreases, indicating that the prediction performance of the model in the later stage is improved. Compared with the model of nucleation start, the signal time captured by the nucleation end model is later, which is consistent with the actual situation that the nucleation end is distributed in the later position and the front remaining time gradually decreases.

[0148] Further, we get Figure 5 The comparison curve of the real value and the predicted value of the earthquake occurrence time shown in the figure shows that the real value of the remaining time before the earthquake gradually decreases, which means that it is gradually approaching the specific time of the earthquake. The predicted value of the remaining time of the earthquake calculated by substituting the amplitude data of the test data into the trained random forest model shows a downward trend, indicating that the normalized remaining time of the earthquake gradually decreases with the passage of time. Similar to the model test results of the end of nucleation, in the initial stage (approximately 0 to 50,000 seconds), it is not able to fit the real data well. As time goes by (especially between 50,000 seconds and 70,000 seconds), the gap between the predicted value and the real value gradually decreases. This is largely because in some earthquake events used for training, the moment of the end of nucleation and the moment of earthquake occurrence overlap with each other, which is in line with expectations. The final real remaining time can be proportionally restored to the normalized vertical axis by the time interval between the prediction curve and the time when the prediction signal fluctuates, so as to obtain the real time by denormalization.

[0149] Figure 6 The figure shows the curve of the friction coefficient changing with time before the earthquake event, and includes the intersection points of the curve of the friction coefficient changing with time with the start of nucleation, the end of nucleation and the time of earthquake occurrence. Figure 7 The figure shows the change of friction coefficient before an earthquake event and the seismic signals received by different acoustic emission channels. Figure 7The gray line in the figure is a schematic diagram of the change of the acoustic emission signal over time at the same time monitored by 32 acoustic emission sensors. The yellow line represents the time point when the macro friction shows a downward trend, the green line represents the time point when the friction coefficient rebounds slightly before a sharp drop, and the red line represents the time point when the macro friction drops sharply. The label file of the start time of the nucleation phase selects the first time point when the macro friction shows a downward trend before each earthquake event (the time corresponding to the yellow line represents the time when nucleation starts). The label file of the earthquake arrival time selects the time point when the macro friction drops sharply at the time of each earthquake (the time corresponding to the red line represents the time when the earthquake occurs). The label file of the end time of the nucleation phase selects the time point when the macro friction rebounds slightly before each earthquake occurs (the time corresponding to the green line represents the end time of nucleation. The end time of nucleation of many events overlaps with the time of earthquake occurrence, and no rebound phenomenon occurs).

[0150] like Figure 8 As shown, corresponding to the above-mentioned earthquake nucleation phase identification method, an embodiment of the present invention further provides an earthquake nucleation phase identification system, and the above-mentioned earthquake nucleation phase identification system includes:

[0151] The data acquisition module 810 is used to obtain the seismic data to be measured;

[0152] A feature extraction module 820 is used to extract features from the seismic data to be measured to obtain several types of data features;

[0153] A seismic phase moment identification module 830 is used to process the data features using a trained machine learning model to determine a number of seismic phase moments, wherein the trained machine learning model is trained based on earthquake simulation data of different roughnesses;

[0154] The identification result prediction module 840 is used to obtain the earthquake nucleation phase identification result based on the seismic phase moment.

[0155] Specifically, in this embodiment, the specific functions of the above-mentioned earthquake nucleation phase identification system can also refer to the corresponding description in the above-mentioned earthquake nucleation phase identification method, which will not be repeated here.

[0156] Based on the above embodiment, the present invention further provides a terminal, whose principle block diagram can be as follows: Fig. 9 As shown. The terminal can be used to execute the earthquake nucleation phase identification method provided in the above embodiment, which will not be described here for the sake of brevity. The terminal includes: a processor, the processor is coupled to a memory, the memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions stored in the memory, so that the method in the above method embodiment is executed.

[0157] The present invention also provides a computer-readable storage medium on which computer instructions for implementing the method in the above method embodiment are stored.

[0158] For example, when the computer program is executed by a computer, the computer can implement the method in the above method embodiment.

[0159] The embodiment of the present application also provides a computer program product including instructions, which, when executed by a computer, enables the computer to implement the method in the above method embodiment.

[0160] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0161] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0162] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0163] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0164] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0165] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

Claims

1. A method for identifying earthquake nucleation phases, characterized in that: The following steps are involved: Acquire the earthquake data to be measured; Extracting features from the seismic data to be measured to obtain several types of data features; Processing the data features using a trained machine learning model to determine a number of seismic phase moments, wherein the trained machine learning model is trained based on earthquake simulation data with different roughness; Based on the seismic phase moment, obtaining an earthquake nucleation phase identification result; The method of obtaining the seismic data to be measured comprises: collecting the natural seismic data to be measured from a real seismic event; if the format of the natural seismic data to be measured is not a preset data format, converting the natural seismic data to be measured into a preset data format through format conversion; and splicing the format-converted data in chronological order to obtain the seismic data to be measured; and filling the missing part with white noise during splicing, and the filling standard deviation is the average value of the background noise before and after the splicing data; The method of processing the data features using the trained machine learning model to determine a number of seismic phase moments includes: based on a preset sampling factor, a preset step size, and a preset time window, processing the data features using the trained machine learning model to determine a number of seismic phase moments, wherein the seismic phase moments at least include a start time of a nucleation phase, an end time of a nucleation phase, and a time when an earthquake occurs; The machine learning model training method includes: collecting earthquake simulation data of different roughness, the earthquake simulation data including a time data set and an amplitude data set; performing time synchronization and downsampling on the earthquake simulation data to obtain earthquake data samples; performing segmentation and feature extraction on the earthquake data samples according to a preset time window to obtain data feature samples in each time window; generating label data for each time point based on a fault point label file and the time data set, the fault point label file being generated based on a file containing earthquake event data in the earthquake simulation data; constructing a training set based on the data feature samples in each time window and the corresponding label data; constructing an initial machine learning model based on a machine learning algorithm; and training the initial machine learning model using the training set until a preset training termination condition is reached to obtain a trained machine learning model.

2. The earthquake nucleation phase identification method according to claim 1, characterized in that: The feature extraction of the seismic data to be measured is performed to obtain several types of data features, including: Downsampling the seismic data to be measured to obtain key seismic data; Based on a preset time window, the key seismic data is segmented to obtain key seismic data segments; Feature extraction is performed on key seismic data segments corresponding to each time window to obtain several types of data features, which at least include statistical features and envelope features.

3. A machine learning model training method, characterized in that: The steps used to implement the earthquake nucleation phase identification method according to any one of claims 1 to 2 include: Collecting earthquake simulation data of different roughness, wherein the earthquake simulation data includes a time data set and an amplitude data set; Performing time synchronization and downsampling on the seismic simulation data to obtain seismic data samples; Segmenting and extracting features from the seismic data samples according to a preset time window to obtain data feature samples within each time window; Generate label data for each time point based on a fault point label file and the time data set, wherein the fault point label file is generated based on a file containing earthquake event data in the earthquake simulation data; Construct a training set based on the data feature samples and corresponding label data in each time window; Based on the machine learning algorithm, build the initial machine learning model; The initial machine learning model is trained using the training set until a preset training termination condition is reached to obtain a trained machine learning model.

4. The machine learning model training method according to claim 3, characterized in that: The segmenting and feature extraction of the seismic data samples according to the preset time windows to obtain data feature samples in each time window includes: Segmenting the seismic data sample according to a preset time window to obtain a plurality of seismic data sample segments; Feature extraction is performed on the amplitude data in the amplitude data set corresponding to each segment of the seismic data sample to obtain data feature samples in each time window.

5. The machine learning model training method according to claim 3, characterized in that: The generating of label data at each time point based on the fault point label file and the time data set includes: Based on the fault point label file, determining a number of fault points corresponding to the earthquake event; The time corresponding to each of the fault points is compared with the time corresponding to the time data in the time data set to generate label data for each time point.

6. The machine learning model training method according to claim 5, characterized in that: After generating the label data at each time point, the following steps are included: Downsampling the label data at each time point to obtain optimized label data; The data feature samples corresponding to each time window are aligned with the optimized label data to update the label data.

7. An earthquake nucleation phase identification system, characterized in that: The system is applied to implement the steps of the earthquake nucleation phase identification method according to any one of claims 1 to 2, and comprises: A data acquisition module, used to obtain the seismic data to be measured; A feature extraction module is used to extract features from the seismic data to be measured to obtain several types of data features; A seismic phase moment identification module, used to process the data features using a trained machine learning model to determine a number of seismic phase moments, wherein the trained machine learning model is trained based on earthquake simulation data of different roughnesses; An identification result prediction module, used for obtaining an earthquake nucleation phase identification result based on the seismic phase moment; The earthquake nucleation phase identification system is also used to collect natural earthquake data to be measured from real earthquake events. If the format of the natural earthquake data to be measured is not a preset data format, the natural earthquake data to be measured is converted into a preset data format through format conversion, and the data after format conversion are spliced ​​in chronological order to obtain the earthquake data to be measured; when splicing, white noise is used to fill the missing part, and the filling standard deviation is the average value of the background noise before and after the splicing data; The earthquake nucleation phase identification system is also used to process the data features using a trained machine learning model based on a preset sampling factor, a preset step size and a preset time window to determine a number of seismic phase moments, wherein the seismic phase moments at least include the start time of the nucleation phase, the end time of the nucleation phase and the time when the earthquake occurs; The earthquake nucleation phase identification system is also used to collect earthquake simulation data of different roughness, and the earthquake simulation data includes a time data set and an amplitude data set; the earthquake simulation data is time synchronized and downsampled to obtain earthquake data samples; the earthquake data samples are segmented and feature extracted according to a preset time window to obtain data feature samples in each time window; based on the fault point label file and the time data set, label data for each time point is generated, and the fault point label file is generated based on a file containing earthquake event data in the earthquake simulation data; a training set is constructed based on the data feature samples in each time window and the corresponding label data; an initial machine learning model is constructed based on a machine learning algorithm; the initial machine learning model is trained using the training set until a preset training termination condition is reached to obtain a trained machine learning model.

8. An earthquake nucleation phase identification terminal, characterized in that: include: A memory for storing executable instructions; A processor, used to call and run the executable instructions in the memory to execute the steps of the earthquake nucleation phase identification method as described in any one of claims 1-2, and / or to execute the steps of the machine learning model training method as described in any one of claims 3-6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions. When the program instructions are executed by the processor, the steps of the earthquake nucleation phase identification method as described in any one of claims 1-2 are implemented, and / or the steps of the machine learning model training method as described in any one of claims 3-6 are executed.

Citation Information

Patent Citations

  • Time-domain-based convolutional neural network model for seismic phase recognition and application thereof

    CN112884134A