Seismic event automatic classification method based on three-branch convolutional neural network and transfer learning
Through the three-branch convolutional neural network and transfer learning method, combined with seismic waveform, spectrum and amplitude ratio characteristics, the limitations of traditional seismic classification methods are solved, and a more accurate and generalized seismic event classification is achieved.
Patent Information
- Application Number
- CN202510351028.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional seismic classification methods rely on artificial features and are difficult to fully describe the complex characteristics of earthquake events, resulting in insufficient classification accuracy and generalization capabilities, especially when facing non-natural earthquake events, they are prone to overfitting and misjudgment.
A three-branch convolutional neural network is used to process seismic waveform, spectrum and amplitude ratio characteristics respectively, and optimize the model through transfer learning and fuse multimodal features for seismic event classification.
It improves the accuracy of earthquake event classification and cross-regional generalization capabilities, reduces misjudgment of non-natural earthquake events, and provides a more comprehensive description of earthquake events.
Smart Images

Figure CN120296492A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of seismic signal processing and artificial intelligence, and particularly to an automatic seismic event classification method based on a three - branch convolutional neural network and transfer learning. Background Art
[0002] Seismic event classification is a key link in seismic monitoring and disaster assessment, and its accuracy directly affects fields such as earthquake prediction, underground fault research, and nuclear explosion monitoring. However, traditional seismic classification methods have many limitations and are difficult to meet the complex requirements of current large - scale seismic data processing.
[0003] Traditional seismic classification methods mainly rely on manually extracted features, such as waveform parameters, spectral features, etc. Although these features can reflect the characteristics of seismic events to a certain extent, their generalization ability is poor. On the one hand, the process of manually extracting features requires a large amount of experience and knowledge of domain experts and is easily affected by subjective factors, making it difficult to ensure the stability and consistency of feature selection. On the other hand, seismic events have significant regional differences and complexities, and the features extracted manually are often difficult to comprehensively and accurately describe the features of different regional and different - type seismic events, thus limiting the generalization performance of the classification model. With the rapid growth of seismic monitoring data, traditional methods have difficulty coping with the processing requirements of large - scale data, and both their classification efficiency and accuracy face severe challenges.
[0004] Non - natural seismic events, such as blasting and collapse, have a relatively low occurrence frequency and scarce sample numbers, resulting in a serious data imbalance problem when constructing a seismic classification model. During the training process of deep - learning models, they often overfit to majority - class samples and have poor classification performance for minority - class samples. This data imbalance phenomenon not only affects the model's ability to identify non - natural seismic events but also may lead to misjudgments in actual applications, bringing potential risks to seismic monitoring and disaster assessment.
[0005] In response to the above problems, classification models with single features (such as waveforms or spectra) show obvious limitations. On the one hand, single features are difficult to comprehensively describe the complex characteristics of seismic events, resulting in the classification model being sensitive to regional geological differences. Factors such as geological structures and seismic wave propagation paths in different regions may cause differences in the feature manifestations of the same - type seismic events in different regions, and single - feature models are difficult to capture such differences, thus affecting their cross - regional generalization ability. On the other hand, when facing the classification task of multiple - type seismic events, single - feature models often have difficulty distinguishing different event types with similar features, leading to a decline in classification accuracy.
[0006] In recent years, the application of convolutional neural networks (CNNs) in earthquake classification has gradually attracted attention. However, existing CNN models mostly adopt a single-branch structure, that is, they only classify a single feature (such as waveform or spectrum). Although this structure can extract the features of earthquake events to a certain extent, it fails to effectively fuse multi-modal earthquake features (such as waveform, spectrum, physical features, etc.), thus limiting the classification performance of the model. Summary of the Invention
[0007] The purpose of the present invention is to provide an automatic earthquake event classification method based on a three-branch convolutional neural network and transfer learning, so as to solve the problem that single features are difficult to comprehensively describe the complex characteristics of earthquake events, resulting in inaccurate classification.
[0008] To achieve the above purpose, the present invention provides an automatic earthquake event classification method based on a three-branch convolutional neural network and transfer learning, including:
[0009] S1. Preprocess the original three-component seismic waveform data to obtain denoised, normalized, and time-aligned waveform data;
[0010] S2. Extract spectrum data according to the preprocessed waveform data, and calculate the amplitude ratio of Pg waves and Sg waves;
[0011] S3. Input the preprocessed waveform data, spectrum data, and amplitude ratio into a three-branch convolutional neural network to obtain waveform branch features, spectrum branch features, and amplitude ratio features;
[0012] S4. Concatenate or weighted-fuse the output features of the three branches to form a comprehensive feature, perform classification processing on the comprehensive feature, and output the classification probabilities of earthquake, blasting, and collapse events.
[0013] Further, the original three-component seismic waveform data includes multiple earthquake event datasets, and each earthquake event dataset includes earthquake, blasting, and collapse events.
[0014] Further, the S1 step specifically includes:
[0015] S11. De-mean and trend-correct the original three-component seismic waveform data to eliminate baseline drift and DC components;
[0016] S12. Perform smoothing processing using a Hanning window, and the window length accounts for 5% of the total waveform length to reduce the truncation effect;
[0017] S13. Filter the waveform using a 2Hz high-pass filter to remove low-frequency noise;
[0018] S14. Perform amplitude normalization on the waveform to ensure that the amplitude ranges of different events are consistent;
[0019] S15. Cut a fixed waveform length, and randomly start the cutting time 5-20 seconds before the arrival of the P wave to ensure the maximum signal-to-noise ratio.
[0020] Furthermore, the step of extracting spectrum data specifically includes: segmenting the waveform using a 2-second sliding window with a 75% overlap rate; performing discrete Fourier transform on the waveform data in each window, and calculating the spectrum within the range of 0.5-50 Hz.
[0021] Furthermore, the amplitude ratio of the Pg wave to the Sg wave is calculated by the following formula:
[0022]
[0023] Among them, Pz is the amplitude of the vertical component of the Pg wave, Sz is the amplitude of the vertical component of the Sg wave, and Nz is the amplitude of the vertical component of the noise.
[0024] Furthermore, the window of the Pg wave and the Sg wave is from the first 5% to the last 60% of the Sg-Pg time difference, with a total length limit of 3 seconds; the first 60% of the Sg-Pg time difference is used as background noise.
[0025] Furthermore, the S3 step specifically includes:
[0026] Waveform branch processing: The preprocessed waveform data is processed through 4 layers of convolution, the convolution output is flattened, and the features are further extracted through the fully connected layer to obtain a 420-dimensional feature vector; each layer of convolution is followed by a 4×1 maximum pooling layer, and a Dropout layer is added after each layer starting from the second convolution layer;
[0027] Spectrum branch processing: The spectrum data is processed through 4 layers of convolution, the convolution output is flattened, and the features are further extracted through the fully connected layer to obtain a 420-dimensional feature vector; each convolution layer is followed by a 2×2 maximum pooling layer, and a Dropout layer is added after the fully connected layer;
[0028] Amplitude ratio branch processing: The amplitude ratio is mapped into a 420-dimensional feature vector through a fully connected layer.
[0029] Furthermore, the S4 step specifically includes:
[0030] S41, waveform branch features, spectrum branch features and amplitude ratio features are concatenated into a 1260-dimensional vector;
[0031] S42, pass the concatenated features through the Dropout layer to suppress overfitting;
[0032] S43. Input the features processed by S42 into the fully connected layer, and output the classification probabilities of earthquake, explosion and collapse through the Softmax function.
[0033] Further, it further includes: S0. Conduct transfer learning training on the three-branch convolutional neural network:
[0034] S01. Train a three-branch CNN pre-training model based on the initial dataset;
[0035] S02. Combine the pre-training model and the target domain dataset, unfreeze at least part of the convolutional layers of the waveform and spectrum branches, and adjust the parameters of the fully connected layer based on the target domain dataset to form an optimized dedicated model for the target area.
[0036] Further, it further includes: S0'. Augment the original three-component seismic waveform data: randomly offset the waveform start time 5 - 10 times to generate waveform data with different noise information.
[0037] Beneficial effects achieved by the present invention:
[0038] The automatic earthquake event classification method based on a three-branch convolutional neural network and transfer learning proposed by the present invention effectively improves the limitations of traditional earthquake classification methods and has significant technical advantages and practical application value.
[0039] First, by preprocessing the original three-component seismic waveform data, waveform data that is denoised, normalized, and time-aligned is obtained, significantly improving the data quality. The denoising process effectively removes high-frequency noise and low-frequency drift, retaining the effective seismic signal, making subsequent feature extraction and model training more accurate. The normalization and time-alignment processes eliminate the amplitude differences and time-axis offsets between different stations or events, ensuring the consistency and comparability of the data, and providing a high-quality data basis for subsequent model training.
[0040] Second, spectral data is extracted from the preprocessed waveform data, and the amplitude ratios of Pg waves and Sg waves are calculated, further enriching the feature information. The spectral data reflects the characteristics of the seismic signal in the frequency domain and can capture the frequency component differences of different earthquake events. The amplitude ratio feature is sensitive to the focal depth and event type and helps to distinguish different earthquake events. The extraction of these multi-modal features provides a more comprehensive description of earthquake events for the model and enhances the model's ability to distinguish different earthquake events.
[0041] Then, the preprocessed waveform data, spectral data, and amplitude ratio are input into a three-branch convolutional neural network to obtain waveform branch features, spectral branch features, and amplitude ratio features respectively. The three-branch structure enables the model to process different features separately, improving the diversity and effectiveness of feature extraction. The waveform branch focuses on learning the spatial and temporal features of the waveform data, the spectral branch focuses on learning spectral features, and the amplitude ratio branch directly processes the amplitude ratio features. Each branch processes features independently, preserving the differences between different features and avoiding the limitations of a single-feature model, which helps to understand seismic events more comprehensively.
[0042] Finally, the output features of the three branches are concatenated or weighted and fused to form comprehensive features, and the comprehensive features are classified to output the classification probabilities of seismic, blasting, and collapse events. Feature fusion combines the features of multiple branches, forming a more comprehensive representation of seismic events, which helps to improve the classification accuracy. Description of the Drawings
[0043] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0044] Figure 2 It is a schematic diagram of the process of step S1 of the present invention;
[0045] Figure 3 It is a schematic diagram of the process of step S4 of the present invention;
[0046] Figure 4 It is a schematic diagram of the overall process of the present invention with pre-steps S0 and S0' added;
[0047] Figure 5 It is a schematic diagram of the process of step S0 of the present invention;
[0048] Figure 6 It is a schematic diagram of the amplitude ratio calculation window of the present invention;
[0049] Figure 7 It is an architecture diagram of the three-branch CNN network model of the present invention.
[0050] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent; for better illustration of this embodiment, some components in the drawings may be omitted, enlarged, or reduced, and do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted; the same or similar reference numerals correspond to the same or similar components; the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent. Detailed Embodiments
[0051] It should be noted that, without conflict, the embodiments in the present application and the technical features in the embodiments may be combined with each other. The detailed descriptions in the specific embodiments should be understood as the explanatory descriptions of the purpose of the present application and should not be regarded as an improper limitation of the present application.
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not used to limit the scope of the present application.
[0053] In the embodiments of the present application, the term "comprising", "including", or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising such element.
[0054] The following will introduce and describe the technical solutions of the present invention in detail with reference to specific accompanying drawings.
[0055] As Figure 4 shown, the complete process in this embodiment is as follows:
[0056] First, construct a three-branch CNN model, as Figure 7 shown, including a waveform branch, a spectrum branch, and an amplitude ratio branch.
[0057] Among them, the waveform branch: 4 layers of convolution (the number of filters is 64 → 128 → 256 → 512), 4×1 max pooling, ReLU activation, and Dropout (0.3) is added to the last 3 layers, and a fully connected layer (420 nodes). The waveform branch is used to extract waveform features (i.e., a 420-dimensional feature vector, representing a high-level abstraction of the spatio-temporal features of the waveform).
[0058] The spectrum branch: 4 layers of convolution (the number of filters is 18 → 36 → 68 → 68), 2×2 max pooling, ReLU activation, and Dropout (0.5) is added after the fully connected layer (420 nodes). The spectrum branch is used to extract spectrum features (i.e., a 420-dimensional feature vector, representing a deep expression of the spectrum pattern).
[0059] The amplitude ratio branch: directly processes scalar features with a fully connected layer. The amplitude ratio branch is used to extract amplitude features (i.e., 420-dimensional feature values, representing a quantization index of the physical attributes of the event).
[0060] Second, perform model training.
[0061] The model input is the US ComCat dataset (earthquake and blasting events). First, a three-branch CNN model is trained on the ComCat dataset to learn general feature expressions. The initial learning rate is set to 0.001 and dynamically adjusted (halved when the loss stagnates). The processed result is a pre-trained model.
[0062] III. Perform fine-tuning on the target domain (as Figure 5 shown, in this embodiment, the Chinese "Ditong" 2.0 dataset is used for illustration).
[0063] First, freeze all convolutional layers. Then, unfreeze the high-level convolutional layers of the waveform and spectrum branches, and keep the first 2 shallow convolutional layers of the waveform and spectrum branches frozen. Subsequently, fine-tune the parameters of the fully connected layer based on the target domain data. The initial learning rate is set to 0.001 and dynamically adjusted (halved when the loss stagnates). The processed result is an optimized region-specific model.
[0064] IV. The region-specific model performs earthquake classification. (as Figure 1 shown)
[0065] 1. Data preprocessing
[0066] 1.1 Waveform processing
[0067] As Figure 2 shown, the input is the original three-component seismic waveform data (length ≥ 60 seconds, sampling rate 100 Hz). First, the waveform data is de-meaned and detrended to eliminate baseline drift and DC components. Subsequently, Hanning window smoothing is performed, with the window length accounting for 5% of the total waveform length, to reduce the truncation effect. Then, a 2 Hz high-pass filter is used to filter the waveform to remove low-frequency noise. Finally, amplitude normalization is performed on the waveform to ensure consistent amplitude ranges for different events. The preprocessed waveform data is truncated to a fixed length (6001 sampling points, 60 seconds), and the truncation time starts randomly 5 - 20 seconds before the P-wave arrival time to maximize the signal-to-noise ratio. The processed waveform data represents a denoised, normalized, and time-aligned waveform signal, providing high-quality input for subsequent feature extraction.
[0068] 1.2 Spectrum calculation
[0069] The input is the preprocessed waveform data. First, the waveform is segmented using a 2-second sliding window with a 75% overlap rate to ensure the continuity of spectrum calculation. Subsequently, the discrete Fourier transform (DFT) is performed on the waveform data within each window to calculate the spectrum in the range of 0.5 - 50 Hz. To enhance the resolution of the Rg wave (for shallow-source events), the frequency resolution is increased during spectrum calculation. The processed spectrum data is a spectrogram of 117×100×3, representing the energy distribution of the signal in the frequency domain and capable of effectively capturing the differences in event frequency characteristics (such as the Rg wave characteristics of blasting events).
[0070] 1.3 Amplitude Ratio Feature Extraction
[0071] The input is the preprocessed waveform data. As Figure 6 shown, where N is the noise window and T is the Sg - Pg arrival time difference. First, calculate the amplitude ratio of the Pg wave and the Sg wave, and the formula is as follows: where Pz is the amplitude of the Pg wave in the vertical component, Sz is the amplitude of the Sg wave in the vertical component, and Nz is the amplitude of the noise in the vertical component. The window selection is based on the Sg - Pg arrival time difference. The length of the Pg window is from the first 5% to the last 60%, and the total length is limited to 3 seconds; the noise window intercepts the first 60% of the Sg - Pg arrival time difference as the background noise. The single - station amplitude ratio feature takes the median of the amplitude ratios of all the Z - component records of the event to reduce the random error of the single - station records. The processed amplitude ratio feature is a scalar value, representing the physical characteristics of the event source mechanism, and can effectively distinguish natural earthquakes, blasts, and collapse events.
[0072] Previously, in order to alleviate the data imbalance problem, the method of randomly offsetting the waveform start time can be used to generate multiple sets of noise - enhanced data. In the specific operation, the waveform start time is randomly offset 5 - 10 times to generate waveform data with different noise information. The processed data is the balanced training set, representing the expanded sample size, and can significantly improve the classification performance of the model on small - sample events (such as blasts and collapses).
[0073] 2. Three - branch Convolutional Neural Network Processing (as Figure 2 shown)
[0074] 2.1 Waveform Branch Processing
[0075] The input is the preprocessed waveform data (three - component waveform of 6001×1×3). First, the waveform data undergoes 4 - layer convolution processing. The number of filters is 64, 128, 256, and 512 respectively. After each layer of convolution, a 4×1 max - pooling layer is connected to highlight the significant features and reduce the computational amount. Starting from the second layer of convolution, a Dropout layer (0.3) is added after each layer to suppress overfitting. Subsequently, the convolution output is flattened and further feature - extracted through a fully - connected layer (420 nodes). The output of the processed waveform branch is a 420 - dimensional feature vector.
[0076] 2.2 Spectrum Branch Processing
[0077] The input is spectrogram data (three-component spectrum of 117×100×3). First, the spectral data undergoes 4 layers of convolution processing, with the number of filters being 18, 32, 68, and 68 respectively. After each layer of convolution, a 2×2 max pooling layer is connected. Subsequently, the convolution output is flattened and further feature extraction is performed through a fully connected layer (420 nodes). A Dropout layer (0.5) is added after the fully connected layer to enhance the generalization ability of the model. The output of the processed spectral branch is a 420-dimensional feature vector.
[0078] 2.3 Processing of amplitude ratio branch
[0079] The input is amplitude ratio features (scalar values). The amplitude ratio features are directly mapped to 420-dimensional feature values through a fully connected layer.
[0080] 2.4 Feature fusion and classification
[0081] The inputs are waveform branch features (420-dimensional), spectral branch features (420-dimensional), and amplitude ratio features (420-dimensional). First, the three-way features are concatenated into a 1260-dimensional vector, and then a Dropout layer (0.5) is used to further suppress overfitting. Finally, the concatenated features are input into a fully connected layer, and the probability distributions of three categories (earthquake, blasting, collapse) are output through the Softmax function.
[0082] Furthermore, to improve the accuracy, robustness, and generalization ability of the model in the earthquake event classification task, this embodiment proposes steps for optimizing classification performance:
[0083] 1. Data augmentation: Randomly shift the waveform start time (5 - 10 times) to generate multiple sets of noise-enhanced data; oversample non-natural events (blasting, collapse) to balance the event category ratio.
[0084] 2. Physical feature enhancement: Apply high-frequency band-pass filtering (10 - 18 Hz) to recalculate the amplitude ratio of Pg wave and Sg wave; combine waveform, spectral, and amplitude ratio features to construct a multi-modal input.
[0085] 3. Cross-region data fusion: Integrate multi-region datasets (such as ComCat, "Ditong" 2.0, Shanxi, Northeast); combine blasting and collapse events according to the magnitude distribution ratio to ensure data balance; retrain the model after data augmentation.
[0086] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.
Claims
1. An automatic seismic event classification method based on a three-branch convolutional neural network and transfer learning, characterized in that, Including: S1. Preprocess the original three-component seismic waveform data to obtain denoised, normalized, and time-aligned waveform data; S2. Extract spectral data based on the preprocessed waveform data and calculate the amplitude ratio of Pg wave and Sg wave; S3. Input the preprocessed waveform data, spectral data, and amplitude ratio into a three-branch convolutional neural network to obtain waveform branch features, spectral branch features, and amplitude ratio features; S4. Concatenate or weighted fuse the output features of the three branches to form a comprehensive feature, perform classification processing on the comprehensive feature, and output the classification probabilities of seismic, blasting, and collapse events.
2. The automatic earthquake event classification method based on a three-branch convolutional neural network and transfer learning according to claim 1, wherein The original three-component seismic waveform data includes multiple seismic event data sets, and each seismic event data set includes seismic, blasting, and collapse events.
3. The automatic earthquake event classification method based on a three-branch convolutional neural network and transfer learning according to claim 2, characterized in that The specific steps of step S1 include: S11. Remove the mean value and perform trend correction on the original three-component seismic waveform data to eliminate baseline drift and DC components; S12. Perform smoothing processing using a Hanning window, and the window length accounts for 5% of the total waveform length to reduce the truncation effect; S13. Filter the waveform using a 2Hz high-pass filter to remove low-frequency noise; S14. Perform amplitude normalization processing on the waveform to ensure that the amplitude ranges of different events are consistent; S15. Intercept a fixed waveform length, and the interception time randomly starts 5 - 20 seconds before the arrival time of the P wave to maximize the signal-to-noise ratio.
4. The automatic earthquake event classification method based on a three-branch convolutional neural network and transfer learning according to claim 1, characterized in that The specific steps of extracting spectral data include: Segment the waveform using a 2-second sliding window with a 75% overlap rate; Perform discrete Fourier transform on the waveform data within each window and calculate the spectrum within the range of 0.5 - 50Hz.
5. The automatic earthquake event classification method based on a three-branch convolutional neural network and transfer learning according to claim 1, characterized in that The amplitude ratio of Pg wave and Sg wave is calculated by the following formula: where Pz is the amplitude of the Pg wave in the vertical component, Sz is the amplitude of the Sg wave in the vertical component, and Nz is the amplitude of the noise in the vertical component.
6. The automatic seismic event classification method based on a three-branch convolutional neural network and transfer learning according to claim 5, characterized in that The window of the Pg wave and Sg wave is from the first 5% to the last 60% of the Sg - Pg arrival time difference, and the total length is limited to 3 seconds; The first 60% of the Sg - Pg arrival time difference is used as background noise.
7. The automatic seismic event classification method based on a three-branch convolutional neural network and transfer learning according to claim 1, characterized in that, The specific steps of step S3 include: Waveform branch processing: Pass the preprocessed waveform data through 4 layers of convolution processing, flatten the convolution output, and further extract features through a fully connected layer to obtain a 420-dimensional feature vector; Among them, each layer of convolution is followed by a 4×1 max pooling layer, and a Dropout layer is added starting from the second layer of convolution; Spectral branch processing: Pass the spectral data through 4 layers of convolution processing, flatten the convolution output, and further extract features through a fully connected layer to obtain a 420-dimensional feature vector; Among them, each layer of convolution is followed by a 2×2 max pooling layer, and a Dropout layer is added after the fully connected layer; Amplitude ratio branch processing: Map the amplitude ratio to a 420-dimensional feature vector through a fully connected layer.
8. The automatic seismic event classification method based on a three-branch convolutional neural network and transfer learning according to claim 7, characterized in that The specific steps of step S4 include: S41. Concatenate the waveform branch features, spectral branch features, and amplitude ratio features into a 1260-dimensional vector; S42. Pass the concatenated features through a Dropout layer to suppress overfitting; S43. Input the features processed in S42 into a fully connected layer and output the classification probabilities of seismic, blasting, and collapse through the Softmax function.
9. The automatic seismic event classification method based on a three-branch convolutional neural network and transfer learning according to claim 1, characterized in that Also including: S0. Conduct transfer learning training on the three-branch convolutional neural network: S01. Train the three-branch CNN pre-training model based on the initial dataset; S02. Combine the pre-training model and the target domain dataset, unfreeze at least part of the convolutional layers of the waveform and spectrum branches, and adjust the parameters of the fully connected layer based on the target domain dataset to form an optimized model dedicated to the target area.
10. The automatic seismic event classification method based on a three-branch convolutional neural network and transfer learning according to claim 1, characterized in that It also includes: S0'. Augment the original three-component seismic waveform data: randomly shift the waveform start time 5 - 10 times to generate waveform data with different noise information.