Seismic data augmentation method and system based on adversarial network and semi-supervised learning
Through adversarial networks and semi-supervised learning methods, seismic data is generated and labeled, which solves the problem of low efficiency in seismic data labeling, realizes efficient seismic data set construction and labeling, and supports seismic signal processing.
Patent Information
- Application Number
- CN202510735142.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing seismic data annotation efficiency is low, the manual annotation cost is high, and it is difficult to quickly update and expand the seismic data set, which hinders the application of deep learning technology in seismic signal processing.
A method based on adversarial networks and semi-supervised learning is used to collect data through seismic detectors. The generator generates unlabeled seismic data, adjusts the parameters through the discriminator, and combines the data annotation module to perform pseudo-label annotation to form a new seismic dataset.
It significantly improves the efficiency of seismic dataset construction, reduces annotation costs, improves the quality of datasets and annotation accuracy, and provides strong data support for seismic signal processing.
Smart Images

Figure CN120610309A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of seismic data processing, and in particular relates to a seismic data augmentation method and system based on adversarial networks and semi-supervised learning. Background Art
[0002] Deep learning technology has been widely used in seismic signal processing, mainly by learning from large amounts of labeled seismic data to improve the accuracy of seismic signal analysis.
[0003] Existing seismic data is mainly collected by seismometers and relies on manual labeling of seismic data. This method is not only inefficient, but also easily interfered with by human factors, making it difficult to ensure labeling consistency. In order to use deep learning technology to analyze seismic signals, a large amount of manpower and time are required to produce seismic datasets. In addition, the high cost of manual labeling makes frequent updates and expansions of seismic data training sets almost infeasible. Once a new earthquake event occurs, traditional manual labeling often cannot be completed in a short period of time, which hinders the rapid analysis of new seismic data by deep learning technology. How to significantly improve the efficiency of dataset construction while ensuring the quality of labeling has become a key difficulty in deep learning technology in seismic signal processing.
[0004] In recent years, generative adversarial networks (GANs) and semi-supervised learning techniques have demonstrated tremendous potential in data generation and automatic labeling. GANs can effectively generate data with realistic features, while semi-supervised learning techniques can automatically label large amounts of data using only a small number of labeled samples, providing new insights into seismic data augmentation. Summary of the Invention
[0005] To solve the problems existing in the existing technology, the present invention provides a seismic data augmentation method and system based on adversarial networks and semi-supervised learning. While ensuring data quality, it greatly improves the efficiency of seismic dataset construction and provides strong data support and technical guarantee for seismic signal processing.
[0006] To achieve the above object, the present invention provides the following solutions:
[0007] A seismic data augmentation method based on adversarial networks and semi-supervised learning, the method comprising:
[0008] Seismic signals are collected by geophones to obtain time series data, and real seismic data is obtained after data preprocessing;
[0009] Generate unlabeled earthquake data based on real earthquake data and the constructed data generation module;
[0010] The constructed data annotation module is used to annotate the unlabeled earthquake data and obtain earthquake data with pseudo labels;
[0011] We will rely on manually labeled earthquake data and pseudo-labeled earthquake data to form a new earthquake dataset and realize earthquake data augmentation.
[0012] Preferably, the method for generating unlabeled seismic data based on real seismic data and a constructed data generation module includes:
[0013] In the training phase, the generator is used to generate earthquake data x1; the discriminator is used to compare the similarity between the real earthquake data σ and the earthquake data x1, and the parameters of the generator are adjusted according to the comparison results;
[0014] Use the trained data generation module to generate unlabeled earthquake data x n .
[0015] Preferably, the method of generating seismic data x1 using the generator includes:
[0016] Input the random noise z into the multi-layer perceptron, obtain the initial time series data x through mapping, perform instance normalization on the initial time series data, and obtain the normalized time series data x′;
[0017] Use a lightweight neural network composed of a pooling layer and a two-layer perceptron to scan the normalized time series data x′ as a whole, divide the sequence into P segments, and output a size weight matrix
[0018] The preset N size sets of different sizes F={F1,F2,...,F N}, where each size corresponds to the length of a partition block, and the size set F is represented as a row vector
[0019] Multiply the size weight matrix W by the row vector F to obtain the dynamic partition size F corresponding to the qth time series data q ';
[0020] The dynamic partition size F corresponding to each time series data q 'composes the sequence F';
[0021] Cut the input sequence x′ according to the sequence F′ sliding to form several time series data x″ of different lengths;
[0022] Apply the attention mechanism structure to each segment to extract features of different scales;
[0023] The extracted features of different scales are aggregated by weighted summation to obtain the aggregated feature Y fuse ;
[0024] For the aggregated feature Y fuse Use 1D deconvolution to restore the seismic data x1 of the same dimension as the initial time series data x, where x1 represents the seismic data generated by the generator (G).
[0025] Preferably, the method of using the discriminator to compare the similarity between the real earthquake data σ and the earthquake data x1 includes:
[0026] Input the generated earthquake data x1 and the real earthquake data σ;
[0027] Use a short time window with a convolution kernel size of 3×3 and a step size of 1 to extract local information and obtain local features of the data;
[0028] Use a medium time window with a convolution kernel size of 9×9 and a step size of 3 to extract features from global information and obtain the global features of the data;
[0029] A long-term window with a convolution kernel size of 15×15 and a step size of 5 is used to obtain the long-term dependence of the data;
[0030] The local features of the data, the global features of the data and the long-term dependencies of the data are concatenated using the concat operation to obtain a complete seismic data feature vector;
[0031] The complete earthquake data feature vector is input into the fully connected layer for feature mapping, and the similarity between the generated earthquake data x1 and the real earthquake data σ is compared through the sigmoid function.
[0032] Preferably, the method of labeling unlabeled earthquake data using the constructed data labeling module to obtain earthquake data with pseudo labels includes:
[0033] In the training phase, in addition to using earthquake data σ, earthquake data x is also added to the training set. n Part of the data x n ′, seismic data σ and data x n 'Weighted sum of the loss weights after network training and back propagate to update the network parameters;
[0034] Use the trained data annotation module to annotate the unlabeled earthquake data x n Marking is performed to obtain earthquake data with pseudo labels.
[0035] Preferably, in the training phase, the method for processing the seismic data σ includes:
[0036] The earthquake data σ is manually labeled to obtain the true label δ. The earthquake data σ is labeled using the label prediction network to predict the label and obtain the predicted label result δ1. The predicted label δ1 is cross entropy with the true label δ:
[0037] H(δ,δ1)=-∫δ(σ)logδ1(σ)dσ
[0038] Among them, δ1 represents the prediction result of the earthquake data σ label; δ represents the label of the earthquake data σ manually labeled; δ(σ) represents the probability density function of the label δ; δ1(σ) represents the probability density function of the label δ1.
[0039] Preferably, during the training phase, the seismic data x n Part of the data x n 'The treatment methods include:
[0040] For the same seismic data, add n random noises and perform n different enhancements to obtain the enhanced seismic data x i ″,i∈(1,2,...,n); predict the label by using the label prediction network to obtain the earthquake data x i ″Label prediction result λ i ,i∈(1,2,...,n);
[0041] Use the confidence threshold M to analyze the earthquake data x i Filter the label prediction results:
[0042]
[0043] Among them, λ i Represents earthquake data x i "Prediction label; comparison() represents the comparison function; Represents the predicted label after comparison with the threshold M; the confidence threshold M is adaptively and dynamically adjusted:
[0044]
[0045] Among them, M min represents the initial confidence threshold; M max represents the final confidence threshold; c represents the current amount of training data; M total represents the total amount of training data; κ represents the flexibility factor of confidence adjustment; represents the control parameter of confidence smoothing; ε represents the amplification factor for adjusting the amount of training batch data;
[0046] After filtering, the same data Do mean square error mse:
[0047]
[0048] According to the mean square error mse, the total mean square error MSE is obtained:
[0049]
[0050] Where m represents the total number of data samples; n j Indicates the number of samples after random enhancement of the j-th data.
[0051] Preferably, the earthquake data σ and data x n 'The method of back-propagating the weighted sum of the loss weights after network training to update the network parameters includes:
[0052] The obtained cross entropy H(δ,δ1) and mean square error MSE weighted summation are used as the loss function of the label prediction network:
[0053] L=ρ·H(δ,δ1)+(1-ρ)·MSE
[0054] Among them, ρ represents the weighting coefficient and L represents the loss function.
[0055] The present invention also provides a seismic data augmentation system based on adversarial networks and semi-supervised learning, the system is used to implement the above method, the system includes: an acquisition module, a generation module, a labeling module and a synthesis module;
[0056] The acquisition module is used to collect seismic signals through seismic detectors, obtain time series data, and obtain real seismic data through data preprocessing;
[0057] The generating module is used to generate unlabeled seismic data based on real seismic data and the constructed data generating module;
[0058] The labeling module is used to label the unlabeled seismic data using the constructed data labeling module to obtain seismic data with pseudo labels;
[0059] The synthesis module is used to form a new earthquake data set by manually marking earthquake data with labels and earthquake data with pseudo labels, thereby realizing earthquake data augmentation.
[0060] Compared with the prior art, the present invention has the following beneficial effects:
[0061] This paper proposes a seismic data augmentation method and system based on adversarial networks and semi-supervised learning. By constructing a data generation module, the system automatically generates seismic data. Furthermore, by constructing a data annotation module, the system improves the efficiency and accuracy of data annotation, significantly reducing annotation costs and reliance on manual annotation. While ensuring data quality, this method significantly improves the efficiency of seismic dataset construction, providing strong data support and technical assurance for seismic signal processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0063] Figure 1 1 is a schematic flow chart of a method for augmenting seismic data based on adversarial networks and semi-supervised learning in an embodiment of the present invention;
[0064] Figure 2 is a schematic diagram of an adversarial network in an embodiment of the present invention;
[0065] Figure 3 is a schematic diagram of a generator (G) network in an embodiment of the present invention;
[0066] Figure 4 is a schematic diagram of a discriminator (D) network in an embodiment of the present invention;
[0067] Figure 5 2 is a schematic diagram of the semi-supervised learning training process in an embodiment of the present invention. DETAILED DESCRIPTION
[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0069] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0070] Example 1
[0071] like Figure 1As shown, this embodiment provides a seismic data augmentation method based on adversarial networks and semi-supervised learning, including: seismic data acquisition, real seismic data, manual labeling, data generation module, data annotation module and pseudo-label data set. Specifically:
[0072] Seismic signals are collected by seismic geophones to obtain time series data, and real seismic data σ is obtained after data preprocessing.
[0073] The data generation module constructed by the present invention generates a large amount of new earthquake data. In the training phase, the generator is used to generate earthquake data x1; the discriminator compares the similarity between the earthquake data σ and the earthquake data x1, and adjusts the parameters of the generator according to the comparison results. After training, the data generation module can be used to generate a large amount of new unlabeled earthquake data x1. n .
[0074] In order to give the above earthquake data x n In the training phase, in addition to using earthquake data σ, earthquake data x is also added to the training set to prevent overfitting of the model. n A small amount of data x n After training, the data annotation module is used to generate a large amount of new unlabeled earthquake data x for the data generation module. n Marking is performed to obtain a large amount of earthquake data with pseudo labels.
[0075] A new earthquake dataset will be formed by manually labeling earthquake data with labels and obtaining a large amount of pseudo-labeled earthquake data, thereby realizing earthquake data augmentation.
[0076] In this embodiment, data is collected: seismic data is collected using a seismic detector, and the data may contain various noises and abnormal values.
[0077] Data preprocessing operation: Clean the collected seismic data to remove noise and outliers to make the data purer.
[0078] The preprocessed data is used as the real earthquake data σ.
[0079] In this embodiment, the data generation module adopts an adversarial network method, such as Figure 2 As shown in Figure 2, random noise z is input into the generator, which continuously generates earthquake data x1. This data is then fed into the discriminator and compared with real earthquake data σ. The discriminator outputs the degree of similarity between the two, which is fed back to the generator as a judgment result. Based on this feedback, the generator continuously adjusts its parameters to improve the similarity between the generated data and the real earthquake data, thus achieving adversarial network training.
[0080] In the above adversarial network method, the generator (G) is as follows Figure 3 As shown, random noise z is mapped to initial time series data through a multi-layer perceptron (MLP). The data is normalized through an instance normalization layer. The processed data is scanned using a lightweight neural network to adaptively divide the time series into multiple scales. Feature extraction is performed using an attention mechanism, and the extracted features are weighted and aggregated for fusion. The fused features are then deconvolved to restore the data dimension. Specifically:
[0081] Input the random noise z into the multi-layer perceptron (MLP), and obtain the initial time series data through mapping: x = MLP (z), and then perform instance normalization:
[0082]
[0083]
[0084] Among them, x′: represents the normalized time series data; μ(x): represents the mean of the time series data; t: is the length of the time series data; σ(x): represents the standard deviation of the time series data; γ: represents the learnable scaling factor; α: represents the learnable offset factor.
[0085] Next, a lightweight neural network consisting of a pooling layer and a two-layer perceptron (MLP) is used to scan the time series x′ as a whole, divide the sequence into P segments, and output a size weight matrix It is used to adaptively assign the most appropriate partition size to each time series data segment. Indicates the matching degree of the P-th segment corresponding to each preset division size.
[0086] W=softmax(E2·ReLU(E1·Pool(x′)))
[0087] Among them, E1: represents the weight of the first fully connected layer; E2: represents the weight of the second fully connected layer; ReLU(): represents the activation function; Pool(): represents the pooling layer function.
[0088] The preset N sizes of different sizes are set F = {F1, F2, ..., F N}, where each size corresponds to the length of a partition block. The size set F is represented as a row vector Multiply the size weight matrix W output by the neural network with the row vector F to obtain the dynamic partition size F corresponding to the qth time series data q ′:
[0089] F q ′=Wq ·F T ,(q=1,...,P)
[0090] Among them, F q ': indicates the size of the data partition of the qth segment; W q : represents the qth row weight matrix; F T : represents the transpose of the row vector F.
[0091] The dynamic division size corresponding to each time series data is composed of sequence F':
[0092] F′={F1′,F2′,...,F P ′}
[0093] The input sequence x′ is cut according to the sliding sequence F′ to form several time series data x″ of unequal lengths.
[0094] x″=PatchDivisin(x′,F′)={x1′,x2′,...,x P ′}
[0095] Among them, x″: represents the time series data set after being divided into different sizes; PatchDivisin(): represents the division operation; x P ′: represents the time series data after being divided into different sizes.
[0096] After completing the above dynamic multi-scale division of time series data, in order to capture each segment x P ′The internal features are processed by applying the structure of the attention mechanism on each segment.
[0097] Introducing an additional set of adaptability matrices A into the traditional attention triple (Q, K, V), we define a new four-tuple attention mechanism (Q, A, K, V). The adaptability matrix A acts as a proxy for the query matrix, aggregating information from K and V, and then passing the information back to Q:
[0098] Y i =patchAttention(x″)
[0099]
[0100] Among them, Y i : represents the features of the i-th time series data after partitioning; patchAttention(): represents the attention mechanism function; Q: represents the query matrix; A: represents the adaptability matrix; K: represents the key matrix; V: represents the value matrix; B1, B2: represents the difference component; d: represents the dimension of the query and key; K T: Indicates the transposition of matrix K; softmax(): indicates the normalization operation.
[0101] The extracted features of different scales are aggregated by weighted summation:
[0102]
[0103] Among them, w i Represents the weight of the i-th feature; exp(): represents the exponential operation function; g(): represents the activation function; Y fuse Represents the features after aggregation.
[0104] Finally, for Y fuse Use 1D deconvolution to restore the seismic data x1 of the same dimension as the initial time series data x, where x1 represents the seismic data generated by the generator (G).
[0105] In the above adversarial network method, the discriminator (D) is as follows Figure 4 As shown, the generated seismic data x1 and the real seismic data σ are input into the discriminator. A multi-scale CNN method is used, and a short time window with a convolution kernel size of 3×3 and a step size of 1 is used to extract features of local information and learn the local features of the data; a medium time window with a convolution kernel size of 9×9 and a step size of 3 is used to extract features of global information and learn the global features of the data; a long time window with a convolution kernel size of 15×15 and a step size of 5 is used to learn the long-term dependencies of the data. The different scale features extracted by the multi-scale CNN are concatenated using the concat operation to obtain the complete seismic signal data feature vector. The feature vector is input into the fully connected layer for feature mapping, and the similarity between the generated seismic data x1 and the real seismic data σ is compared through the sigmoid function, and the result is fed back to the generator (G). Specifically:
[0106] Input the generated earthquake data x1 and the real earthquake data σ.
[0107] Use a convolution kernel size of 3×3 and a short time window of 1 to extract local information and learn local features of the data:
[0108] h local =Conv1D(η,k=3,s=1),η∈{x1,σ}
[0109] Among them, h local : represents the local information extracted in a short time window; η: represents the seismic data; k: represents the convolution kernel size; s: represents the convolution step size.
[0110] Use a medium time window with a convolution kernel size of 9×9 and a step size of 3 to extract features of global information and learn the global features of the data:
[0111] h global =Conv1D(η,k=9,s=3),η∈{x1,σ}
[0112] Among them, h global : Represents the global information extracted by the medium time window.
[0113] Use a long-term window with a convolution kernel size of 15×15 and a stride of 5 to learn the long-term dependencies of the data:
[0114] h long-term =Conv1D(η,k=15,s=5),η∈{x1,σ}
[0115] Among them, h long-term : Indicates long-term dependency information.
[0116] The different scale features extracted by multi-scale CNN are concatenated using the concat operation to obtain the complete seismic data feature vector:
[0117] h concat =Concat(h local ,h global ,h long-term )
[0118] Among them, h concat : Indicates the feature information after the splicing operation; Concat(): indicates the splicing operation.
[0119] Eigenvector h concat Input the fully connected layer for feature mapping:
[0120] β=R(ωh concat +b)
[0121] Among them, β: represents the feature map output; R(): represents the activation function; ω: represents the weight matrix; b: represents the bias vector.
[0122] The similarity between the generated earthquake data x1 and the real earthquake data σ is compared through the sigmoid function:
[0123]
[0124] The generator (G) and the discriminator (D) compete with each other, and according to the adversarial loss:
[0125]
[0126] Among them, min G max DV(D,G): indicates that the goal of the discriminator is to maximize this value (this value represents a probability value, which represents the degree of difference between the generated earthquake data and the real earthquake data. The goal of the discriminator is to distinguish the generated earthquake data from the real earthquake data. The larger the probability value, the greater the difference, and the easier it is to distinguish between the two). Try to distinguish between the real earthquake data and the generated earthquake data (according to the size of the above probability value, the goal of the discriminator is to maximize this value and distinguish according to the degree of difference); the goal of the generator is to minimize this value (this value is also the previous probability value. This probability value represents the degree of difference between the generated earthquake data and the real earthquake data. The goal of the generator is to minimize the degree of difference between the generated earthquake data and the real earthquake data), so that the generated earthquake data is closer to the real earthquake signal (the goal of the generator is to minimize this probability value, which will reduce the difference between the generated earthquake data and the real earthquake data, making the discriminator unable to distinguish between the real earthquake data and the generated earthquake data), so that the discriminator cannot distinguish them.
[0127] The adversarial network is trained to generate unlabeled earthquake data (unlabeled earthquake data refers to earthquake data generated by training the generator using the adversarial network method, and the generated earthquake data is not labeled). The parameters of the generator are adjusted according to the obtained adversarial loss. Specifically, according to the adversarial loss function, the goal of the generator is to reduce the difference between the generated earthquake data and the real earthquake data, so that the discriminator cannot distinguish between the real earthquake data and the generated earthquake data. The goal of the discriminator is to distinguish the generated earthquake data from the real earthquake data. According to the size of the loss value, the generator adjusts its parameters to continuously generate earthquake data that is closer to the real one and deceives the discriminator; the discriminator also adjusts its parameters to better distinguish between the real earthquake data and the generated earthquake data. With such continuous adjustments, the data generated by the generator will become more and more like the real earthquake data, and the generated data can be considered as correct data and used.
[0128] Use the trained data generation module to generate a large amount of new seismic signal data x n .
[0129] In this embodiment, the data annotation module uses a semi-supervised learning approach to label the large amount of seismic data generated by the data generation module. The data annotation module uses two batches of data input: one batch is processed using real seismic data to ensure the accuracy of the model data annotation. The other batch uses a small amount of generated seismic data to prevent model overfitting. Both batches of data use the same label prediction network. Finally, the weighted sum of the loss weights after network training for the two batches of data is back-propagated to update the network parameters.
[0130] like Figure 5As shown in the figure, first, the first batch of data: real earthquake data σ, is manually labeled to obtain the real label; after the earthquake data σ is randomly enhanced, the label prediction network is used to predict the label to obtain the predicted label result. The predicted label δ1 is cross-entropy with the real label δ, and the cross-entropy is used as the loss weight of the first batch of data. Next, in order to prevent overfitting, the second batch of data is added: the earthquake data x is generated n A small amount of data x n ′, for the same seismic data, add n random noises and perform n different enhancements to obtain the enhanced seismic data x i ″,i∈(1,2,...,n); predict the label by using the label prediction network to obtain the label prediction result λ i ,i∈(1,2,...,n). In order to improve the training speed of the label prediction network, an adaptive confidence threshold is used to predict the seismic data x i The label prediction results are screened, and the mean square error of the screened prediction labels is calculated. The total mean square error is used as the loss weight of the second batch of data. Finally, the obtained cross entropy and total mean square error are weighted summed as the loss value of the label prediction network, and back propagation is used to update and optimize the label prediction network parameters. Specifically:
[0131] The module processes the first batch of data: real earthquake data σ.
[0132] Earthquake data σ is manually labeled to obtain the true label δ. The label prediction network is used to predict the label of the earthquake data σ (the label refers to the label predicted by the label prediction network for the real earthquake data, which is the label prediction for the real earthquake data) to obtain the predicted label result δ1. The predicted label δ1 is cross entropy with the true label δ:
[0133] H(δ,δ1)=-∫δ(σ)logδ1(σ)dσ
[0134] Among them, δ1: represents the prediction result of the earthquake data σ label; δ: represents the label of the earthquake data σ manually labeled; δ(σ): represents the probability density function of the label δ; δ1(σ): represents the probability density function of the label δ1.
[0135] To prevent overfitting, the module adds a second batch of data: generate earthquake data x n A small amount of data x n ′, process it.
[0136] In the second batch of data, for the same earthquake data, n random noises are added and n different enhancements are performed to obtain the enhanced earthquake data x i″,i∈(1,2,...,n); predict the label (the label refers to the label predicted by the label prediction network of the earthquake data generated by the generator, and the label prediction is made on the generated earthquake data) by using the label prediction network to obtain the label prediction result λ i ,i∈(1,2,...,n).
[0137] In order to improve the training speed of the label prediction network, the confidence threshold M is used to train the earthquake data x i Filter the label prediction results:
[0138]
[0139] Among them, λ i Represents earthquake data x i "Prediction label; comparison(): represents the comparison function; Represents the predicted label after comparison with the threshold M; the confidence threshold M is adaptively and dynamically adjusted:
[0140]
[0141] Among them, M min : represents the initial confidence threshold; M max : represents the final confidence threshold; c: represents the current amount of training data; M total : represents the total amount of training data; κ: represents the flexibility factor of confidence adjustment; Represents the control parameter of confidence smoothing; ε: represents the amplification factor for adjusting the amount of training batch data.
[0142] For label prediction results greater than the label confidence threshold, it is considered to be the correct prediction label , label prediction results that are less than the label confidence threshold are set to 0. For the same data, when the label prediction network's label prediction results are all greater than the confidence threshold, it indicates that the network is strong and only requires fine-tuning. However, when some results are greater than the confidence threshold and some are less than the confidence threshold, it indicates that the network is weak and requires more extensive adjustments. By screening with this confidence threshold, the error of low-confidence prediction results can be amplified during error calculation, thereby improving network training speed and effectiveness.
[0143] After processing the same data Do mean square error mse:
[0144]
[0145] The total mean square error MSE after processing the second batch of data:
[0146]
[0147] Among them, m: represents the total number of data samples; n j : Indicates the number of samples after random enhancement of the j-th data.
[0148] After training the network with the above two batches of data, the weighted sum of the obtained cross entropy H(δ,δ1) and mean square error MSE is used as the loss value of semi-supervised learning training:
[0149] L=ρ·H(δ,δ1)+(1-ρ)·MSE
[0150] Where, ρ: represents the weighting coefficient. L: represents the loss function.
[0151] The loss value measures the gap between the training prediction label and the actual result. Back propagation uses the loss value to update and optimize the label prediction network parameters, minimize the prediction error, and improve the performance of the data labeling module.
[0152] In this embodiment, a trained data labeling module is used to label the earthquake data generated by the data generation module, thereby obtaining a large amount of earthquake data with pseudo labels.
[0153] In this embodiment, the manually labeled real earthquake data and the generated large amount of pseudo-labeled earthquake data are combined to form a new earthquake data set, thereby realizing earthquake data augmentation.
[0154] Example 2
[0155] The present invention also provides a seismic data augmentation system based on adversarial networks and semi-supervised learning, the system is used to implement the method of embodiment 1, the system includes: an acquisition module, a generation module, a labeling module and a synthesis module;
[0156] The acquisition module is used to collect seismic signals through seismic detectors, obtain time series data, and obtain real seismic data through data preprocessing;
[0157] A generation module, used for generating unlabeled seismic data based on real seismic data and the constructed data generation module;
[0158] The labeling module is used to label the unlabeled earthquake data using the constructed data labeling module to obtain earthquake data with pseudo labels;
[0159] The synthesis module is used to form a new earthquake dataset by combining earthquake data with labels manually marked and earthquake data with pseudo labels, thereby realizing earthquake data augmentation.
[0160] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A seismic data augmentation method based on adversarial networks and semi-supervised learning, characterized in that: The method comprises: Seismic signals are collected by geophones to obtain time series data, and real seismic data is obtained after data preprocessing; Generate unlabeled earthquake data based on real earthquake data and the constructed data generation module; The constructed data annotation module is used to annotate the unlabeled earthquake data and obtain earthquake data with pseudo labels; We will rely on manually labeled earthquake data and pseudo-labeled earthquake data to form a new earthquake dataset and realize earthquake data augmentation.
2. The method according to claim 1, characterized in that The method for generating unlabeled seismic data based on real seismic data and the constructed data generation module includes: In the training phase, the generator is used to generate earthquake data x1; the discriminator is used to compare the similarity between the real earthquake data σ and the earthquake data x1, and the parameters of the generator are adjusted according to the comparison results; Use the trained data generation module to generate unlabeled earthquake data x n .
3. The method according to claim 2, characterized in that The method of using the generator to generate earthquake data x1 includes: Input the random noise z into the multi-layer perceptron, obtain the initial time series data x through mapping, perform instance normalization on the initial time series data, and obtain the normalized time series data x′; Use a lightweight neural network composed of a pooling layer and a two-layer perceptron to scan the normalized time series data x′ as a whole, divide the sequence into P segments, and output a size weight matrix The preset N size sets of different sizes F={F1,F2,...,F N }, where each size corresponds to the length of a partition block, and the size set F is represented as a row vector Multiply the size weight matrix W by the row vector F to obtain the dynamic partition size F corresponding to the qth time series data q '; The dynamic partition size F corresponding to each time series data q 'composes the sequence F'; Cut the input sequence x′ according to the sequence F′ sliding to form several time series data x″ of different lengths; Apply the attention mechanism structure to each segment to extract features of different scales; The extracted features of different scales are aggregated by weighted summation to obtain the aggregated feature Y fuse ; For the aggregated feature Y fuse Use 1D deconvolution to restore the seismic data x1 of the same dimension as the initial time series data x, where x1 represents the seismic data generated by the generator (G).
4. The method according to claim 3, characterized in that Methods for using the discriminator to compare the similarity between the real earthquake data σ and the earthquake data x1 include: Input the generated earthquake data x1 and the real earthquake data σ; Use a short time window with a convolution kernel size of 3×3 and a step size of 1 to extract local information and obtain local features of the data; Use a medium time window with a convolution kernel size of 9×9 and a step size of 3 to extract features from global information and obtain the global features of the data; A long-term window with a convolution kernel size of 15×15 and a step size of 5 is used to obtain the long-term dependence of the data; The local features of the data, the global features of the data and the long-term dependencies of the data are concatenated using the concat operation to obtain a complete seismic data feature vector; The complete earthquake data feature vector is input into the fully connected layer for feature mapping, and the similarity between the generated earthquake data x1 and the real earthquake data σ is compared through the sigmoid function.
5. The method according to claim 1, wherein The method of labeling unlabeled earthquake data using the constructed data labeling module to obtain earthquake data with pseudo labels includes: In the training phase, in addition to using earthquake data σ, earthquake data x is also added to the training set. n Part of the data x n ′, seismic data σ and data x n 'Weighted sum of the loss weights after network training and back propagate to update the network parameters; Use the trained data annotation module to annotate the unlabeled earthquake data x n Marking is performed to obtain earthquake data with pseudo labels.
6. The method according to claim 5, characterized in that During the training phase, the processing methods for seismic data σ include: The earthquake data σ is manually labeled to obtain the true label δ. The earthquake data σ is labeled using the label prediction network to predict the label and obtain the predicted label result δ1. The predicted label δ1 is cross entropy with the true label δ: H(δ,δ1)=-∫δ(σ)logδ1(σ)dσ Among them, δ1 represents the prediction result of the earthquake data σ label; δ represents the label of the earthquake data σ manually labeled; δ(σ) represents the probability density function of the label δ; δ1(σ) represents the probability density function of the label δ1.
7. The method according to claim 6, characterized in that In the training phase, the earthquake data x n Part of the data x n 'The treatment methods include: For the same seismic data, add n random noises and perform n different enhancements to obtain the enhanced seismic data x i ″,i∈(1,2,...,n); predict the label by using the label prediction network to obtain the earthquake data x i ″Label prediction result λ i ,i∈(1,2,...,n); Use the confidence threshold M to analyze the earthquake data x i Filter the label prediction results: Among them, λ i Represents earthquake data x i "Prediction label; comparison() represents the comparison function; Represents the predicted label after comparison with the threshold M; the confidence threshold M is adaptively and dynamically adjusted: Among them, M min represents the initial confidence threshold; M max represents the final confidence threshold; c represents the current amount of training data; M total represents the total amount of training data; κ represents the flexibility factor of confidence adjustment; represents the control parameter of confidence smoothing; ε represents the amplification factor for adjusting the amount of training batch data; After filtering, the same data Do mean square error mse: According to the mean square error mse, the total mean square error MSE is obtained: Where m represents the total number of data samples; n j Indicates the number of samples after random enhancement of the j-th data.
8. The method according to claim 7, characterized in that Seismic data σ and data x n 'The method of back-propagating the weighted sum of the loss weights after network training to update the network parameters includes: The obtained cross entropy H(δ,δ1) and mean square error MSE weighted summation are used as the loss function of the label prediction network: L=ρ·H(δ,δ1)+(1-ρ)·MSE Among them, ρ represents the weighting coefficient and L represents the loss function.
9. A seismic data augmentation system based on adversarial networks and semi-supervised learning, the system being used to implement the method according to any one of claims 1 to 8, characterized in that: The system includes: an acquisition module, a generation module, a labeling module and a synthesis module; The acquisition module is used to collect seismic signals through seismic detectors, obtain time series data, and obtain real seismic data through data preprocessing; The generating module is used to generate unlabeled seismic data based on real seismic data and the constructed data generating module; The labeling module is used to label the unlabeled seismic data using the constructed data labeling module to obtain seismic data with pseudo labels; The synthesis module is used to form a new earthquake data set by manually marking earthquake data with labels and earthquake data with pseudo labels, thereby realizing earthquake data augmentation.
Citation Information
Patent Citations
Time sequence prediction method based on interactive multi-scale recurrent neural network
CN111027672A
Novel seismic data generation deep learning sample label method and seismic data adaptive partitioning lossless fitting method
CN115437010A
Image classification method, system and equipment based on dynamic semi-supervised deep learning
CN116188896A
Three-stage semi-supervised voiceprint recognition method based on adaptive expansion strategy
CN116543773A
Semi-supervised micro-seismic first arrival intelligent pickup method combining SimMatch and improved TransUGA
CN116990860A
Cited By
Seismic phase identification method and system fusing meta learning and transfer learning
CN121956119A