Self-supervised earthquake denoising method and system fusing noise estimation module and channel attention mechanism
By fusing the noise estimation module and channel attention mechanism in seismic data denoising, combined with pixel downsampling technology, the problems of noise model complexity, data annotation dependence and insufficient feature expression in the existing technology are solved, and better denoising effect and generalization are achieved.
Patent Information
- Application Number
- CN202510177197.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The problems of noise model complexity, data labeling dependence and insufficient feature expression in real seismic data denoising have not been effectively solved, resulting in poor denoising effect.
The self-supervised seismic denoising method is adopted that combines the noise estimation module and the channel attention mechanism, combined with pixel downsampling technology, denoising is performed through the noise prediction module and the self-supervised blind spot network (BSN), and only real noise-containing data is used for training.
It improves the generalization and practicality of seismic data denoising, enhances the ability to remove complex noise, and avoids the problems of noise residues and effective signal loss in traditional methods.
Smart Images

Figure BDA0005275902240000031 
Figure BDA0005275902240000061 
Figure BDA0005275902240000151
Abstract
Description
Technical Field
[0001] The present invention relates to the field of oil and gas exploration and seismic data processing, and provides a self-supervised seismic denoising method and system integrating a noise estimation module and a channel attention mechanism. Background Art
[0002] In the field of seismic exploration, the quality of received data directly determines the accuracy of underground structural analysis. However, seismic signals are susceptible to various noise interferences during the acquisition process. Among them, random noise has become a core issue affecting the quality of seismic data due to its wide distribution and overlap with the effective signal frequency band. This type of noise not only masks weak reflection information, but also leads to the accumulation of errors in subsequent inversion and interpretation. Therefore, efficient seismic denoising technology is a key link in oil and gas exploration and data processing.
[0003] Traditional seismic denoising methods are mainly based on physical model assumptions, such as using the spatial differences between noise and seismic events (such as predictive filtering technology) or the sparsity of seismic reflections (such as sparse representation denoising). Although such methods are effective in certain scenarios, their performance is highly dependent on artificially preset prior conditions, and parameters need to be repeatedly adjusted for data in different work areas, making it difficult to adapt to complex noise environments. In addition, traditional methods have limited ability to extract nonlinear features of seismic signals, which can easily cause effective signal loss or residual noise.
[0004] In recent years, deep learning-based seismic denoising methods have significantly improved denoising performance. Models such as fully convolutional networks (FCN), U-Net, DnCNN, and Transformer have achieved nonlinear mapping between noise and signal through end-to-end learning. However, existing deep learning methods have the following bottlenecks:
[0005] 1. Supervised learning relies on disadvantages: Most models require a large amount of paired “noisy-clean” data for supervised training, but it is difficult to obtain sufficient high-quality noise-free data in actual exploration, which limits the generalization of the model.
[0006] 2. Limitations of synthetic data: Existing methods often use Gaussian white noise simulation training, but real seismic noise has spatial correlation and complex statistical characteristics (such as non-stationarity and heteroscedasticity), and the model trained with synthetic data has significantly reduced effect on actual noise removal.
[0007] 3. Insufficient feature extraction: Traditional convolutional networks find it difficult to balance local details and global dependencies, and lack dynamic perception of feature weights between channels, resulting in inaccurate noise estimation, especially in weak signal areas, which can easily lead to over-smoothing or structural distortion.
[0008] In response to the above problems, some studies have tried to introduce self-supervised learning to avoid clean data dependence, such as the Blind Spot Network (BSN) to achieve pixel-level denoising through masked convolution. However, BSN assumes that noise is independent in space, which is inconsistent with the spatial correlation of real seismic noise, and its limited receptive field makes it difficult to model long-range dependencies, resulting in incomplete removal of complex noise.
[0009] In summary, the existing technology has not effectively solved the core problems in real seismic data denoising, such as the complexity of noise models, data annotation dependence and insufficient feature expression. There is an urgent need for an innovative method that takes into account self-supervised training, adaptive modeling of noise characteristics and multi-scale feature fusion. Summary of the invention
[0010] The purpose of the present invention is to provide a self-supervised seismic denoising method that integrates a noise estimation module and a channel attention mechanism, combined with pixel downsampling technology, to solve the problems in the prior art that the actual noise removal effect is poor and the original data is easily removed during denoising. At the same time, only real noisy data can be used during training, which has stronger generalization and practicality than using only synthetic data with Gaussian noise.
[0011] In order to achieve the above object, the present invention adopts the following technical solution:
[0012] The present invention provides a self-supervised seismic denoising method integrating a noise estimation module and a channel attention mechanism, comprising the following steps:
[0013] Step 1: construct a seismic data set, including slicing the real noisy seismic data and the synthetic noisy seismic data to obtain multiple noisy data blocks;
[0014] Step 2: Build a self-supervised denoising network, including:
[0015] A noise prediction module (CNNest), which consists of multiple 1×1 convolutional layers and ReLU activation functions, is used to output a noise level prediction map with the same size as the input data, and the noise level prediction map represents the noise covariance matrix of each pixel;
[0016] Self-supervised Blind Spot Network (BSN), which includes a dual-branch structure with integrated channel attention mechanism:
[0017] The first branch uses a dilated convolution module with center mask convolution, referred to as the DC module, to extract local features;
[0018] The second branch uses a dilated Transformer block combined with a channel attention mechanism, referred to as a DTB module. The DTB module implements global feature modeling through a channel attention mechanism, and its feed-forward layer contains a depthwise separable convolution.
[0019] Step 3: Apply pixel random downsampling PD operation to the input data to destroy the noise spatial correlation, and input the processed data into the self-supervised network for training, use the Bayesian prediction formula to fuse the output of the noise prediction module with the intermediate features of the self-supervised blind spot network, and optimize the network parameters through back propagation to obtain a denoising model;
[0020] Step 4: Input the seismic data to be denoised into the trained model for denoising, and restore the final denoising result through the inverse PD operation.
[0021] In the above method, step 2 comprises the following steps:
[0022] Step 2.1: Construct a noise prediction module, which includes five 1x1 convolutional layers connected in sequence. The first four layers are followed by a ReLU nonlinear activation function, and the fifth layer outputs a noise covariance matrix of the same size as the input noisy data.
[0023] Step 2.2: Construct a dual-branch architecture of the self-supervised blind spot network BSN, where
[0024] The first branch includes the upper and lower paths:
[0025] Upper path: includes a 3×3 center mask convolutional layer and 9 DC modules in series. The DC module here uses dilated convolution with a stride of 2 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolutional layer and a 3×3 dilated convolutional layer to extract local features.
[0026] Lower path: It contains a DC module with a 5×5 center mask convolution layer and 9 dilated convolution layers with S=3. The DC module here uses dilated convolution with a stride of 3 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolution layer and a 3×3 dilated convolution layer to extract local features.
[0027] The second branch includes a 5×5 center mask convolution layer and 9 DTB modules connected in series. The DTB module generates a global feature map of channel interaction through a channel attention mechanism, and uses deep separable convolution and feedforward network to enhance long-range dependencies. The channel attention mechanism is implemented by calculating the dot product operation of the query matrix Q, the key matrix K and the value matrix V. The outputs of the first and second branches are concatenated after removing the blind spot information, and multi-scale features are fused through a 1×1 convolution layer. At the same time, the output of the noise prediction module is Bayesian predicted with the fused features to form the final BSN output result.
[0028] In the above method, step 3 comprises the following steps:
[0029] Step 3.1: Perform pixel random downsampling PD processing with a step factor of s=2 on the input real noisy seismic data, split the input data into 2×2 sub-blocks and arrange them in mosaic, and generate four sub-block data to break the noise spatial correlation;
[0030] Step 3.2: Input the PD-processed sub-block data into the noise prediction module to obtain a 1×1 noise covariance matrix for each pixel, and expand the matrix dimension to match the input data size;
[0031] Step 3.3: Feed the expanded noise covariance matrix and the sub-block data processed by PD into the self-supervised blind spot network BSN;
[0032] Step 3.4: Calculate the network prediction error based on the Bayesian formula, which is:
[0033]
[0034] Explanation of symbols:
[0035] y: noisy image input, the target to be denoised;
[0036] E[·]: average over all training data;
[0037] μ: clean signal predicted by the network, corresponding to D-BSN output;
[0038] Σ μ : covariance matrix of the predicted signal;
[0039] ∑ n : The noise covariance matrix output by the noise prediction module;
[0040] tr(·): matrix trace operation, constraining ∑ μ The amplitude of
[0041] Step 3.5: Optimize the noise prediction module and BSN network parameters through the back propagation algorithm, obtain the trained denoising model weights and save them, and finally obtain the trained model.
[0042] In the above method, step 4 comprises the following steps:
[0043] Step 4.1: Load the denoising model weights trained in claim 3, and input the noisy seismic data to be processed into the model;
[0044] Step 4.2: Perform pixel random downsampling PD processing with a step factor of s=2 on the input noisy data to form four discrete sub-block data, and feed them into the denoising model;
[0045] Step 4.3: The denoised sub-block data after processing is output by the noise prediction module of the denoising model and the self-supervised blind spot network BSN, and then each sub-block is subjected to an inverse PD operation according to the mosaic arrangement during the original PD processing, and is merged into the final denoised seismic data of the full size.
[0046] The present invention also provides a self-supervised seismic denoising system integrating a noise estimation module and a channel attention mechanism, comprising:
[0047] A seismic data set construction module includes slicing real noisy seismic data and synthetic noisy seismic data to obtain multiple noisy data blocks;
[0048] Self-supervised denoising network module, including:
[0049] A noise prediction module (CNNest), which consists of multiple 1×1 convolutional layers and ReLU activation functions, is used to output a noise level prediction map with the same size as the input data, and the noise level prediction map represents the noise covariance matrix of each pixel;
[0050] Self-supervised Blind Spot Network (BSN), which includes a dual-branch structure with integrated channel attention mechanism:
[0051] The first branch uses a dilated convolution module with center mask convolution, referred to as the DC module, to extract local features;
[0052] The second branch uses a dilated Transformer block combined with a channel attention mechanism, referred to as a DTB module. The DTB module implements global feature modeling through a channel attention mechanism, and its feed-forward layer contains a depthwise separable convolution.
[0053] The data processing and training unit applies a pixel random downsampling (PD) operation to the input data to destroy the spatial correlation of the noise, and inputs the processed data into the self-supervised network for training, fuses the output of the noise prediction module with the intermediate features of the self-supervised blind spot network using the Bayesian prediction formula, and optimizes the network parameters through back propagation to obtain a denoising model;
[0054] The denoising and recovery unit inputs the seismic data to be denoised into the trained model for denoising, and recovers the final denoising result through the inverse PD operation.
[0055] In the above system, the implementation of the self-supervised denoising network module includes the following steps:
[0056] Step 2.1: Construct a noise prediction module, which includes five 1x1 convolutional layers connected in sequence. The first four layers are followed by a ReLU nonlinear activation function, and the fifth layer outputs a noise covariance matrix of the same size as the input noisy data.
[0057] Step 2.2: Construct a dual-branch architecture of the self-supervised blind spot network BSN, where
[0058] The first branch includes the upper and lower paths:
[0059] Upper path: includes a 3×3 center mask convolutional layer and 9 DC modules in series. The DC module here uses dilated convolution with a stride of 2 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolutional layer and a 3×3 dilated convolutional layer to extract local features.
[0060] Lower path: It contains a DC module with a 5×5 center mask convolution layer and 9 dilated convolution layers with S=3. The DC module here uses dilated convolution with a stride of 3 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolution layer and a 3×3 dilated convolution layer to extract local features.
[0061] The second branch includes a 5×5 center mask convolution layer and 9 DTB modules connected in series. The DTB module generates a global feature map of channel interaction through a channel attention mechanism, and uses deep separable convolution and feedforward network to enhance long-range dependencies. The channel attention mechanism is implemented by calculating the dot product operation of the query matrix Q, the key matrix K and the value matrix V. The outputs of the first and second branches are concatenated after removing the blind spot information, and multi-scale features are fused through a 1×1 convolution layer. At the same time, the output of the noise prediction module is Bayesian predicted with the fused features to form the final BSN output result.
[0062] In the above system, the implementation of the data processing and training unit includes the following steps:
[0063] Step 3.1: Perform pixel random downsampling PD processing with a step factor of s=2 on the input real noisy seismic data, split the input data into 2×2 sub-blocks and arrange them in mosaic, and generate four sub-block data to break the noise spatial correlation;
[0064] Step 3.2: Input the PD-processed sub-block data into the noise prediction module to obtain a 1×1 noise covariance matrix for each pixel, and expand the matrix dimension to match the input data size;
[0065] Step 3.3: Feed the expanded noise covariance matrix and the sub-block data processed by PD into the self-supervised blind spot network BSN;
[0066] Step 3.4: Calculate the network prediction error based on the Bayesian formula, which is:
[0067]
[0068] Explanation of symbols:
[0069] y: noisy image input, the target to be denoised;
[0070] E[·]: average over all training data;
[0071] μ: clean signal predicted by the network, corresponding to D-BSN output;
[0072] Σ μ : covariance matrix of the predicted signal;
[0073] ∑ n : The noise covariance matrix output by the noise prediction module;
[0074] tr(·): matrix trace operation, constraining Σ μ The amplitude of
[0075] log|·|: Logarithmic determinant, used to measure the complexity of the noise model.
[0076] Step 3.5: Optimize the noise prediction module and BSN network parameters through the back propagation algorithm, obtain the trained denoising model weights and save them, and finally obtain the trained model.
[0077] In the above system, the implementation of the denoising and restoration unit includes the following steps:
[0078] Step 4.1: Load the denoising model weights trained in claim 3, and input the noisy seismic data to be processed into the model;
[0079] Step 4.2: Perform pixel random downsampling PD processing with a step factor of s=2 on the input noisy data to form four discrete sub-block data, and feed them into the denoising model;
[0080] Step 4.3: The denoised sub-block data after processing is output by the noise prediction module of the denoising model and the self-supervised blind spot network BSN, and then each sub-block is subjected to an inverse PD operation according to the mosaic arrangement during the original PD processing, and is merged into the final denoised seismic data of the full size.
[0081] Because the present invention adopts the above technical means, it has the following beneficial effects:
[0082] 1. The present invention uses PD to convert real seismic noise into AWGN-like noise, which is consistent with the basic assumption of BSN.
[0083] 2. The present invention adopts a noise prediction module and a self-supervised network BSN with an integrated channel attention mechanism. By estimating the noise intensity to limit the denoising type, good noise-free data characteristics can be extracted in both real seismic data and synthetic seismic data, thereby improving the denoising performance and generalization of deep learning for seismic data.
[0084] 3. Noise adaptive modeling: Through the noise prediction module (CNNest) constructed by 1×1 convolutional layers, the noise covariance matrix is learned pixel by pixel using a full convolutional architecture, which solves the problem of mismatch between the traditional synthetic noise model and the statistical characteristics of real seismic noise. This module dynamically corrects the blind spot network output through the Bayesian formula to achieve accurate estimation of non-stationary and heteroscedastic noise, so that the denoising intensity is adaptively matched with the local noise level, avoiding the loss of weak signals caused by uniform denoising.
[0085] 4. Feature enhancement under blind spot constraints: A dual-branch self-supervised blind spot network (BSN) is designed to force the exclusion of pixel information through center mask convolution, breaking the reliance of traditional supervised learning on clean data. The DC module uses a dilated convolution stack with a step size of 2 / 3 to expand the receptive field to 115×115 pixels while maintaining the blind spot constraint, solving the problem of insufficient local feature extraction of conventional convolution and significantly improving the ability to capture complex noise patterns.
[0086] 4. Global-local feature collaboration: Innovatively integrate the DTB module in BSN to achieve cross-channel global interaction through the channel attention mechanism. Specifically, replace the spatial attention with the channel attention in the Transformer block to avoid violating the blind spot constraint, and use the point multiplication operation of the query matrix (Q) and the key matrix (K) to establish long-range dependencies, compensating for the modeling defect of the pure CNN architecture for the continuity of seismic signals.
[0087] 5. Cracking the spatial correlation of noise: Introducing pixel random downsampling (PD) technology, through mosaic block reorganization with a step length of s=2, the original noise correlation distance is expanded from 1-2 pixels in the neighborhood to 2-4 pixels. According to actual measurements, this method reduces the noise spatial correlation coefficient from 0.89 to 0.06, forcing the BSN noise independence assumption to be met, and solving the core contradiction between the spatial correlation of real seismic noise and the conflict of algorithm assumptions.
[0088] 6. Multi-scale feature fusion optimization: The noise estimation module and the BSN network are jointly optimized through the Bayesian prediction formula. This loss function enables the synthetic data pre-training model to quickly converge to the real noise distribution, effectively alleviating the domain shift problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 It is a flowchart of the present invention;
[0090] Figure 2 It is a schematic diagram of a prediction module of the present invention;
[0091] Figure 3 This is a schematic diagram of the DTB module in the present invention;
[0092] Figure 4This is a schematic diagram of a DC module with S=2 in the present invention;
[0093] Figure 5 It is a schematic diagram of the PD operation and the inverse PD operation under the random sampling of pixels in the present invention;
[0094] Figure 6 For the present invention Figure 2 Schematic diagram of denoising results of earthquake records;
[0095] Figure 7 For the present invention Figure 2 Schematic diagram of the removed noise. DETAILED DESCRIPTION
[0096] The following is a detailed description of the embodiments of the present invention. Although the present invention will be described and illustrated in conjunction with some specific embodiments, it should be noted that the present invention is not limited to these embodiments. On the contrary, modifications or equivalent substitutions made to the present invention should all be included in the scope of the claims of the present invention.
[0097] In addition, in order to better illustrate the present invention, numerous specific details are given in the following specific embodiments. It will be understood by those skilled in the art that the present invention can also be implemented without these specific details.
[0098] The present invention first proposes a self-supervised network BSN that uses a noise prediction module and an integrated channel attention mechanism, uses a DC module to capture the local characteristics of seismic data, and uses a DTB module to capture the global characteristics of seismic data, so as to better obtain the implicit characteristics of the seismic data itself. On the basis of predicting the noise intensity of seismic data, pixel shuffling and downsampling are used to improve the generalization of the network.
[0099] To this end, the present invention provides an implementation method, a self-supervised seismic denoising method integrating a noise estimation module, comprising the following steps:
[0100] Step 1: Collect real noisy seismic data, slice the noisy seismic data to obtain multiple noisy seismic data blocks, then add synthetic seismic data of the same size and manually add Gaussian white noise to construct a seismic data set;
[0101] Step 2: Build a self-supervised blind spot network based on the channel attention mechanism and prediction network, and combine the processed noise estimation module with the blind spot network;
[0102] Step 3: Apply a pixel random downsampling method to the noise module to reduce spatial correlation, and feed it into the entire network, and use the training set to train the seismic data denoising model to obtain a trained denoising model for the real noise of the seismic data;
[0103] Step 4: Input the test set into the trained denoising model and merge it into the denoised seismic data.
[0104] Further, the step 1 comprises the following steps:
[0105] Step 1.1: Use the real noisy seismic data and the synthetic noisy seismic data to slice, with a slice size of 128×128, as the seismic dataset. The seismic dataset includes a training set, a validation set, and a test set.
[0106] Further, the step 2 comprises the following steps:
[0107] Step 2.1: Preprocess the noisy seismic data before inputting it into the network. Here, the data is normalized and rotated, symmetric, and other operations are performed to improve the robustness of the algorithm.
[0108] Step 2.2: Noise estimation module construction
[0109] The noisy seismic data are fed into the noise prediction module CNNest. The noise estimation module CNNest consists of five 1x1 convolutional layers with 16 channels, and ReLU nonlinearity is deployed in all convolutional layers except the last one.
[0110] Assuming that the noise is conditionally pixel-independent given the underlying clean data, using CNNest, an FCN architecture with 1x1 convolution, to learn the noise model is beneficial to learning the information around each pixel. Benefiting from all 1x1 convolution layers, it can be guaranteed that the noise level at a location depends only on the input value at the same location. Therefore, CNNest takes seismic noise as input and approximates a noise level function in its multivariate heteroscedastic Gaussian model through learning. Each pixel of the input data outputs a 1x1 covariance matrix, which is subsequently used for Bayesian prediction with the results of the BSN output to obtain denoised seismic data.
[0111] Step 2.3: Construction of Self-Supervised Blind Spot Network (BSN)
[0112] Under the assumption that spatial noise is irrelevant, the BSN network excludes the pixel itself from the receptive field of each pixel, preventing it from learning its own identity. It can then be trained with the same seismic noise data as input and target to learn to remove spatially independent noise at the pixel level. However, since the network cannot see blind spot information, BSN will face greater information loss under the condition of self-supervision.
[0113] To address the above problems, the self-supervised network BSN based on the channel attention mechanism of the present invention starts from a 1×1 convolution and enters two groups of identical network branches, each of which contains a 3×3 or 5×5 center mask convolution layer and multiple dilated convolution modules with a step size of 2 or 3, namely DC modules. Finally, the feature maps of the two branches are connected.
[0114] A binary mask m of size 3×3 and 5×5 is introduced, 0 is assigned to the central element of m, and 1 is assigned to other elements, and the center mask convolution is realized by the element-wise product of the mask and the convolution kernel of the same size. Further stacking of center mask convolution layers will break the blind spot requirement, so a DC module is added to maintain the blind spot requirement. The DC module refers to adding an interval S between the original convolution kernels to enhance the receptive field, while the number of parameters only increases linearly. Each DC module used in the present invention contains a 3×3 or 5×5 dilated convolution with a dilation coefficient of S, where S=2 and S=3 are used for the upper and lower paths of the network, respectively. Each weight loads the pre-trained model weights of the pre-trained non-integrated BSN.
[0115] Step 2.4: Integrate the DTB module based on the channel attention mechanism
[0116] Since the receptive field of the blind spot network is limited and the CNN-based BSN cannot fully capture long-distance dependencies, this limits its performance. Therefore, the DTB module is introduced to make up for this shortcoming and enhance the global modeling capability by introducing the dilated Transformer block. The core of the DTB module is to design a Transformer block that can exchange global information while meeting the blind spot requirements. Specifically, it includes two core components: the self-attention layer and the feedforward layer. For the self-attention layer, the channel attention mechanism is adopted instead of the spatial attention mechanism to avoid information exchange between adjacent pixels; for the feedforward layer, the depth-separable convolution is introduced to reduce the computational cost and enhance the local context.
[0117] In order to achieve global information exchange without violating the blind spot requirements, the DTB module adopts a channel attention mechanism. This mechanism does not rely on spatial position information, but achieves global perception through interaction between channels, and can obtain global information of seismic data. In order to achieve global information exchange without violating the blind spot requirements, the DTB module adopts a channel attention mechanism. This mechanism does not rely on spatial position information, but achieves global perception through interaction between channels. Specifically, the Query (Q), Key (K), and Value (V) matrices are first calculated, and then the attention map of channel interaction is obtained through a dot product operation. The DC module in step 2.3 is changed to a DTB module, layer normalization is added to the feedforward processing, and a 5×5 center mask convolution layer is selected.
[0118] The results after the DTB module are then fused with the results after the DC module to generate the network output. The integrated network can effectively make a comprehensive judgment on the model prediction results and obtain better results than the simple blind spot network model.
[0119] Further, the step 3 comprises the following steps:
[0120] Step 3.1: BSN's assumption that spatial noise is uncorrelated does not match the actual situation of seismic data, so PD is introduced, which means creating mosaics by downsampling seismic data with a specific step size factor s, increasing the actual distance between noise signals and breaking the spatial correlation of noise. After verification, the step size factor s = 2 is selected here. BSN is applied to the downsampled seismic data that meets its assumptions after processing, and the PD inverse operation is then performed after denoising to reconstruct the full-size output.
[0121] Step 3.2: Input the PD-processed seismic data into the noise prediction module, expand the length, width and dimension of the obtained results to make them the same size as the input seismic data, feed the results and seismic data into BSN for training, and tune them through the preset loss function. Among them, the loss function uses the Bayesian formula to comprehensively predict the limitations of the noise model and the results of denoising. Finally, a trained model is obtained.
[0122] Further, the step 4 comprises the following steps:
[0123] Step 4.1: Load the model weights saved after training into the model, and then input the real seismic data containing noise into the model loaded with the model weights for denoising to obtain denoised seismic data.
[0124] Example 1
[0125] Combine the following Figure 1-5 The present invention is described in detail.
[0126] like Figure 1 As shown in FIG. 1 , a self-supervised seismic denoising method integrating a noise estimation module and a channel attention mechanism comprises the following steps:
[0127] Step 1: Use real seismic data and synthetic noisy seismic data to slice and create a data set. The size of the slice is 128*128.
[0128] Step 2: Input the seismic data into the network for training. The network has a noise prediction module and a BSN with integrated channel attention mechanism. The noise prediction module is as follows: Figure 2 As shown in the figure, there are two basic modules of BSN: DTB module as Figure 3 As shown, the DC module is Figure 4 As shown;
[0129] The specific steps are:
[0130] Step 2.1: Preprocess the noisy seismic data before inputting it into the network. Here, the data is normalized and rotated to improve the robustness of the algorithm.
[0131] Step 2.2: Perform PD operation on the data with a step size of S = 2, and the obtained data is ready to be input into the prediction network and BSN, where the PD operation is as follows Figure 5 shown.
[0132] Step 2.3: Input the data into the noise prediction module to train the noise prediction module.
[0133] The noisy seismic data is input into the noise prediction module. First, the seismic noise model is learned through four consecutive 1x1 convolutional layers and ReLU layers. The output channel of the convolutional layer is 16 and the step size is 1. Finally, a 1x1 convolutional layer is passed to further extract the local features of the noise. The output channel of this convolutional layer is 1, and a 1x1 covariance matrix is output for each pixel of the input data. Finally, a noise level prediction image with the same length and width as the input data is obtained.
[0134] Step 2.4: Enter the data into the BSN section containing the DC module.
[0135] The BSN of the present invention starts from a 1×1 convolution and enters two parts, a part containing a DC module and a part containing a DTB module. The part containing the DC module is divided into two groups of branches, called upper and lower paths. The first group of branches, i.e., the upper path, contains a DC module containing a 3×3 center mask convolution layer and 9 dilated convolution layers with S=2, and the second group of branches, i.e., the lower path, contains a DC module containing a 5×5 center mask convolution layer and 9 dilated convolution layers with S=3. The prepared seismic data is input into the upper and lower paths at the same time.
[0136] A binary mask m of size 3×3 is introduced, 0 is assigned to the central element of m and 1 is assigned to other elements, and the center mask convolution is realized by the element-wise product of the mask and the convolution kernel of the same size.
[0137] Further stacking of center mask convolutional layers will break the blind spot requirement, so a DC module is added to maintain the blind spot requirement. The DC module refers to adding a spacing S between the original convolution kernels to enhance the receptive field, while the number of parameters only increases linearly.
[0138] The upper path passes through 9 DC modules, each of which contains three branches, called the upper, middle and lower branches. The upper branch contains: a 1×1 convolution plus a layer of ReLU, two 3×3 dilated convolutions, and each layer is followed by a layer of ReLU with a dilation coefficient of 2; the middle branch contains: a 1×1 convolution plus a layer of ReLU, a 3×3 dilated convolution plus a layer of ReLU with a dilation coefficient of 2; the lower branch contains a 1×1 convolution plus a layer of ReLU. The step size of the 1×1 convolution is 1, the step size of the 3×3 dilated convolution is 1, and the number of channels of each convolution layer is 96.
[0139] The lower path passes through 9 DC modules, each of which contains three branches, called the upper, middle and lower branches. The upper branch contains: a 1×1 convolution plus a layer of ReLU, two 3×3 dilated convolutions, each followed by a layer of ReLU with a dilation factor of 3; the middle branch contains: a 1×1 convolution plus a layer of ReLU, a 3×3 dilated convolution plus a layer of ReLU with a dilation factor of 3; the lower branch contains a 1×1 convolution plus a layer of ReLU. The step size of the 1×1 convolution is 1, the step size of the 3×3 dilated convolution is 1, and the number of channels of each convolution layer is 96.
[0140] The input seismic data passes through each DC module, and the results obtained from the upper, middle and lower branches are stacked. The stacked results are passed through a layer of 1×1 convolution and a layer of ReLU, where the number of channels in the convolution layer is 96 and the step size is 1. The result is added to the output of the previous DC module to get the final result and then sent to the next DC module.
[0141] Finally, the feature maps obtained from the upper and lower paths are stacked, and four 1×1 convolutional layers are deployed to extract deep network features and generate the BSN partial output containing the DC module. The number of channels in the convolutional layer is 32 and the stride is 1.
[0142] Step 2.5: Enter the data into the BSN section containing the DTB module.
[0143] The prepared seismic data is input into the part containing the DTB module. This part contains a 5×5 center mask convolution layer and 9 DTB modules with 3×3 dilated convolution layers of S=3. Each DTB module first passes through a layer normalization layer and then feeds into three branches at the same time. Among them, each branch contains a 3×3 dilated convolution layer with S=3, the number of output channels of the convolution layer is 96, the stride is 1, and the channel attention mechanism is adopted to achieve global perception through the interaction between channels. First, the Query (Q), Key (K) and Value (V) matrices are calculated, and then the attention map of channel interaction is obtained by the dot multiplication operation. The attention map of channel interaction is added to the original data, and the result is passed through layers of normalization layers. The features extracted by the expanded depth-wise convolution pass through the gating unit for nonlinearity, which contains two branches. Each branch contains a 3×3 dilated convolution layer with S=3, the number of output channels of the convolution layer is 96, and the stride is 1. And the gating unit is the element-by-element product of two parallel paths, one of which is activated by the GELU unit. The final result is added to the input data to obtain the final output result of the BSN part containing the DTB module.
[0144] Step 3: Train the BSN.
[0145] The output results of the BSN part containing the DTB module and the BSN part containing the DC module are added together, and the Bayesian prediction is performed with the result of the noise prediction module. The final feature extraction is performed through 4 1×1 convolutional layers to obtain the final BSN network output result. Training and optimization are then performed based on this process. The optimizer is Adam.
[0146] Step 4: De-noise the seismic data.
[0147] The specific steps are:
[0148] Step 3.1: Load the model weights saved after training into the model, and then input the noisy seismic data into the model loaded with the model weights for denoising to obtain denoised seismic data.
[0149] Step 3.2: Perform an inverse PD operation on the denoised seismic data to restore the impact of the PD operation during input and obtain the final seismic denoising data.
[0150] In summary, a self-supervised seismic denoising method integrating noise estimation module and channel attention mechanism is proposed to denoise noisy seismic data. Figure 7 As shown, the horizontal axis is the number of seismic traces, and the vertical axis is the number of sampling points. The denoising results show that this method can better restore the original seismic data and remove noise, proving the correctness of the method.
[0151] Example 2
[0152] The present invention also provides a self-supervised seismic denoising system integrating a noise estimation module and a channel attention mechanism, comprising:
[0153] A seismic data set construction module includes slicing real noisy seismic data and synthetic noisy seismic data to obtain multiple noisy data blocks;
[0154] Self-supervised denoising network module, including:
[0155] A noise prediction module (CNNest), which consists of multiple 1×1 convolutional layers and ReLU activation functions, is used to output a noise level prediction map with the same size as the input data, and the noise level prediction map represents the noise covariance matrix of each pixel;
[0156] Self-supervised Blind Spot Network (BSN), which includes a dual-branch structure with integrated channel attention mechanism:
[0157] The first branch uses a dilated convolution module with center mask convolution, referred to as the DC module, to extract local features;
[0158] The second branch uses a dilated Transformer block combined with a channel attention mechanism, referred to as a DTB module. The DTB module implements global feature modeling through a channel attention mechanism, and its feed-forward layer contains a depthwise separable convolution.
[0159] The data processing and training unit applies a pixel random downsampling (PD) operation to the input data to destroy the spatial correlation of the noise, and inputs the processed data into the self-supervised network for training, fuses the output of the noise prediction module with the intermediate features of the self-supervised blind spot network using the Bayesian prediction formula, and optimizes the network parameters through back propagation to obtain a denoising model;
[0160] The denoising and recovery unit inputs the seismic data to be denoised into the trained model for denoising, and recovers the final denoising result through the inverse PD operation.
[0161] In the above system, the implementation of the self-supervised denoising network module includes the following steps:
[0162] Step 2.1: Construct a noise prediction module, which includes five 1x1 convolutional layers connected in sequence. The first four layers are followed by a ReLU nonlinear activation function, and the fifth layer outputs a noise covariance matrix of the same size as the input noisy data.
[0163] Step 2.2: Construct a dual-branch architecture of the self-supervised blind spot network BSN, where
[0164] The first branch includes the upper and lower paths:
[0165] Upper path: includes a 3×3 center mask convolutional layer and 9 DC modules in series. The DC module here uses dilated convolution with a stride of 2 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolutional layer and a 3×3 dilated convolutional layer to extract local features.
[0166] Lower path: It contains a DC module with a 5×5 center mask convolution layer and 9 dilated convolution layers with S=3. The DC module here uses dilated convolution with a stride of 3 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolution layer and a 3×3 dilated convolution layer to extract local features.
[0167] The second branch includes a 5×5 center mask convolution layer and 9 DTB modules connected in series. The DTB module generates a global feature map of channel interaction through a channel attention mechanism, and uses deep separable convolution and feedforward network to enhance long-range dependencies. The channel attention mechanism is implemented by calculating the dot product operation of the query matrix Q, the key matrix K and the value matrix V. The outputs of the first and second branches are concatenated after removing the blind spot information, and multi-scale features are fused through a 1×1 convolution layer. At the same time, the output of the noise prediction module is Bayesian predicted with the fused features to form the final BSN output result.
[0168] In the above system, the implementation of the data processing and training unit includes the following steps:
[0169] Step 3.1: Perform pixel random downsampling PD processing with a step factor of s=2 on the input real noisy seismic data, split the input data into 2×2 sub-blocks and arrange them in mosaic, and generate four sub-block data to break the noise spatial correlation;
[0170] Step 3.2: Input the PD-processed sub-block data into the noise prediction module to obtain a 1×1 noise covariance matrix for each pixel, and expand the matrix dimension to match the input data size;
[0171] Step 3.3: Feed the expanded noise covariance matrix and the sub-block data processed by PD into the self-supervised blind spot network BSN;
[0172] Step 3.4: Calculate the network prediction error based on the Bayesian formula, which is:
[0173]
[0174] Explanation of symbols:
[0175] y: noisy image input, the target to be denoised;
[0176] E[·]: average over all training data;
[0177] μ: clean signal predicted by the network, corresponding to D-BSN output;
[0178] Σ μ : covariance matrix of the predicted signal;
[0179] ∑ n : The noise covariance matrix output by the noise prediction module;
[0180] tr(·): matrix trace operation, constraining ∑ μ The amplitude of
[0181] log|·|: Logarithmic determinant, used to measure the complexity of the noise model.
[0182] Step 3.5: Optimize the noise prediction module and BSN network parameters through the back propagation algorithm, obtain the trained denoising model weights and save them, and finally obtain the trained model.
[0183] In the above system, the implementation of the denoising and restoration unit includes the following steps:
[0184] Step 4.1: Load the denoising model weights trained in claim 3, and input the noisy seismic data to be processed into the model;
[0185] Step 4.2: Perform pixel random downsampling PD processing with a step factor of s=2 on the input noisy data to form four discrete sub-block data, and feed them into the denoising model;
[0186] Step 4.3: The denoised sub-block data after processing is output by the noise prediction module of the denoising model and the self-supervised blind spot network BSN, and then each sub-block is subjected to an inverse PD operation according to the mosaic arrangement during the original PD processing, and is merged into the final denoised seismic data of the full size.
[0187] The invention adds DTB module and DC module to BSN to optimize the network structure and enhance the feature extraction capability of seismic data. At the same time, the denoising performance is optimized through the seismic data noise intensity prediction module. In addition, the generalization of the network is enhanced through PD and inverse PD operations. A self-supervised seismic denoising method integrating noise estimation module and channel attention mechanism is proposed. The denoising result obtained after the method is applied to noisy seismic data is better.
[0188] The above are only representative embodiments of the present invention in many specific application scopes, and do not constitute any limitation on the protection scope of the present invention. Any technical solutions formed by transformation or equivalent replacement fall within the protection scope of the present invention.
Claims
1. A self-supervised seismic denoising method integrating noise estimation module and channel attention mechanism, characterized in that: The following steps are involved: Step 1: construct a seismic data set, including slicing the real noisy seismic data and the synthetic noisy seismic data to obtain multiple noisy data blocks; Step 2: Build a self-supervised denoising network, including: A noise prediction module, composed of multiple 1×1 convolutional layers and ReLU activation functions, is used to output a noise level prediction map with the same size as the input data, wherein the noise level prediction map represents the noise covariance matrix of each pixel; Self-supervised blind spot network, including a dual-branch structure with integrated channel attention mechanism: The first branch uses a dilated convolution module with center mask convolution, referred to as the DC module, to extract local features; The second branch uses a dilated Transformer block combined with a channel attention mechanism, referred to as a DTB module. The DTB module implements global feature modeling through a channel attention mechanism, and its feed-forward layer contains a depthwise separable convolution. Step 3: Apply pixel random downsampling PD operation to the input data to destroy the noise spatial correlation, and input the processed data into the self-supervised network for training, use the Bayesian prediction formula to fuse the output of the noise prediction module with the intermediate features of the self-supervised blind spot network, and optimize the network parameters through back propagation to obtain a denoising model; Step 4: Input the seismic data to be denoised into the trained model for denoising, and restore the final denoising result through the inverse PD operation.
2. The method according to claim 1, characterized in that Step 2 includes the following steps: Step 2.1: Construct a noise prediction module, which includes five 1x1 convolutional layers connected in sequence. The first four layers are followed by a ReLU nonlinear activation function, and the fifth layer outputs a noise covariance matrix of the same size as the input noisy data. Step 2.2: Construct a dual-branch architecture of the self-supervised blind spot network BSN, where The first branch includes the upper and lower paths: Upper path: includes a 3×3 center mask convolutional layer and 9 DC modules in series. The DC module here uses dilated convolution with a stride of 2 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolutional layer and a 3×3 dilated convolutional layer to extract local features. Lower path: It contains a DC module with a 5×5 center mask convolution layer and 9 dilated convolution layers with S=3. The DC module here uses dilated convolution with a stride of 3 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolution layer and a 3×3 dilated convolution layer to extract local features. The second branch includes a 5×5 center mask convolution layer and 9 DTB modules connected in series. The DTB module generates a global feature map of channel interaction through a channel attention mechanism, and uses deep separable convolution and feedforward network to enhance long-range dependencies. The channel attention mechanism is implemented by calculating the dot product operation of the query matrix Q, the key matrix K and the value matrix V. The outputs of the first and second branches are concatenated after removing the blind spot information, and multi-scale features are fused through a 1×1 convolution layer. At the same time, the output of the noise prediction module is Bayesian predicted with the fused features to form the final BSN output result.
3. The method according to claim 1, characterized in that The step 3 comprises the following steps: Step 3.1: Perform pixel random downsampling PD processing with a step factor of s=2 on the input real noisy seismic data, split the input data into 2×2 sub-blocks and arrange them in mosaic, and generate four sub-block data to break the noise spatial correlation; Step 3.2: Input the PD-processed sub-block data into the noise prediction module to obtain a 1×1 noise covariance matrix for each pixel, and expand the matrix dimension to match the input data size; Step 3.3: Feed the expanded noise covariance matrix and the sub-block data processed by PD into the self-supervised blind spot network BSN; Step 3.4: Calculate the network prediction error based on the Bayesian formula, which is: Explanation of symbols: y: noisy image input, the target to be denoised; E[·]: average over all training data; μ: clean signal predicted by the network, corresponding to D-BSN output; ∑ μ : covariance matrix of the predicted signal; ∑ n : The noise covariance matrix output by the noise prediction module; tr(·): matrix trace operation, constraining Σ μ The amplitude of Step 3.5: Optimize the noise prediction module and BSN network parameters through the back propagation algorithm, obtain the trained denoising model weights and save them, and finally obtain the trained model.
4. The method according to claim 1, characterized in that: The step 4 comprises the following steps: Step 4.1: Load the denoising model weights trained in claim 3, and input the noisy seismic data to be processed into the model; Step 4.2: Perform pixel random downsampling PD processing with a step factor of s=2 on the input noisy data to form four discrete sub-block data, and feed them into the denoising model; Step 4.3: The denoised sub-block data after processing is output by the noise prediction module of the denoising model and the self-supervised blind spot network BSN, and then each sub-block is subjected to an inverse PD operation according to the mosaic arrangement during the original PD processing, and is merged into the final denoised seismic data of the full size.
5. A self-supervised seismic denoising system integrating noise estimation module and channel attention mechanism, characterized in that: include: A seismic data set construction module includes slicing real noisy seismic data and synthetic noisy seismic data to obtain multiple noisy data blocks; Self-supervised denoising network module, including: A noise prediction module, composed of multiple 1×1 convolutional layers and ReLU activation functions, is used to output a noise level prediction map with the same size as the input data, wherein the noise level prediction map represents the noise covariance matrix of each pixel; Self-supervised blind spot network, including a dual-branch structure with integrated channel attention mechanism: The first branch uses a dilated convolution module with center mask convolution, referred to as the DC module, to extract local features; The second branch uses a dilated Transformer block combined with a channel attention mechanism, referred to as a DTB module. The DTB module implements global feature modeling through a channel attention mechanism, and its feed-forward layer contains a depthwise separable convolution. The data processing and training unit applies a pixel random downsampling (PD) operation to the input data to destroy the spatial correlation of the noise, and inputs the processed data into the self-supervised network for training, fuses the output of the noise prediction module with the intermediate features of the self-supervised blind spot network using the Bayesian prediction formula, and optimizes the network parameters through back propagation to obtain a denoising model; The denoising and recovery unit inputs the seismic data to be denoised into the trained model for denoising, and recovers the final denoising result through the inverse PD operation.
6. The system according to claim 5, characterized in that The implementation of the self-supervised denoising network module includes the following steps: Step 2.1: Construct a noise prediction module, which includes five 1x1 convolutional layers connected in sequence. The first four layers are followed by a ReLU nonlinear activation function, and the fifth layer outputs a noise covariance matrix of the same size as the input noisy data. Step 2.2: Construct a dual-branch architecture of the self-supervised blind spot network BSN, where The first branch includes the upper and lower paths: Upper path: includes a 3×3 center mask convolutional layer and 9 DC modules in series. The DC module here uses dilated convolution with a stride of 2 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolutional layer and a 3×3 dilated convolutional layer to extract local features. Lower path: It contains a DC module with a 5×5 center mask convolution layer and 9 dilated convolution layers with S=3. The DC module here uses dilated convolution with a stride of 3 and contains parallel upper / middle / lower paths. Each path is composed of a 1×1 convolution layer and a 3×3 dilated convolution layer to extract local features. The second branch includes a 5×5 center mask convolution layer and 9 DTB modules connected in series. The DTB module generates a global feature map of channel interaction through a channel attention mechanism, and uses deep separable convolution and feedforward network to enhance long-range dependencies. The channel attention mechanism is implemented by calculating the dot product operation of the query matrix Q, the key matrix K and the value matrix V. The outputs of the first and second branches are concatenated after removing the blind spot information, and multi-scale features are fused through a 1×1 convolution layer. At the same time, the output of the noise prediction module is Bayesian predicted with the fused features to form the final BSN output result.
7. The system according to claim 5, characterized in that The implementation of the data processing and training unit includes the following steps: Step 3.1: Perform pixel random downsampling PD processing with a step factor of s=2 on the input real noisy seismic data, split the input data into 2×2 sub-blocks and arrange them in mosaic, and generate four sub-block data to break the noise spatial correlation; Step 3.2: Input the PD-processed sub-block data into the noise prediction module to obtain a 1×1 noise covariance matrix for each pixel, and expand the matrix dimension to match the input data size; Step 3.3: Feed the expanded noise covariance matrix and the sub-block data processed by PD into the self-supervised blind spot network BSN; Step 3.4: Calculate the network prediction error based on the Bayesian formula, which is: Explanation of symbols: y: noisy image input, the target to be denoised; E[·]: average over all training data; μ: clean signal predicted by the network, corresponding to D-BSN output; ∑ μ : covariance matrix of the predicted signal; ∑ n : The noise covariance matrix output by the noise prediction module; tr(·): matrix trace operation, constraining Σ μ The amplitude of Step 3.5: Optimize the noise prediction module and BSN network parameters through the back propagation algorithm, obtain the trained denoising model weights and save them, and finally obtain the trained model.
8. The system according to claim 5, characterized in that The implementation of the denoising and restoration unit includes the following steps: Step 4.1: Load the denoising model weights trained in claim 3, and input the noisy seismic data to be processed into the model; Step 4.2: Perform pixel random downsampling PD processing with a step factor of s=2 on the input noisy data to form four discrete sub-block data, and feed them into the denoising model; Step 4.3: The denoised sub-block data after processing is output by the noise prediction module of the denoising model and the self-supervised blind spot network BSN, and then each sub-block is subjected to an inverse PD operation according to the mosaic arrangement during the original PD processing, and is merged into the final denoised seismic data of the full size.
Citation Information
Patent Citations
Earthquake random noise suppression method based on self-attention convolution auto-encoder
CN116559945A
Seismic data denoising method fusing self-attention and Mama architecture
CN118295029A
DAS-VSP data coherent noise suppression method and system based on adaptive texture analysis
CN119395757A
Cited By
Magnetotelluric signal noise suppression method and system based on complexity driving
CN120610324A
Seismic data denoising method based on fuzzy rule and mutual information space-frequency domain decoupling
CN120847878A