VSP shaft wave suppression method based on attention mechanism and multi-scale convolution

By introducing attention mechanism and multi-scale convolution method in VSP data processing, the problem of wellbore wave noise suppression is solved, high signal-to-noise ratio and efficient data processing are achieved, and high-precision seismic imaging under complex geological conditions is supported.

CN120144935APending Publication Date: 2025-06-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510288327.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively suppress wellbore wave noise when processing VSP data, resulting in a reduced signal-to-noise ratio and low data processing efficiency.

Method used

The VSP wellbore wave suppression method based on attention mechanism and multi-scale convolution is adopted to extract feature information of different scales through multi-scale convolution, and the feature channel is weighted and optimized using the adaptive attention channel mechanism to accurately identify and suppress the wellbore wave and its reflected wave.

Benefits of technology

It significantly improves the signal-to-noise ratio, reduces residual noise, improves data processing efficiency, and achieves high-precision seismic imaging under complex geological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144935A_ABST
    Figure CN120144935A_ABST
Patent Text Reader

Abstract

The invention discloses a VSP wellbore wave suppression method based on an attention mechanism and multi-scale convolution. The method comprises the following steps of S1, data preparation and preprocessing; s2, data normalization and data set construction: forming a sample data set by taking noise-containing VSP original data as sample data and wellbore wave noise data as a sample label; s3, designing a U-Net model, and performing model training by using the sample data set; the U-Net model includes an input layer, an encoder, a jump connection, a decoder, and an output layer. The method can accurately identify and effectively suppress wellbore waves and reflected waves thereof in a complex VSP wave field, significantly improve the signal-to-noise ratio, reduce residual noise, and improve the data processing efficiency at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of oil and gas exploration, and particularly relates to a VSP wellbore wave suppression method based on an attention mechanism and multi-scale convolution. Background Art

[0002] Under the current background, with the increasing consumption of global oil and gas resources and the rapid development of the oil and gas exploration field in China, the exploration complexity is continuously increasing, and the demand for high-quality exploration methods is also expanding. In the deep and complex exploration in China, VSP uses sensors in deep underground wells to collect seismic signals, and achieves high-resolution underground structure imaging at a relatively low cost. However, during the acquisition process of VSP data, noise interference will be generated due to differences in device configuration and sensor types. Common regular noise will seriously interfere with the effective signals in seismic data, reduce the signal-to-noise ratio of seismic data, and have a huge impact on the accuracy of subsequent inversion and imaging. Therefore, suppressing strong noise in VSP data and obtaining effective signals with high signal-to-noise ratio have become the research focus in the field of signal processing.

[0003] Wellbore waves are common interfering waves in VSP data, usually a kind of guided wave that propagates along the interface between the fluid in the wellbore and the well wall in multiple groups. During the propagation process, the wellbore waves will be reflected multiple times when reaching the bottom and liquid surface of the well. Its particle vibration trajectory is elliptical, and the frequency is usually high. Since its one-dimensional propagation attenuation is slow, the wellbore wave noise with strong energy causes significant interference to signal processing. Traditional methods mainly rely on the physical mechanism of noise propagation, and extract the parameters of noise for its separation (F-K filtering, τ-p filtering, etc.). The disadvantages are that it is difficult to extract parameters, there is a lot of manual intervention, the efficiency is low, and the actual implementation effect of separating complex wave field signals is limited.

[0004] In recent years, with the development of artificial intelligence technology, deep learning methods can learn the internal laws and representation levels of sample data. At present, artificial intelligence technology has been widely used in seismic signal denoising. Sparse coding or dictionary learning, convolutional neural networks, and hybrid methods of the two have been widely used to suppress random noise. For coherent noise with relatively small feature differences, the method of supervised learning using deep neural networks is mostly used. The above intelligent denoising methods perform well in processing surface seismic signals. The research on VSP noise suppression based on deep learning is mostly in the initial exploration stage, and most deep learning methods do not fully consider the kinematic and dynamic characteristics of noise during the denoising process. Many studies mainly rely on theoretical data to verify the feasibility of the algorithm, lacking in-depth analysis based on actual data. Most studies use synthetic data for network training, and the selection of samples and labels is biased towards idealization, lacking the support of actual work area data. This leads to the fact that in a complex wave field environment, the generalization ability and actual application effect of the model still need to be improved, and it is difficult to achieve the target noise suppression effect in actual work area VSP data. Summary of the Invention

[0005] The object of the present invention is to overcome the deficiencies of the prior art and provide a VSP wellbore wave suppression method based on the attention mechanism and multi-scale convolution. The present invention can accurately identify and effectively suppress wellbore waves and their reflected waves in a complex VSP wave field, significantly improve the signal-to-noise ratio and reduce residual noise, and at the same time improve the data processing efficiency.

[0006] The object of the present invention is achieved by the following technical solutions: A VSP wellbore wave suppression method based on the attention mechanism and multi-scale convolution, comprising the following steps:

[0007] S1. Data preparation and preprocessing, including the following sub-steps:

[0008] S1-1. Adopt a forward simulation method based on a velocity model to generate noise-free VSP data by solving the elastic wave equation.

[0009] S1-2. Generate Gaussian noise by random perturbation and add the Gaussian noise to the VSP data; then, according to the trend and energy relationship of the wellbore wave, generate wellbore wave noise and add the wellbore wave noise to the VSP data to obtain mixed VSP data containing wellbore wave noise and Gaussian noise. At the same time, use the wellbore wave noise as the label of the mixed VSP data.

[0010] S1-3. Obtain the VSP data of the actual work area and perform data preprocessing on the obtained VSP to obtain wellbore wave noise data, and use the obtained wellbore wave noise as the sample label of the VSP data in the actual work area.

[0011] Use the mixed VSP data obtained in S1-2 and the VSP in the actual work area together as the original data, and the corresponding wellbore wave noise as the data label.

[0012] S1-4. Use the method of sliding slicing to process the data and the data label respectively to obtain a plurality of sample data and the sample label corresponding to each sample data.

[0013] S2. Data normalization and dataset construction: Form a sample dataset with each sliced sample and the corresponding sample label; normalize the sample and label data of the entire wave field by the maximum absolute value of the amplitude.

[0014] S3. Design a U-Net model and use the sample dataset to train the model; the U-Net model includes an input layer, an encoder, a decoder, and an output layer.

[0015] Input layer: Convert the input two-dimensional data samples into a four-dimensional data structure, where the dimensions correspond to batch size, number of channels, data height, and width respectively.

[0016] Encoder: Responsible for feature extraction and downsampling of data; multiple consecutive feature extraction modules are designed in the encoder, and each feature extraction module includes a convolutional layer, an SE module, and a max pooling layer; the convolutional layer in the first feature extraction module uses two parallel 5×3 convolutional kernels to extract features from the input data and adjust the number of channels; the convolutional layers in the remaining feature extraction modules use multi-scale convolutional residual modules, and the multi-scale convolutional residual modules select four parallel convolutional kernels with sizes of 3×1, 5×3, 7×5, and 1

[0017] ×1 to extract features from the input data respectively; the features obtained by weighted fusion of the three convolutional kernels of 3×1, 5×3, and 7×5 are added element-wise to the features extracted by the 1×1 convolutional kernel as the output of the multi-scale convolutional residual module.

[0018] After the last max pooling layer, the output features are resized through a multi-scale convolutional residual module and an SE module. To maintain the richness of features and adjust the output channel information, ensure that the subsequent decoder can restore the correct spatial dimensions and channels.

[0019] Decoder: Gradually map the deep features extracted by the encoder back to the spatial resolution of the original input to generate the final output; the decoder includes multiple consecutive feature decoding layers, and the number of feature decoding layers is the same as the number of feature extraction modules; each feature decoding layer sequentially includes upsampling, feature fusion, and feature map concatenation.

[0020] Denote the feature extraction modules and feature decoding layers as 1, 2, …, N from input to output. Then the output features of the 1st to Nth feature extraction modules are fused with the features output by upsampling in the 1st to Nth feature decoding layers respectively, and then input into the feature fusion layer connected to the upsampling.

[0021] The feature decoding layer performs multiple bilinear interpolation upsamplings through upsampling, and the number of channels is reduced through the Conv2d and BatchNorm2d modules after upsampling; feature fusion combines the features of the encoder with skip connections and the features output by upsampling, and the feature fusion is achieved through torch.cat; feature map concatenation further extracts and reconstructs features by applying a multi-scale convolutional residual module to the feature fusion result, and finally outputs after channel attention optimization through the SE module.

[0022] The output layer maps the last feature map in the network to the desired output space. Through a 1×1 convolution operation, the last feature map of the network is mapped to the desired output space and normalized through batch normalization to improve the stability of training and accelerate convergence;

[0023] Adopt MSE, MAE and R 2 The weighted loss of these three parts is used as the overall optimization objective of the U-Net model; after obtaining the total loss function, the gradient of the loss with respect to the model parameters is calculated through the error backpropagation algorithm, and the AdamW optimizer is used to update the model parameters based on the calculated gradient to minimize the loss and improve the model performance, thereby training the deep learning model.

[0024] The beneficial effects of the present invention are as follows: Aiming at the problem of borehole wave interference in VSP data processing, the present invention proposes a VSP borehole wave intelligent suppression method combining an attention mechanism and multi-scale convolution. This method processes the input data using convolution kernels of different sizes (such as 3x1, 5x3, 7x5) respectively to extract feature information of different scales. Through multi-scale feature fusion, the overall architecture enhances the model's perception ability of various details, enabling small-size convolution kernels to extract local waveform features and large-size convolution kernels to extract global kinematic trend features of borehole waves. In addition, in order to improve the model's attention to borehole wave noise in the original noisy data, an adaptive attention channel mechanism is designed to weight and optimize each feature channel, further highlighting the key features of borehole waves. By training the network to learn the mapping relationship between the noisy image and the clean image, the model can effectively restore the details and structure of the noisy image, thereby minimizing the difference between the denoised image and the network output and achieving the purpose of accurate denoising. During this process, the residual learning mechanism helps the model retain the effective signal and accurately suppress the borehole wave noise.

[0025] Through experimental verification of simulated data and actual work area data, the results show that this method can accurately identify and effectively suppress borehole waves and their reflected waves in complex VSP wave fields, significantly improve the signal-to-noise ratio and reduce residual noise, while improving the data processing efficiency. This provides strong technical support for high-precision seismic imaging under complex geological conditions. In the application of a certain well in the actual work area, the significant advantages of this method are further verified:

[0026] (1) Multi-scale feature extraction with multi-size convolution kernels

[0027] The present invention extracts multi-scale features by using multi-size convolution kernels, which can capture local waveform details and global trend features simultaneously. Small-size convolution kernels focus on local waveform changes, while large-size convolution kernels extract global kinematic trends. This multi-scale processing method improves the model's perception ability of different-level features, can accurately identify borehole wave noise, retain the effective signal, and enhance the suppression effect of borehole wave noise.

[0028] (2) High fidelity and high waveform consistency

[0029] The predicted signals of this method have significant advantages in terms of high fidelity and waveform consistency. By learning the characteristics of borehole waves and accurately identifying and extracting them, the network can maximize the preservation of the original characteristics of the signals during the denoising process, ensuring that the waveform after denoising is highly consistent with the real signal. Compared with traditional borehole wave suppression methods, this network not only effectively eliminates interference such as borehole wave noise, but also avoids signal distortion and phase shift, providing high-quality seismic data and reliable support for subsequent signal analysis and imaging. Description of the Drawings

[0030] Figure 1 is the algorithm flowchart of the intelligent suppression method for VSP borehole waves;

[0031] Figure 2 is the schematic diagram of the velocity model;

[0032] Figure 3 is the schematic diagram of two randomly selected synthesized data;

[0033] Figure 4 is the schematic diagram of two randomly selected groups of samples and labels;

[0034] Figure 5 is the schematic diagram of five randomly selected groups of sliced noisy data and borehole wave noise;

[0035] Figure 6 is the network architecture diagram of the U-Net model;

[0036] Figure 7 is the schematic diagram of the multi-scale convolution kernel module;

[0037] Figure 8 is the schematic diagram of the SE attention module;

[0038] Figure 9 is the schematic diagram of the skip connection;

[0039] Figure 10 Schematic diagram of refined processing of samples, borehole wave noise and effective signals and energy statistics;

[0040] Figure 11 is the schematic diagram of the original noisy data, borehole waves and effective signals and energy statistics under the traditional filtering method;

[0041] Figure 12 is the schematic diagram of the U-Net predicted residual data, predicted borehole wave data, predicted effective signals and energy statistics;

[0042] Figure 13This method is for predicting prediction residual data, predicted borehole wave data, schematic diagrams of predicted effective signals, and energy statistics. Specific implementation manners

[0043] An intelligent borehole wave suppression method is proposed. By introducing a multi-scale convolution structure and an attention mechanism, comprehensive extraction and fusion of global kinematic features and local dynamic features are achieved. Specifically, small convolution kernels are used to capture local waveform details, while large convolution kernels are dedicated to extracting global trend features, thereby realizing effective representation of multi-level features. In addition, by introducing an adaptive attention channel mechanism, different feature channels are weighted and optimized, significantly highlighting the key features of borehole waves, ensuring that the detailed information of seismic signals is retained, and achieving accurate identification and suppression of borehole waves and their reflected waves. This design effectively improves the signal-to-noise ratio of the model, reduces noise residues while retaining the high fidelity of the signals. The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0044] As Figure 1 shown, a VSP borehole wave suppression method based on an attention mechanism and multi-scale convolution of the present invention includes the following steps:

[0045] S1. Data preparation and preprocessing, including the following sub-steps:

[0046] S1-1. The diversity of training samples and the accuracy of label data are crucial for the network performance and generalization ability. To enrich the training data set, the present invention adopts a forward simulation method based on a velocity model to generate noise-free VSP data by solving the elastic wave equation; this process ensures the diversity and authenticity of the data, thereby improving the robustness and prediction accuracy of the model.

[0047] In this embodiment, an inclined layered medium containing 6 different velocity layers is constructed (as Figure 2 shown). When simulating VSP data with variable well-source distances, the well is placed at the position of the red frame, the horizontal coordinate is 3500 meters, the depth range is from 400 meters to 3500 meters, the source is located at 5000 meters, a P-wave source with a main frequency of 30 Hz is used, a Ricker wavelet is used for excitation, and 161 geophones are arranged at 20-meter intervals underground. Elastic wave simulation of the underground two-dimensional medium is used for wave field forward modeling to obtain the vibration components of the wave field along the Z direction (vertically downward) and the X direction (horizontally to the right). Figure 3 The component data of two sets of synthetic data (1 and 2) are shown, where a, b, c, and d are the X components of synthetic data 1, the X component of synthetic data 2, the Z component of synthetic data 1, and the Z component of synthetic data 2, respectively.

[0048] Subsequently, based on these synthetic data, noises of different intensities are added, and relevant parameters are adjusted as needed to simulate complex geological conditions and noise interference characteristics.

[0049] S1-2. Generate Gaussian noise by random perturbation and add the Gaussian noise to the VSP data; then, based on the trend and energy relationship of the borehole wave, generate borehole wave noise and add the borehole wave noise to the VSP data at different intensities to obtain mixed VSP data containing borehole wave noise and Gaussian noise. At the same time, take the borehole wave noise as the label of the mixed VSP data. In this embodiment, two selected groups of sample and label results are as Figure 4 shown, where a and b are two groups of synthetic data samples containing borehole wave noise, and c and d are the corresponding borehole wave noise labels respectively.

[0050] S1-3. Obtain the VSP data of the actual work area and perform data preprocessing on the obtained VSP to avoid the influence of bad channels, etc.; in the present invention, the F-K filtering method is used to extract the borehole wave noise data from the VSP data; in order to solve the dispersion problem in the F-K transform domain filtering, the τ-p transform domain filtering is used to solve the dispersion phenomenon, so as to obtain refined borehole wave noise data, and take the obtained borehole wave noise as the sample label of the VSP data in the actual work area;

[0051] Take the mixed VSP data obtained in S1-2 and the VSP data of the actual work area together as the original data, and the corresponding borehole wave noise as the data label; the two types of sample data can not only ensure the data volume but also increase the accuracy of simulating the real situation with actual data.

[0052] S1-4. Use the method of sliding slices to process the data and the data label respectively to obtain multiple sample data and the sample label corresponding to each sample data; considering that the VSP data usually consists of multiple receivers, forming a time series matrix D∈R L×N , where L represents the number of time sampling points, N represents the number of seismic channels, and the time dimension is usually greater than the number of seismic channels. The data is presented as a spatio-temporal two-dimensional matrix, with the time axis as the long axis and the space axis as the short axis, and the lengths of the two are significantly different, and the overall structure presents a narrow and long strip shape.

[0053] In this embodiment, rectangular slices (H×W) are used for operation. Each slice contains H rows (number of time sampling points) and W columns (number of seismic channels). The sliding step size in the X direction is M x and the sliding step size in the Y direction is M y , and the extraction formula for each small block is:

[0054] P i,j =D[j:j + H, i:i + W] (1)

[0055] D[j:j+H, i:i+W] represents a sub-matrix extracted from matrix D, which includes rows from j to j+H-1 and columns from i to i+W-1. By adjusting the sliding step size, more training samples can be generated, thus enhancing the diversity of the data. The schematic diagram of the sliced data is as shown in Figure 5 shown, demonstrating how to extract multiple sub-samples from the original data through different step size settings for further training.

[0056] S2. Data normalization and dataset construction: Form a sample dataset with each sliced sample and its corresponding sample label; normalize the sample and label data of the entire wave field by the maximum absolute value of the amplitude, normalizing the amplitude of the data to the range of [-1, 1]; specifically, this embodiment includes the following sub-steps:

[0057] S2-1. Input data of samples and labels: The sliding slice extracts small regions by gradually moving on the input data, effectively capturing local pixel relationships and spatial features, thereby enhancing the model's recognition ability of local patterns. The network slices the sampling time and the number of seismic traces with a 50% sliding step size, and the overlap degree between each slice is 50%. In the sliced data, the original noisy data serves as the input sample, and the wellbore wave noise data serves as the corresponding label. Each two-dimensional data sample is extended to a four-dimensional structure, where the dimensions are batch size, number of channels, data height, and width respectively. By corresponding the samples and labels one by one and shuffling the order of data loading, the randomness of the training process is ensured, and the data can be efficiently input into the convolutional neural network for training and prediction.

[0058] S2-2. K-Fold cross-validation: K-Fold cross-validation divides the dataset into K non-overlapping subsets. Each time, K-1 subsets are used for training, and the remaining one subset is used as the validation set for model evaluation. This method can more comprehensively evaluate the performance of the model, reduce the data division bias error to improve the generalization ability of the model. In addition, through multiple validations, K-Fold can effectively utilize each part of the dataset, avoid overfitting of the model on the training dataset, and provide more reliable model evaluation results. In the Kth validation, subset D k is used as the validation set, and the remaining D\D k is used as the training set, where D is the total subset.

[0059]

[0060] This process is repeated K times, and each subset is used as the validation set once, ensuring that each sample participates in both training and validation. Finally, the evaluation results of all validation sets are aggregated to obtain the overall performance of the model. In this embodiment, K = 3 is selected, that is, the dataset is divided into a training set (66.67%) and a validation set (33.33%) through K-fold cross-validation (K = 3), split idx is the partitioning index, and the training set and validation set are expressed as:

[0061]

[0062] X[:split idx represents taking the part from index 0 to split idx -1 from the dataset X, that is, the training set, and X[split idx :] represents taking the part from index split idx to N (the total number of indexes in the dataset), that is, the validation set.

[0063] The dataset class forms training pairs by corresponding sample data and label data one by one, and uses DataLoader to load the dataset for training and validation. While shuffling the loading order, the data is loaded according to the batch size (Batch). The batch data is expressed as follows:

[0064]

[0065] Among them, Batch k represents the data of the k-th batch, x high,i and x low,i represent the sample and label data pairs respectively, B represents the size of each batch (Batch size), k represents the index of the batch, that is, the current is the k-th batch, N represents the total number of samples in the dataset, represents the total number of batches and rounds up.

[0066] Data loading order shuffling formula:

[0067] Shuffle((X) = Permute(X, π) π is a randomly permuted index (5)

[0068] Permute(X, π) means performing a random permutation of X given by π, where X is the input data; π is a randomly arranged index sequence used to shuffle the order of X, and Permute(X, π) means rearranging X in the order given by π;

[0069] S2-3. Amplitude maximum normalization: Normalization is to scale the data into a standard interval to ensure that different features have the same dimension, thereby improving the stability and efficiency of training. Seismic data usually contains positive and negative amplitudes, indicating the direction of fluctuations. Normalizing the data to the range [-1, 1] can preserve the positive and negative information of the amplitudes, helping the model to more accurately capture the physical characteristics of the data and focus on its relative changes. Calculate the absolute maximum amplitudes in the full wavefields of the input samples and labels respectively, and use the maximum values to normalize the global data of the samples and labels to the range [-1, 1]. The relative changes of the data are more evenly distributed across the entire value range, and the model can better capture the relative differences. The input data D ∈ R M×N , where N is the number of traces and M is the number of time sampling points. Calculate the absolute maximum amplitudes of the samples / labels according to Equation (6), and the normalization is shown in Equation (7):

[0070] M = max(|D i , k|) (6)

[0071]

[0072] |D i,k | represents the absolute value of the element in the i-th row and k-th column of the matrix D, which is the absolute value of the amplitude in the wavefield. M is its absolute maximum value. If M = 0, return the original wavefield data. If M ≠ 0, normalize each data to obtain the normalized matrix D'. This normalization method can ensure that the relative changes of the wavefield data are evenly distributed in the value range [-1, 1], thereby enhancing the model's ability to capture relative differences.

[0073] S3. Design a U-Net model and use the sample dataset to train the model; The U-Net model includes an input layer, an encoder, skip connections, a decoder, and an output layer; This model takes the original noisy data and wellbore wave noise data as inputs, uses a structure that combines U-Net, an attention mechanism (SE module), and a multi-scale convolutional residual module, uses residual learning, and performs denoising by learning the features of the wellbore wave at different scales and fusing the features. The specific structure is as Figure 6 shown.

[0074] To better extract the global trend of the wellbore wave and the local waveform changes, enhance the model's perception ability of various details through multi-scale feature fusion, and realize the extraction of local waveform features by small-size convolutional kernels and the extraction of global kinematic trend features of the wellbore wave by large-size convolutional kernels. In addition, to improve the model's attention to the wellbore wave noise in the original noisy data, an adaptive attention channel mechanism is used to weight and optimize each feature channel to further highlight the key features of the wellbore wave.

[0075] In order to suppress wellbore noise more accurately, the model learns the mapping relationship between noisy images and clean images, and minimizes the difference between the network output and the actual denoised image during training. Combining these improvements, the model achieves accurate suppression of wellbore noise while retaining the details and structure of the target signal.

[0076] Input layer: The main task of the input layer is to receive the sliced ​​noisy VSP data and the corresponding wellbore wave noise as sample data and corresponding sample labels respectively. The sample and label data are one-to-one corresponding to ensure the validity of the input and the accuracy of the training process. The input two-dimensional data sample is converted into a four-dimensional data structure, where the dimensions correspond to the batch size, the number of channels, the data height, and the width;

[0077] The encoder part is responsible for feature extraction and downsampling of the data. It extracts richer high-level features by gradually reducing the spatial size of the data and increasing the number of channels. Each layer extracts and fuses local and global wellbore wave features through multi-resolution convolution operations to enhance the model's ability to learn waveform features and trend features. In addition, combined with the SE attention mechanism, the model's attention to key features is improved by adaptively adjusting the weights of each channel. The downsampling process helps to retain more spatial information, allowing the encoder to gradually learn more abstract feature expressions from low levels to high levels as it goes deeper layer by layer, providing richer feature maps for the decoding stage.

[0078] In VSP data, seismic records are usually composed of multiple receivers, forming a two-dimensional matrix of time and space, in which the number of signal points in the time dimension is much larger than the spatial dimension, so the time axis is defined as the long axis and the spatial axis is the short axis, and the overall data presents a narrow strip shape. To better adapt to this data feature, this model uses a rectangular convolution kernel to improve feature extraction capabilities, especially in capturing global wellbore wave trends and local waveform details.

[0079] The encoder part is designed with multiple sequential feature extraction modules, each of which includes a convolutional layer, a SE module and a maximum pooling layer; the convolutional layer in the first feature extraction module uses two parallel 5×3 convolution kernels to extract features of the input data and adjust the number of channels; the convolutional layers in the remaining feature extraction modules use multi-scale convolution residual modules, and the multi-scale convolution residual modules use a parallel convolution (Parallel Convolutions) structure.

[0080] like Figure 7As shown, the multi-scale convolutional residual module selects four parallel convolutional kernels with sizes of 3×1, 5×3, 7×5, and 1×1 to extract features from the input data respectively. Different-scale features are extracted through multi-scale convolution. Among them, the small-size convolutional kernel (3×1) is mainly used for learning the dynamic features of local waveforms, the 5×3 convolutional kernel can extract medium-scale features and provide a balance between local details and overall trends, while the large-size convolutional kernel (7×5) is used for learning the kinematic features of global borehole waves. To fully fuse different-scale features, the extraction results of the three convolutional kernels of 3×1, 5×3, and 7×5 are weighted and fused according to the weight ratio of 5:3:1 to enhance the expression ability of multi-scale information. The 1×1 convolution is a standard convolutional sub-kernel that can strengthen global feature learning. The features obtained by weighted fusion of the three convolutional kernels of 3×1, 5×3, and 7×5 are added element-wise to the features extracted by the 1×1 convolutional kernel, which serves as the output of the multi-scale convolutional residual module.

[0081] Figure 7 Among them, ⊕ represents element-wise addition, that is, the input x is directly added to the output to enhance the gradient flow and alleviate the problem of gradient disappearance. The present invention also adds a combination of batch normalization (BN) and hyperbolic tangent activation function (Tanh) (BN+Tanh) after each convolutional kernel to accelerate convergence and improve training stability. By combining multi-scale convolution and standard convolution, this model can more comprehensively extract the spatio-temporal features of VSP data and improve the ability to identify and model borehole waves.

[0082] The present invention uses three different convolutional kernels, which can learn from two aspects: the dynamic features of local waveforms and the kinematic features of the global borehole wave trend, thereby enhancing the model's ability to learn physical features at different scales, and enabling the model to synchronously learn the physical features of borehole waves through feature fusion. During the training process, maintaining the fidelity and similarity of waveforms is the key goal. Therefore, weight ratios of 5:3:1 are assigned to convolutional kernels of different sizes, and after parallel feature extraction, they are fused to further improve the model's ability to express and process multi-scale features.

[0083] To enhance the attention to key features, an SE module (Squeeze and Excitation) is added after each feature extraction module. The structure of the SE Model module is as Figure 8As shown, the SE module extracts global context information through the channel attention mechanism, performs global average pooling to extract global context information, and adjusts the channel weights to enhance the attention to key features. The SE module mainly consists of three parts: Squeeze (global information compression), Excitation (channel weight calculation), and Scale (channel reweighting). The original dimension of the data features is H*W*C. First, spatial feature compression is performed on the feature map. Global information compression mainly uses global average pooling to reduce the dimension of the input feature map, extract global context information, average all spatial positions of each channel, and implement global average pooling in the spatial dimension to obtain a 1*1*C feature map, where C is the number of channels. Through learning with the FC fully connected layer, a feature map with channel attention is obtained, and its dimension is still 1*1*C. The feature map with channel attention of 1*1*C and the original input feature map of H*W*C, F scale are multiplied by the weight coefficient channel by channel to enhance key features, and finally a feature map with channel attention is output. The SE module adaptively adjusts the importance of each channel through the channel attention mechanism, mainly obtaining: global context information, that is, the importance of each channel in the entire feature map. Enhancing the expression ability of key features, such as the main vibration mode of the borehole wave. Suppressing irrelevant information, reducing the influence of noise on the result, and improving the robustness of the model.

[0084] To reduce the size of the feature map to reduce the computational complexity and also retain key information to a certain extent, max pooling (MaxPool2d) is used after each SE Model module to extract the maximum value in each local area of the feature map, thereby effectively compressing the data and retaining the most representative features.

[0085]

[0086] window(i,j) is a 2x2 window centered on (i,j), x(p,q) is the pixel value within the window, and the final size of the feature map is 1 / 2 of the input, and the width and height change as

[0087] After the last max pooling layer, the output features are resized through a multi-scale convolutional residual module and an SE module to maintain the richness of the features and adjust the output channel information to ensure that the subsequent decoder can recover the correct spatial size and channels.

[0088] Skip connection: To effectively retain low-level feature information, the feature maps output by each feature extraction module in the encoder part are fused with the feature maps output by the corresponding upsampling layer in the decoder part through skip connections ( Figure 9) During the encoding stage of the network, the encoder focuses on extracting wellbore wave features from the original noisy data, while the decoder restores the spatial resolution of the image by gradually upsampling. During the decoding process, skip connections fuse the output of the encoder with the upsampling results in the decoder, usually by concatenation or summation, enhancing the ability to express details. This helps slow down the loss of information in the deep network and maintain higher accuracy when restoring the spatial resolution.

[0089] y = F(x) + x (9)

[0090] where x is the input of the skip connection, F(x) is the output calculated through the intermediate layer, and y is the final output containing the original input and the transformed kernel.

[0091] Decoder: The main task of the decoder is to gradually map the deep features extracted by the encoder back to the spatial resolution of the original input, thereby generating the final output; the decoder includes multiple sequentially connected feature decoding layers, and the number of feature decoding layers is the same as the number of feature extraction modules; each feature decoding layer sequentially includes upsampling, feature fusion, and feature map concatenation. The encoder continuously reduces the spatial resolution of the feature map during the feature extraction process, while the task of the decoder is to gradually restore the spatial resolution of the original input. Therefore, the upsampling operation needs to be carried out in multiple steps, and the size of the feature map is increased in each step.

[0092] Denote the feature extraction module and the feature decoding layer as 1, 2, …, N in order from input to output. Then, the output features of the 1st to Nth feature extraction modules are respectively fused with the features output by upsampling in the 1st to Nth feature decoding layers, that is, the skip connections mentioned above, and then input into the feature fusion layer connected to the upsampling.

[0093] The feature decoding layer performs bilinear interpolation upsampling multiple times through upsampling. After each upsampling, a Conv(1×1) convolution group composed of Conv2d and BatchNorm2d modules (such as upconv1, etc.) is used to reduce the number of channels; Feature fusion combines the corresponding feature maps of the encoder with skip connections and the feature maps output by upsampling (skip connections), and the fusion of features is achieved through torch.cat. On the feature fusion result after each upsampling, feature maps are concatenated on the feature fusion result and then multi-scale convolutional residual modules are applied to further extract and reconstruct features. These modules ensure that the decoder can gradually recover the spatial resolution of the data through convolutional operations, while fusing low-level and high-level features. The output of each feature decoding layer is output after channel attention optimization by the SE module. The SE module adaptively adjusts the channel weights according to global information, enhances the expression ability of key features, and reduces the interference of irrelevant information. The specific calculation formula is as follows:

[0094] x concat =Concat(x up , x encoder ) (10)

[0095] x out =Activation(Conv2D(x concat ) (11)

[0096] x SE (c)=x(c).σ(W 2 *ReLU(W 1 *GAP(x))) (12)

[0097] Among them, x concat represents the fused feature map after upsampling and skip connection, x up represents the feature map after upsampling, x encoder represents the feature map of the corresponding layer in the encoder, x out represents the output map of the corresponding layer in the encoder; x SE (c) represents the c-th channel of the feature map after being processed by the SE module, x(c) represents the c-th channel of the input feature map, σ represents the Sigmoid activation function, W 2 represents the weight matrix of the fully connected layer, which is used to map the feature generated by W 1 and GAP operation to the final channel weight, W 1 represents the weight matrix of another fully connected layer, which is usually used to transform the feature map (compressed global feature) obtained by the GAP operation into an intermediate representation, and GAP(x) represents global average pooling.

[0098] The output layer maps the last feature map in the network to the desired output space through a 1×1 convolution operation. At the same time, it normalizes the input data of each layer through a BatchNorm2d (batch normalization) to reduce gradient vanishing and explosion, improve the stability of training and accelerate convergence, and enhance the generalization ability of the model. It converts the multi-channel feature map into the target output noise. Assuming the input feature map is F with a size of B×C×H×W (where C is the number of input channels), the conversion formula of the output layer is, represents the output feature (prediction result) after convolution, W conv is the weight matrix of the convolution kernel, b conv is the bias term of the convolution layer, which is used to adjust the values of each channel of the output feature map.

[0099]

[0100] Finally, the predicted noise signal will be converted into two-dimensional data. At this time, the distribution range of the predicted borehole wave data is in the normalized interval of [-1,1]. To ensure that the model output not only conforms to the data characteristics, in order to restore the original amplitude of the predicted output data, it is necessary to multiply by the maximum amplitude M used in the normalization process, so as to restore the data to its original amplitude range.

[0101] D = D'×M (14)

[0102] D' is the normalized predicted output, M is the absolute maximum amplitude of the corresponding label data in the normalization process, and D is the data after denormalization, that is, the target predicted output is restored to the original amplitude range.

[0103] Loss function design: Use the weighted loss of MSE, MAE, and R 2 as the overall optimization objective of the U-Net model. The calculation methods of the three losses are as follows:

[0104]

[0105]

[0106]

[0107] Among them, NR and NT represent the number of geophones and the number of signal samplings respectively, m represents the m-th sampling point in the time series, E[S(x j , t m ) represents the average value of all data in a single shot gather record, x j represents the spatial position of the j-th geophone, that is, the j-th trace in the VSP data; t m represents the m-th time sampling point, corresponding to the sampling interval; S(x j , tm ) represents the label signal, i.e., the true value of the seismic signal at position x j and time t m ; S pre represents the predicted signal obtained through the model, and the magnitude of this value is close to 0, S pre and S represent the predicted and target gather records. A gather represents the seismic waveform data recorded by all receiving points (geophones) after being triggered by a single seismic source (shot point). The target gather is the amplitude value of the label gather record, and the predicted gather is the amplitude value of the predicted gather record. R 2 reflects the goodness of fit of the prediction, and its value range is (-∞, 1].

[0108] The total loss is as follows:

[0109]

[0110] Weight mse 、Weight mae 、 are the weights of the losses of the three parts of MSE, MAE, and R 2 respectively. After obtaining the total loss function Total loss , the gradient of the loss with respect to the model parameters is calculated through the error backpropagation algorithm, and the AdamW optimizer is used to update the model parameters based on the calculated gradient to minimize the loss and improve the model performance, thereby training the deep learning model. Finally, the trained U-Net model is used for suppressing YSP borehole wave noise.

[0111] In view of the deficiencies in the extraction of the global kinematic trend characteristics and local waveform dynamic characteristics of borehole waves in the present invention, as well as problems such as poor physical feature fusion effect and limited sample quantity, a VSP borehole wave intelligent suppression method based on the attention mechanism and multi-scale feature convolution is proposed for suppressing borehole waves in various VSP data. Based on the original U-Net network architecture, this method combines a multi-scale convolutional network and uses the attention mechanism to assign weights to the extraction and fusion of different features. Through the learning and fusion of multi-scale features, it can automatically complete the accurate identification, extraction, and suppression of borehole waves under the condition of few samples without the need for manual setting of filtering parameters. The specific process of model training is as follows:

[0112] a) Model evaluation: Record the best model in each round of training, evaluate the model using the data of the actual work area. First, normalize the input original data and then call the model. After obtaining the predicted output, use denormalization to return to the original data range, and check its performance and noise suppression effect in the actual scenario based on data comparison and analysis.

[0113] b) Hyperparameter tuning: Adjust hyperparameters such as learning rate, batch size, decay rate, etc., to find the best training configuration, improve model performance, and compare the model training situations under different conditions.

[0114] c) Repeated optimization: According to the results of model evaluation, iterate the training process, continuously optimize hyperparameters and model structure, and improve the applicability and accuracy of the model in the actual work area.

[0115] d) Loss monitoring: During the training process, continuously record the training and test losses and save the losses in a.txt document to analyze the convergence of the model and improve the training strategy.

[0116] e) Unified input format: The input samples (including noisy raw data) and labels (borehole wave noise) of the network should correspond one by one and ensure the same data format. Here, the.npy format is selected to read into the model for training.

[0117] Model generalization verification: To verify the borehole wave suppression effect of the model of the present invention in the data of new work areas, use the optimal model obtained by training to conduct generalization verification on VSP data. That is, to check whether the model can correctly judge whether there are borehole waves and their reflected waves in the wave field in unseen data, and effectively suppress the borehole wave noise on the basis of accurate identification. In addition, by comparing the predicted noise with the effective signal after suppression, evaluate the waveform fidelity and noise suppression level, so as to verify the adaptability and reliability of the model in various VSP data. The experiment also includes tests for different noise intensities and complex wave field conditions to comprehensively analyze the stability and robustness of the model, and ensure that it can meet the actual application requirements and improve the accuracy and efficiency of borehole wave suppression.

[0118] Select a group of untrained VSP data for verification and comparison using the model. Analyze the global trend characteristics and local waveform characteristics of the predicted borehole wave noise in the time domain, and at the same time complete the comparison of numerical transformation by calculating the energy transformation before and after prediction through the square of the amplitude value. As Figure 10 shown, it is a schematic diagram of a selected group of original data (samples) containing borehole waves and borehole reflection wave noise, borehole wave noise data (labels), and effective signal data after suppressing borehole wave noise. The model uses the samples and labels as data inputs, and the effective signal data is the data obtained by subtracting the noise from the samples. The energy of the borehole wave noise obtained using the traditional filtering method (F-K filtering) is 43.34% lower than that of the refined label, and the remaining borehole wave energy residue in the effective signal is large and the noise suppression is not clean. The results are as Figure 11 shown. Using U-Net physics for noise suppression, we can obtain Figure 12From the displayed predicted noise, predicted noise residuals, and effective signal data, it is not difficult to find that the fidelity of predicting borehole wave noise has been improved to a certain extent compared with traditional methods at this time. However, the global feature extraction of borehole reflection waves is poor at this time, and the borehole reflection waves cannot be well identified, showing great limitations.

[0119] The prediction results obtained by the present invention are as Figure 13 shown. Figure 13 It shows the schematic diagram and energy statistics of the residual data predicted by the method of the present invention, the predicted borehole wave data, and the predicted effective signal. After comparing with the residual wave fields and energies predicted by three different methods, it is found that the residual energy obtained by the present invention is weaker, indicating that the fidelity of the borehole wave is higher. In addition, this method shows better performance in the identification, prediction, and suppression of borehole waves and borehole reflection waves. From the analysis results of the energy statistics data, compared with the refined borehole wave noise and effective signals, this method shows higher fidelity when predicting noise and also has a certain ability to reconstruct effective signals. In summary, through comparison with traditional filtering methods and U-Net methods, the present invention performs better in many aspects. It can not only simply and quickly achieve high-fidelity signal extraction and suppression without manual parameter selection, but also has strong robustness and effectiveness in practical applications.

[0120] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A VSP wellbore wave suppression method based on attention mechanism and multi-scale convolution, characterized in that: The following steps are involved: S1. Data preparation and preprocessing, including the following sub-steps: S1-1. Using the forward modeling method based on the velocity model, noise-free VSP data are generated by solving the elastic wave equation; S1-2, using random disturbance to generate Gaussian noise, and adding the Gaussian noise to the VSP data; then generating wellbore wave noise based on the trend and energy relationship of the wellbore wave, and adding the wellbore wave noise to the VSP data, to obtain mixed VSP data containing wellbore wave noise and Gaussian noise, and using the wellbore wave noise as a label for the mixed VSP data; S1-3, obtaining VSP data of the actual work area, and performing data preprocessing on the obtained VSP to obtain wellbore wave noise data, and using the obtained wellbore wave noise as a sample label of the VSP data of the actual work area; The mixed VSP data obtained by S1-2 and the VSP of the actual work area are used as the original data, and the corresponding wellbore wave noise is used as the data label; S1-4, using the sliding slicing method to process the data and data labels respectively, to obtain multiple sample data and sample labels corresponding to each sample data; S2. Data normalization and data set construction: The sample data set is formed by each sliced ​​sample and the corresponding sample label; the sample and label data of the entire wave field are normalized by the maximum absolute value of the amplitude; S3. Design a U-Net model and use a sample data set to train the model. The U-Net model includes an input layer, an encoder, a decoder, and an output layer. Input layer: converts the input two-dimensional data samples into a four-dimensional data structure, where the dimensions correspond to batch size, channel number, data height, and width; Encoder: responsible for data feature extraction and downsampling; the encoder designs multiple sequential feature extraction modules, each of which includes a convolutional layer, a SE module, and a maximum pooling layer; The convolution layer in the first feature extraction module uses two parallel 5×3 convolution kernels to extract features from the input data and adjust the number of channels; The convolution layers in the remaining feature extraction modules use a multi-scale convolution residual module. The multi-scale convolution residual module uses four parallel convolution kernels of sizes 3×1, 5×3, 7×5, and 1×1 to extract features from the input data respectively. The features after weighted fusion of the three convolution kernels 3×1, 5×3, and 7×5 are added element by element to the features extracted by the 1×1 convolution kernel as the output of the multi-scale convolution residual module. After the last maximum pooling layer, a multi-scale convolution residual module and SE module are used to resize the output features. Decoder: maps the deep features extracted by the encoder back to the spatial resolution of the original input step by step to generate the final output; the decoder includes multiple consecutive feature decoding layers, and the number of feature decoding layers is the same as the number of feature extraction modules; Each feature decoding layer includes upsampling, feature fusion, and feature map concatenation in sequence; The feature extraction modules and feature decoding layers are denoted as 1, 2, ..., N from input to output. The output features of the 1st to Nth feature extraction modules are fused with the upsampled output features in the 1st to Nth feature decoding layers, respectively, and then input into the feature fusion layer connected to the upsampled output. The feature decoding layer performs multiple bilinear interpolation upsampling through upsampling, and after upsampling, the channel dimension is reduced through Conv2d and BatchNorm2d modules; feature fusion combines the features of the encoder of the jump connection and the features of the upsampling output, and realizes feature fusion through torch.cat; feature map splicing is applied to the feature fusion result and then the multi-scale convolution residual module is applied to further extract and reconstruct the features, and finally the SE module is used to optimize the channel attention and output; The output layer maps the last feature map in the network to the desired output space through a 1×1 convolution operation and normalizes it through batch normalization; Using MSE, MAE and R 2 The three-part weighted loss is used as the overall optimization goal of the U-Net model. After obtaining the total loss function, the error back propagation algorithm is used to calculate the gradient of the loss to the model parameters, and the AdamW optimizer is used to update the model parameters based on the calculated gradient to minimize the loss and improve the model performance, thereby training the deep learning model.

Citation Information

Cited By

  • Rock mechanical parameter prediction method based on improved convolutional neural network

    CN121350591A