Low frequency continuation method for seismic exploration data based on transformer
By using Transformer networks to extend seismic exploration data at low frequencies and employing a self-attention mechanism for global feature modeling, the problem of missing low-frequency data in full waveform inversion is solved, thereby improving prediction accuracy and generalization while reducing computational costs.
Patent Information
- Application Number
- CN202410767622.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-14
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-06-14
AI Technical Summary
Existing full-waveform inversion methods are prone to getting trapped in local minima when low-frequency data is missing, which leads to a decrease in the accuracy of seismic signal processing and interpretation. Furthermore, low-frequency extension methods based on convolutional neural networks are difficult to utilize the global information of long-wavelength time series signals, which limits the prediction accuracy and generalization.
We employ Transformer networks to extend low-frequency seismic exploration data, utilize self-attention mechanisms for global feature modeling, and combine the SEAM velocity model training dataset to design multi-head self-attention, window Transformer modules, and sliding window Transformer modules to fit the high-dimensional nonlinear mapping relationship of low-frequency data.
It improves the accuracy and generalization of full waveform inversion, reduces computational costs, and enables adaptive and accurate prediction of missing low-frequency data.
Smart Images

Figure CN118859317B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of seismic exploration and artificial intelligence, and particularly relates to a low-frequency continuation method for seismic exploration data of a Transformer. BACKGROUND
[0002] Oil and gas resources are an important material basis for the development of modern society, the "blood vessels" of industrial economic operation, and play an irreplaceable role in ensuring national energy security, meeting production and living needs, and achieving high-quality economic and social development. At present, China's oil and gas resources have a high degree of dependence on foreign countries, and the degree of dependence on foreign oil and gas will remain on the rise for a long time. At the same time, some oilfields developed earlier in China are also facing problems such as declining production, declining resource grade, and rising acquisition costs, while unconventional oil and gas resource areas with development potential often have complex environments and high acquisition costs. Therefore, improving the level of oil and gas resource development to increase production and reduce dependence on imports has important research value.
[0003] Seismic exploration is the most important method in geophysical exploration and the most effective method for solving oil and gas exploration problems. It has been widely used in the world's oil and gas exploration due to its high precision, high resolution, large detection depth, and rich information. With the growth of China's economic level, the demand for oil and gas resources and mineral resources is also increasing, and seismic exploration is gradually developing in multiple dimensions and large depths. However, the complex exploration environment limits the precision of underground velocity modeling, which seriously affects the seismic data static correction, velocity analysis, and final imaging results in the exploration area, and further affects the interpreter's determination of low buffer structures, positioning of stratum collapse, and identification of fine oil reservoir structures. Therefore, how to improve the precision of underground velocity modeling will be a key issue in complex stratum imaging processing, and accurate reconstruction of stratum velocity models is of great significance for mineral exploration, engineering geophysical exploration, and improving imaging precision in complex areas and increasing oil and gas production.
[0004] Currently, velocity analysis, tomography, and full waveform inversion are the main methods for physical property inversion using collected seismic data. Among them, full waveform inversion comprehensively utilizes the kinematic and dynamic characteristics of seismic wave fields, and is currently recognized as the most precise inversion method in the field of seismic exploration. Full waveform inversion technology fully utilizes the seismic wave propagation information with physical significance when constructing a velocity model, thereby achieving high-precision imaging of complex stratum models. Specifically, full waveform inversion technology constructs an initial model and iteratively updates the earth parameter model (generally, P-wave and S-wave velocities and density) by continuously comparing predicted seismic data with observed seismic data. When the predicted data and observed data are close enough, the earth parameter model at this time is considered to be close enough to the true situation of the earth's interior.
[0005] However, full waveform inversion is strongly nonlinear and extremely dependent on initial conditions, such as initial velocity model and low-frequency components of seismic data. In the actual seismic data acquisition process, due to the filtering effect of the earth, the limitations of artificial sources, and the distortion of geophones, the collected seismic data usually loses a lot of low-frequency components, making the full waveform inversion result easily fall into local minimum, greatly reducing the accuracy of seismic signal processing and interpretation, and thus seriously limiting the practical application of full waveform inversion technology. Based on this, a high-precision and strong generalization seismic exploration data low-frequency continuation method is proposed to alleviate the period jump phenomenon caused by the lack of low-frequency information in full waveform inversion, which has important research significance for promoting the development of deep high-precision inversion imaging.
[0006] In recent years, a large number of scholars have proposed a lot of methods to predict low-frequency data from seismic data with missing low-frequency information to alleviate the period jump phenomenon and improve the inversion accuracy. Shin and Cha (2008) proposed a method to use the low-frequency components of damped wave fields to invert the data in the Laplace-Fourier domain. Wu et al. (2014) used the super-low frequency signals contained in the signal envelope to predict the low-frequency information, realizing the construction of large-scale velocity model. Hu et al. (2017) proposed a waveform mode decomposition method to reconstruct the low-frequency information with physical meaning. Zhang et al. (2017) used blind deconvolution to obtain the reflection impulse response, and convolved it with the pre-set low-frequency wavelet to obtain the corresponding low-frequency component. Li et al. (2018) used a nonlinear smoothing operator to obtain the low-frequency component. Wang et al. (2018) based on the frequency shift theory shifted the high-frequency data to low frequency to realize the prediction of low-frequency component. However, the low-frequency information predicted by the above methods is often based on the nonlinear transformation of the missing low-frequency data, which may not achieve the expected accuracy, and sometimes introduces additional noise.
[0007] Deep learning can fit various complex functions by constructing high-dimensional nonlinear mapping. With the rapid development of computing device performance, deep learning methods have been gradually applied in various fields, and some deep learning-based seismic exploration data low-frequency continuation methods have also been proposed. Sun and Demanet (2020) constructed a training set using a sub-model of the Marmousi velocity model, designed a network based on one-dimensional convolution, and used a training scheme that uses low-frequency information missing seismic traces to predict seismic traces with low-frequency information. The proposed method verifies the effectiveness of the proposed method in predicting low-frequency signals on the Marmousi and BP 2004 benchmark models. Luo et al. (2023) tried to extrapolate based on single-shot data by designing a multi-scale network to obtain features from seismic records at different resolutions. Hu et al. (2019) introduced the ringing information into the original phase data and used the transfer learning method to predict the low-frequency information from the input data. The above methods have achieved certain results in the low-frequency continuation of seismic exploration data, proving the feasibility of deep learning methods for low-frequency data prediction. However, current deep learning methods usually rely on using convolutional layers to extract local features, and it is difficult to utilize the global information correlation in long-wavelength time-series signals such as low-frequency signals. At the same time, when facing long-sampling-time seismic data, the convolutional neural network needs to constantly increase the depth to ensure that enough receptive fields are obtained, thereby bringing greater computational cost.
[0008] In summary, full waveform inversion is currently the most accurate underground velocity modeling method, but the lack of low-frequency data can cause cycle skipping, thereby reducing the inversion accuracy. Currently, low-frequency continuation methods based on convolutional neural networks all construct mappings by extracting local features, lacking the use of long-distance information and global features, thereby limiting the prediction accuracy. SUMMARY
[0009] The purpose of the present application is to provide a seismic exploration data low-frequency continuation method with high precision and high generalization based on the Transformer network, to solve the problem of adaptive and accurate prediction of low-frequency data and further improve the full waveform inversion accuracy. This method considers the long-wavelength and long-time characteristics of seismic exploration data, uses the self-attention mechanism in the Transformer network to realize global modeling of features, and uses a simulated seismic data training set based on the public SEAM velocity model to train the network to fit the high-dimensional nonlinear mapping relationship between the missing low-frequency data and the low-frequency data, thereby realizing the adaptive extrapolation of missing low-frequency seismic exploration data and improving the accuracy of full waveform inversion using the predicted low-frequency information.
[0010] The purpose of the present application is achieved by the following technical solutions:
[0011] A low-frequency continuation method for seismic exploration data based on a Transformer, comprising the following steps:
[0012] a. Construction of a training data set:
[0013] a1. Forward modeling is performed using the SEAM velocity model to generate simulated records for training, and the velocity model is divided into several sub-models according to regions and stratigraphic complexity;
[0014] a2. The resulting sub-models are expanded to the same size as the complete velocity model by interpolation, and a water layer is added to each model;
[0015] a3. The sub-models are sequentially forward modeled using a missing low-frequency Ricker wavelet and a complete-band Ricker wavelet to obtain simulated records;
[0016] a4. The forward modeling data of the missing low-frequency Ricker wavelet are used as high-frequency components, and the low-frequency components are obtained by low-pass filtering the complete-band Ricker wavelet;
[0017] a5. The high-frequency components and the corresponding low-frequency components are split into one-dimensional time series to obtain low-frequency missing data for inputting into the network and low-frequency data for use as labels;
[0018] b. Construction of a test data set:
[0019] The public Marmousi velocity model is used, and the same forward modeling parameters as in step a, such as seismic wavelet, sampling duration, sampling interval, and grid spacing, are set to perform forward modeling to obtain a test data set;
[0020] c. Design a Transformer network for low-frequency continuation based on data characteristics;
[0021] c1. Design a multi-head self-attention module;
[0022] c2. Design a Transformer module;
[0023] c3. Design a window Transformer module and a sliding window Transformer module;
[0024] c4. Design the overall network structure;
[0025] d. Set the network training hyperparameters, and use mean square error as the loss function;
[0026] e. Input the low-frequency missing data into the network for training, calculate the loss function between the output result and the true low-frequency component, update the network parameters using the Adam optimizer to minimize the loss function, and end the training when the loss tends to be stable;
[0027] f, the trained model is used for low frequency information prediction, and test data constructed by using the Marmousi model is input into the trained Transformer model to verify the low frequency continuation effect and generalization of the proposed method.
[0028] Further, the step c1 specifically comprises the following steps:
[0029] c11, the feature map input into the module is evenly divided along the channel dimension according to the set number of attention heads, and the features in each attention head are respectively adjusted in channel number by three linear transformations to obtain Q, K and V three groups of feature maps;
[0030] c12, the transpose matrix of Q is multiplied by K, and the self-correlation coefficient matrix of the feature is generated by the Softmax function;
[0031] c13, the self-correlation coefficient matrix is multiplied by V to obtain the result of the enhanced effective feature;
[0032] c14, the obtained result is adjusted in channel number by linear transformation and added to the original input data to obtain the output feature of a single attention head;
[0033] c15, the output features of all attention heads are combined along the channel dimension to obtain the output of the multi-head self-attention module.
[0034] Further, the step c2 is specifically: first, the input is processed by layer normalization and multi-head attention to obtain the feature map, which is added to the input to realize the effect of residual connection, and then the feature map is processed by layer normalization and multi-layer perception to obtain the output of the Transformer module.
[0035] Further, the step c3 is specifically: the window Transformer module divides the input sequence feature into equal-length sub-sequences using non-overlapping equal-length time windows based on the Transformer module, and processes each sub-sequence using the Transformer module; the sliding window Transformer module shifts the time window to obtain the features of adjacent regions based on the window Transformer module.
[0036] Further, the step c4 is specifically: the input layer is a convolutional layer for expanding the channel number, the output layer is a convolutional layer for restoring the channel number, and the middle is composed of alternately stacked window Transformer modules and sliding window Transformer modules.
[0037] Compared with the prior art, the beneficial effects of the present application are:
[0038] 1. Compared with the classic deep learning model U-Net, the application aims at the key role of global features in the low-frequency continuation of the time-series-based seismic signal, selects the Transformer network model, uses the self-attention mechanism to model the global features to enhance the useful features concerned and suppress irrelevant information, and thus effectively improves the prediction accuracy and generalization;
[0039] 2. The convolutional neural network generally acquires a larger receptive field by increasing the depth to model long-distance information, but when facing long-sampling-time seismic data, the increasing depth will not only bring huge calculation cost, but also cause the network to fall into overfitting and other problems. Although downsampling can alleviate this problem to a certain extent, downsampling will cause information loss and damage to the frequency spectrum of the data, thereby affecting the prediction accuracy. In view of the problem of high calculation cost of the convolutional neural network when facing long-sequence data, the application proposes a windowed Transformer network structure, which effectively reduces the calculation cost under the premise of ensuring the accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0041] Figure 1 Multi-head attention structure diagram;
[0042] Figure 2 Transformer network structure diagram;
[0043] Figure 3 Selection of 9 sub-models for training set generation;
[0044] Figure 4a Time-domain waveform of the Ricker wavelet for generating training set input data;
[0045] Figure 4b Spectrum of the Ricker wavelet for generating training set input data;
[0046] Figure 5a Time-domain waveform of the full-band Ricker wavelet for generating training set label data;
[0047] Figure 5b Spectrum of the full-band Ricker wavelet for generating training set label data;
[0048] Figure 6aComparison of the single-channel prediction result of the test set with the time-domain waveform of the real low-frequency component;
[0049] Figure 6b Comparison of the single-channel prediction result of the test set with the spectrum of the real low-frequency component;
[0050] Figure 7a Comparison of the single-channel prediction result of the test set with the time-domain waveform of the real low-frequency component;
[0051] Figure 7b Comparison of the single-channel prediction result of the test set with the F-K domain of the real low-frequency component. DETAILED DESCRIPTION
[0052] The application will be further described below in conjunction with the embodiments:
[0053] The application will be further described below in conjunction with the embodiments:
[0054] It should be noted that similar reference numerals and letters refer to similar items throughout the accompanying drawings, and therefore, once an item is defined in one drawing, it need not be further defined and explained in subsequent drawings. Meanwhile, in the description of the application, the terms "first", "second", and the like are merely used to distinguish descriptions, and cannot be understood as indicating or implying relative importance.
[0055] Low-frequency signals in seismic data play an important role. Low-frequency signals can maintain a wide effective frequency band, thereby suppressing the influence of wavelet sidelobes and improving the resolution of seismic data. In the process of seismic wave propagation, low-frequency components attenuate slowly and have strong penetration ability, which can improve the imaging quality of deep strata. In fluid detection, "low-frequency ghosting" can be used as a basis for hydrocarbon indication and for oil and gas reservoir detection. In inversion, low-frequency signals are beneficial to initial velocity modeling, thereby avoiding period jumping and improving the accuracy of inversion. Therefore, in the process of seismic data processing, attention should be paid to protecting effective low-frequency signals in the data to form a data volume containing rich low-frequency signals for subsequent data interpretation work. However, in the process of seismic exploration, the actual acquired seismic exploration record lacks sufficient low-frequency information due to the high main frequency of the source signal and the insufficient sensitivity of the geophone to low-frequency signals. At the same time, low-frequency noise such as surface waves commonly existing in seismic records can seriously reduce the signal-to-noise ratio of low-frequency data, thereby limiting the role of low-frequency signals. Therefore, the application realizes the fitting of the complex nonlinear mapping relationship between the missing low-frequency data and the low-frequency data by using the powerful global modeling capability of the Transformer, thereby realizing the adaptive prediction of the low-frequency data.
[0056] The application is based on a low-frequency continuation method for seismic exploration data based on a Transformer, comprising the following steps:
[0057] 1. Build a Python 3.9 runtime environment and install the deep learning library Pytorch 1.10.2 version.
[0058] The code of the application is implemented on the Python language compiler Pycharm, on which the deep learning toolkit Pytorch version 1.10.2 and the Transformer toolkit Timm version 0.6.13 need to be installed, and the devices for deep learning calculation mainly consist of two RTX 3090 graphics cards and a CPU with a model number of Intel Xeon Silver 4210R Processor 13.75M Cache 2.40GHz.
[0059] 2. Design the input and output data form of neural network training. Considering the feature distribution of seismic data and the calculation cost, the input data and output data are set as one-dimensional time series, wherein the dimension of the input data and the output data can be represented as [B, C, L], wherein B represents the number of sequences contained in a batch of input network during the training process, C represents the channel number, which is set to 1 in the input layer and the output layer, and is expanded to more channel numbers in the middle module to obtain more rich features, and L represents the time length of a single one-dimensional time series.
[0060] 3. Construction of training data set. The quality of the training set will greatly affect the training effect of the model. In order to make the training data set closer to the characteristics of the real exploration seismic data, thereby improving the fitting ability and generalization of the model, the application generates a training data set using the classic SEAM velocity model in the field of seismic exploration. A window with a length and width of 1 / 3 of the complete SEAM velocity model is used to non-repeatedly intercept the velocity model according to the region and the complexity of the stratum to obtain 9 simply structured sub-models; and the forward related parameters are set according to the requirements, including the seismic wavelet main frequency f, the model size nz x nx, the grid distance dx, dz, the total sampling time T, and the time sampling interval dt. The obtained sub-model is expanded to the same size as the complete velocity model by interpolation, and a water layer is added to each model. The missing low-frequency Ricker wavelet and the complete frequency band Ricker wavelet are used to simulate the source signal to excite the sub-model to obtain the forward simulation record, wherein the simulation record obtained by the missing low-frequency Ricker wavelet is used as the data of the input network, and the simulation record obtained by the complete frequency band Ricker wavelet is used as the low-frequency component label for training the network after low-pass filtering.
[0061] Specifically, a training data set size grid size of 120x384 is generated; 9 sub-models are uniformly and non-overlappingly intercepted on the SEAM velocity model using a 40x128 grid size window, as shown in Figure 3 The resulting sub-models are converted to the same size grid number as the standard SEAM model by linear interpolation, and a 10-grid-thick water layer is added above each model; the observation system is set to 384 seismic records at each grid point on the water surface as a geophone, and 96 single-shot records are uniformly arranged on the water surface for each model; the specific forward parameters are set as follows: the boundary condition PML is 20, the grid spacing is 20 m, the sampling interval is 0.002 s, and the total sampling time is 3.2 s; the data used for input network is generated using a Ricker wavelet with a main frequency of 7 Hz and a 5 Hz Butterworth high-pass filter, as shown in Figure 4a , Figure 4b The forward record using a full-band Ricker wavelet with a main frequency of 7 Hz as the source signal is filtered by a 5 Hz Butterworth low-pass filter to serve as the label data for the training network.
[0062] 4. Construction of test data set. The classic complete Marmousi velocity model in the field of seismic exploration is used, and the same forward parameters as in step 3 are set, including the same seismic wavelet, sampling duration, sampling interval, and grid spacing.
[0063] 5. Design of multi-head self-attention module. The feature map input into the module is evenly divided along the channel dimension according to the set number of attention heads, and the features within each attention head are adjusted in channel number by three linear transformations to obtain Q, K, and V three groups of feature maps; then, the transpose matrix of Q is multiplied by K to obtain the self-correlation coefficient matrix of the features, and the self-correlation coefficient matrix is multiplied by V to obtain the enhanced result of the effective features, and the obtained result is adjusted in channel number by linear transformation and added to the original input data to obtain the output features of a single attention head; finally, the output features of all attention heads are combined along the channel dimension to obtain the output of the multi-head self-attention module.
[0064] Specifically, first, the feature map input into the module with a size of [16, 64, 2400] is evenly divided into 8 parts along the channel dimension, i.e. [16, 8, 2400], and is input into 8 attention heads respectively, and the features within each attention head are adjusted in channel number by three linear transformations to obtain Q, K, and V three groups of feature maps, then the transpose matrix of Q is multiplied by K to obtain the self-correlation coefficient matrix reflecting the probability distribution of the features, and the self-correlation coefficient matrix is multiplied by V to realize the enhancement of useful features and the suppression of irrelevant features:
[0065] (Q, K, V) = input x (W Q ,W K ,W V )
[0066] Self-attention(Q, K, V) = Softmax(QK T )V
[0067] where W Q ,W K ,W V are three trainable matrices for obtaining linear transformation results Q, K and V of the input, K T represents the transpose matrix of the linear transformation result K. The resulting feature is adjusted in channel number by linear transformation and added to the original input data to obtain the output feature of a single attention head. Finally, the output features of all attention heads are merged along the channel dimension to obtain the output of the multi-head self-attention module.
[0068] 6、Design the Transformer module. As shown in Figure 2 , the Transformer module is composed of a multi-head attention module, a multi-layer perception module and two layer normalization layers. First, the input is processed by layer normalization and multi-head attention to obtain a feature map, which is added to the input to realize the effect of residual connection, and then processed by layer normalization and multi-layer perception to obtain a feature map, which is added to the output of the first residual connection to obtain the output of the Transformer module. The multi-layer perception module is a feedforward neural network, which includes an input layer, a hidden layer and an output layer. The input layer expands the channel number of the input data by 4 times by linear transformation, i.e. inputs the hidden layer, and the features in the hidden layer are nonlinearly transformed by an activation function before entering the output layer. The output layer restores the channel number to the same as the input by linear transformation. The multi-layer perception introduces nonlinear characteristics into the model, thereby enhancing the model's fitting ability for complex mapping relationships. The layer normalization layer normalizes the input to improve the stability and training effect of the model.
[0069] 7、Design the window Transformer module and the sliding window Transformer module. The window Transformer module divides the input sequence feature into equal-length sub-sequences using non-overlapping equal-length time windows based on the Transformer module, and processes each sub-sequence using the Transformer module. Specifically, the input sequence feature is divided into 60 equal-length sub-sequences using non-overlapping time windows with a length of 40. The sliding window Transformer module translates the time window based on the window Transformer module, specifically by half the size of the distance to obtain the features of adjacent regions, thereby further improving the prediction accuracy.
[0070] 8. Design the overall network structure. As shown in Figure 2 Figure, the input layer is a convolutional layer for expanding the number of channels, and the output layer is a convolutional layer for restoring the number of channels, and the middle is composed of alternately stacked window Transformer modules and sliding window Transformer modules.
[0071] Specifically, the input layer is a convolutional layer with an input channel number of 1 and an output channel number of 64, which serves to expand the channels so that the network can obtain more rich features. The output layer is a convolutional layer with an input channel number of 64 and an output channel number of 1, which aims to integrate the features in different channels obtained, and the middle is composed of alternately stacked window Transformer modules and sliding window Transformer modules, each of which has 4 layers.
[0072] 9. Set the network training hyperparameters, randomly select 16000 pairs of sequence data from the training set for training in each training period, set the number of sequences input into the network in each training batch to 16, set the initial learning rate to 0.001, and reduce the learning rate to 0.3 times of the original every 20 training periods, end the training after 100 periods, use mean square error as the loss function, and the loss function L MSE can be expressed as:
[0073]
[0074] where N represents the number of batches of input data; x j represents the input data; f(x j ) represents the prediction result of the model; y j represents the label value; j represents the serial number of the input data and the corresponding label, and its value range is all integers in the interval [1, N].
[0075] 10. Input the training data into the network for training, and use the Adam optimizer to update the network parameters. After multiple training, the model with the lowest training loss and the best fitting effect is taken as the final training result. That is, when the training loss tends to be stable, the training is ended, and the obtained model is taken as the final training result.
[0076] 11. Input the test data constructed by the Marmousi model into the trained Transformer model to verify the low-frequency continuation effect and generalization of the proposed method.
[0077] In order to improve the prediction accuracy of low-frequency information, the application uses a Transformer model to globally model the data to fit the mapping relationship between the missing low-frequency component data and the low-frequency data. Compared with the classical deep learning method, this method improves the utilization of global information, thereby effectively increasing the prediction accuracy. At the same time, the use of window self-attention mechanism effectively saves the calculation cost, so that the method is also superior to the classical deep learning method in calculation efficiency.
[0078] Embodiment 1
[0079] The embodiment provides an application of a low-frequency continuation method for seismic exploration data based on a Transformer to seismic exploration data with missing effective low-frequency components, which is as follows:
[0080] The test set data comes from the forward of the standard Marmousi velocity model, and 48 single-shot records are generated by uniformly arranging the shot points, each shot set containing 192 seismic traces, the total sampling time being 3.2s, the sampling interval being 0.002s, and the input data frequency band width being 5-50Hz.
[0081] Table 1 Test set data parameters
[0082] Velocity model Shot number Trace number Sampling time Sampling interval Input frequency band Tag frequency band Standard Marmousi 48 192 3.2s 0.002s 5-50 Hz 0-5 Hz
[0083] The data with missing low-frequency components is input into the trained Transformer model to obtain the predicted low-frequency components. As shown in Figure 6a 、 6b , the output data and the label have a very high similarity in the time domain waveform and 0-5Hz amplitude spectrum, the prediction value and the label error are extremely small, the prediction accuracy is high, and the missing low-frequency high-frequency data can be used to accurately predict the low-frequency component. As shown in Figure 7a and Figure 7b , the prediction results of the proposed method for single-shot data are highly similar to the single-shot label data in the time domain two-dimensional data and the F-K spectrum, and have good correlation in the horizontal direction, proving that by dividing the single-shot data into multiple one-dimensional time series signals for sequential prediction, the good lateral continuity can still be maintained, thereby realizing accurate recovery of the effective low-frequency information.
[0084] Note that the above merely describes preferred embodiments of the present application and the principles of the technology applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, modifications and substitutions can be made without departing from the scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.
Claims
1. A low-frequency continuation method based on a Transformer seismic exploration data, characterized in that, The method comprises the following steps: a, construction of a training data set: a1, using a SEAM velocity model to generate simulation records for training by forward modeling, dividing the velocity model into several sub-models according to regions and stratum complexity; a2, expanding the obtained sub-models to the same size as the complete velocity model by interpolation, and adding a water layer to each model; a3, using a missing low-frequency Ricker wavelet and a complete-band Ricker wavelet to sequentially perform forward modeling on the sub-models to obtain simulation records; a4, taking the forward modeling data of the missing low-frequency Ricker wavelet as a high-frequency component, and taking a low-frequency component obtained by low-pass filtering the complete-band Ricker wavelet; a5, splitting the high-frequency component and the corresponding low-frequency component into one-dimensional time series to obtain low-frequency missing data for inputting into the network and low-frequency data for use as labels; b, construction of a test data set: Using a public Marmousi velocity model, and setting the same seismic wavelet, sampling duration, sampling interval and grid spacing forward modeling parameters as in step a to perform forward modeling to obtain a test data set; c, designing a Transformer network for low-frequency continuation in combination with data characteristics; c1, designing a multi-head self-attention module; c2, designing a Transformer module; c3, designing a window Transformer module and a sliding window Transformer module; c4, designing an overall network structure; d, setting network training hyperparameters, and taking mean square error as a loss function; e, inputting low-frequency missing data into the network for training, calculating the loss function between the output result and the real low-frequency component, updating the network parameters using an Adam optimizer to minimize the loss function, and ending the training when the loss tends to be stable; f, using the trained model to predict low-frequency information, and inputting the test data constructed using the Marmousi model into the trained Transformer model to verify the low-frequency continuation effect and generalization of the proposed method.
2. The method of claim 1, wherein the method is based on a Transformer seismic data low-frequency continuation method. Step c1 specifically comprises the following steps: c11, dividing the feature map input into the module according to the set number of attention heads along the channel dimension, adjusting the channel number of the features in each attention head by three linear transformations respectively to obtain Q, K and V three groups of feature maps; c12, multiplying the transpose matrix of Q with K to obtain a self-correlation coefficient matrix of the features through a Softmax function; c13, multiplying the self-correlation coefficient matrix with V to obtain the result of the enhanced effective features; c14, adjusting the channel number of the obtained result by linear transformation and adding it to the original input data to obtain the output features of a single attention head; c15, combining the output features of all attention heads along the channel dimension to obtain the output of the multi-head self-attention module.
3. The method of claim 1, wherein the method is based on a Transformer. Step c2 specifically comprises: first performing layer normalization and multi-head attention processing on the input to obtain a feature map, adding the obtained feature map to the input to realize the effect of residual connection, then performing layer normalization and multi-layer perception processing on the obtained feature map, and adding the obtained feature map to the output of the first residual connection to obtain the output of the Transformer module.
4. The method of claim 1, wherein the method is based on a Transformer seismic data low-frequency continuation method. The step c3 specifically comprises: the window Transformer module divides the input sequence features into equal-length sub-sequences using non-overlapping equal-length time windows on the basis of the Transformer module, and processes each sub-sequence using the Transformer module; and the sliding window Transformer module shifts the time window on the basis of the window Transformer module to obtain features of adjacent regions.
5. The method of claim 1, wherein the method is based on a Transformer seismic data low-frequency continuation method. The step c4 specifically comprises: the input layer is a convolutional layer for expanding the number of channels, the output layer is a convolutional layer for restoring the number of channels, and the middle part is composed of alternately stacked window Transformer modules and sliding window Transformer modules.
Citation Information
Patent Citations
System and method for accurate wellbore placement
CA2692907A1
Multi-scale unsupervised seismic wave velocity inversion method based on observation data self-coding
CN114117906A