A CNN-based low-frequency extension multi-scale full waveform inversion method in the frequency domain

Through the CNN-based frequency domain low-frequency expansion multi-scale full waveform inversion method, the low-frequency seismic wave field data is predicted using the convolutional neural network model, which solves the problem of insufficient low-frequency components in land seismic data, and accurately obtains and improves the inversion speed model.

CN116299702BActive Publication Date: 2025-07-29INST OF GEOPHYSICAL & GEOCHEMICAL EXPLORATION CHINESE ACAD OF GEOLOGICAL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310207286.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-07
Publication Date
2025-07-29
Estimated Expiration
2043-03-07

AI Technical Summary

Technical Problem

At this stage, there are fewer low-frequency components in land seismic data, which makes it difficult to accurately obtain the velocity model in full waveform inversion, and is prone to falling into local minimum values, affecting the inversion accuracy.

Method used

The CNN-based frequency domain low-frequency expansion multi-scale full-waveform inversion method is adopted, and the convolutional neural network model is constructed, and the low-frequency seismic data is used to predict the low-frequency seismic wave field data, and the multi-scale full-waveform inversion is performed in combination with high- and low-frequency data to generate wide-frequency seismic data.

Benefits of technology

It effectively avoids the multi-scale full waveform inversion into local minimum values, and improves the accuracy and accuracy of the inversion speed model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116299702B_ABST
    Figure CN116299702B_ABST
Patent Text Reader

Abstract

The present invention discloses a frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN, which specifically relates to the field of geophysical technologies. After establishing a formation velocity model for the work area according to prior geological information, the present invention adds random perturbations to the formation velocity model of the work area based on the discrete cosine transform to generate multiple velocity models, and respectively uses each velocity model to perform forward simulation of seismic records in the frequency domain to obtain seismic records at multiple single frequency points. Sample data is generated by slicing the seismic data at each single frequency point, and a constructed convolutional neural network model is trained using the sample data. The trained convolutional neural network model is used to predict low-frequency seismic wavefield data based on high-frequency seismic wavefield data, and the predicted low-frequency seismic wavefield data is combined with the high-frequency seismic wavefield data to form broadband seismic data for multi-scale full waveform inversion. The present invention effectively avoids the problem that multi-scale full waveform inversion is prone to falling into local minima, and realizes the accurate acquisition of the inversion velocity model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geophysical technologies, and particularly relates to a frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN. Background Art

[0002] As an effective tool for obtaining the velocity model of underground media using seismic data, low-frequency data is very important for the convergence of inversion. However, at present, due to the limitations of hardware equipment and acquisition costs, the low-frequency components in actual land seismic data are less, and the effective frequency of conventional land seismic data is often only above 5 Hz. Starting the inversion from 5 Hz frequency seismic data is difficult for the full waveform inversion of seismic data. The inversion is often prone to falling into local minima and cannot obtain the correct velocity model. As an effective tool for solving non-linear relationships, deep learning has been widely applied in the field of seismic exploration from aspects such as data reconstruction, processing, inversion, and interpretation, such as seismic data missing trace filling, seismic first arrival picking, seismic data low-frequency extension, velocity inversion, fracture identification, etc.

[0003] Therefore, it is urgent to apply deep learning to full waveform inversion to broaden the acquisition of low-frequency seismic data to obtain broadband seismic data, and a frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN is proposed. Summary of the Invention

[0004] Aiming at the problem that the velocity model cannot be accurately obtained during full waveform inversion due to the small amount of low-frequency components in actual land seismic data, the present invention proposes a frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN, effectively avoiding the multi-scale full waveform inversion from falling into local minima, realizing the accurate acquisition of the inversion velocity model, and improving the accuracy of multi-scale full waveform velocity inversion.

[0005] The present invention adopts the following technical solutions:

[0006] A frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN specifically includes the following steps:

[0007] Step 1, establish a layer velocity model of the work area according to the prior geological information of the area, add random perturbations to the layer velocity model of the work area based on the discrete cosine transform to generate multiple velocity models, respectively perform forward simulation of seismic records in the frequency domain using each velocity model to generate frequency-domain seismic forward records, obtain seismic records of multiple single-frequency points, and generate sample data by slicing the seismic data of each single-frequency point;

[0008] Step 2, construct a convolutional neural network model based on the deep learning platform;

[0009] Step 3: Input the training set and the validation set into the convolutional neural network model, and adjust the weights and biases of the convolutional neural network model by training the convolutional neural network model to obtain the trained convolutional neural network model;

[0010] Step 4: Use the trained convolutional neural network model for low-frequency prediction. Input the high-frequency seismic wavefield data in the frequency-domain seismic wavefield data into the trained convolutional neural network model, and use the trained convolutional neural network model to predict the low-frequency seismic wavefield data of each single frequency point in the frequency-domain seismic wavefield data;

[0011] Step 5: Use the low-frequency seismic wavefield data of each single frequency point in the predicted frequency-domain seismic wavefield data combined with the high-frequency seismic wavefield data to construct broadband seismic data from low frequency to high frequency, and perform multi-scale full waveform inversion using the broadband seismic data to obtain the multi-scale full waveform inversion result based on low-frequency extension in the frequency domain.

[0012] Preferably, in step 1, the following steps are specifically included:

[0013] Step 1.1: Establish a layer velocity model for the work area according to the prior geological information of the region;

[0014] Step 1.2: Add random perturbations to the layer velocity model of the work area based on the discrete cosine transform to obtain a velocity model with random perturbations added;

[0015] Step 1.3: Repeat step 1.2 multiple times to generate multiple velocity models with random perturbations added;

[0016] Step 1.4: Use each velocity model to perform forward simulation of frequency-domain seismic records to generate frequency-domain seismic forward records, obtain seismic records of multiple single frequency points, divide the single frequency point seismic records into high-frequency parts and low-frequency parts according to frequency, after equally spacing the high-frequency parts and low-frequency parts of the single frequency point seismic records, use the high-frequency parts of the single frequency point seismic records after segmentation as training data to form a training set, and use the low-frequency parts of the single frequency point seismic records after segmentation as label data to form a validation set.

[0017] Preferably, in step 1.2, the following steps are specifically included:

[0018] Step 1.2.1: Perform discrete cosine transform on the layer velocity model of the work area to obtain a DCT matrix, as shown in formula (1):

[0019]

[0020] Among them,

[0021]

[0022] In the formula, DCT klis the layer velocity model of the work area after discrete cosine transform; c k and c l are both normalization constants; M is the number of longitudinal grids in the discrete grid of the layer velocity model of the work area, N is the number of transverse grids in the discrete grid of the layer velocity model of the work area; m is the serial number of the longitudinal grid in the discrete grid of the layer velocity model, 0 ≤ m ≤ M - 1 and m is an integer, n is the serial number of the transverse grid in the discrete grid of the layer velocity model, 0 ≤ n ≤ N - 1 and n is an integer; k and l are both coefficients, 0 ≤ k ≤ M - 1 and k is an integer, 0 ≤ l ≤ N - 1 and l is an integer;

[0023] Step 1.2.2, normalize the DCT matrix, and generate the Φ matrix by using the normalized DCT matrix. The Φ matrix includes M × N vectorized DCT matrices, as shown in formula (3):

[0024] Φ = [DCT1(:) DCT2(:)... DCT M×N (:)] (3)

[0025] In the formula, DCT M×N (:) is the M × Nth vectorized DCT matrix;

[0026] After discretizing the Φ matrix, the calculation formula of the DCT coefficient matrix is as shown in formula (4):

[0027]

[0028] In the formula, m is a model parameter, is the DCT coefficient matrix;

[0029] Perform SVD decomposition on the eigenvalues of the DCT matrix in the Φ matrix, and calculate the DCT coefficient matrix as shown in formula (5):

[0030]

[0031] In the formula, λ is a diagonal matrix composed of the eigenvalues of the Φ matrix, u is the first singular value decomposition matrix, v is the second singular value decomposition matrix, and T is the transpose matrix;

[0032] Step 1.2.3, add a random number obeying the normal distribution to the DCT coefficient matrix to generate a random perturbation, and generate the DCT coefficient matrix with random perturbation added, as shown in formula (6):

[0033]

[0034] In the formula, is the random perturbation;

[0035] Step 1.2.4, multiply the DCT coefficient matrix with random perturbations by the Φ matrix to obtain the velocity model with random perturbations, as shown in Equation (7):

[0036]

[0037] where m p is the velocity model with random perturbations.

[0038] Preferably, in Step 1.4, the frequency of the high-frequency part of the single-frequency point seismic wave is 5 - 10 Hz, and the frequency of the low-frequency part is 0.1 - 5 Hz.

[0039] Preferably, in Step 2, the convolutional neural network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a Flatten layer, a first fully connected layer, and a second fully connected layer;

[0040] A BN layer and a Max_Pooling layer are provided after each of the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer;

[0041] The first convolutional layer is provided with 16 3×3 convolutional kernels, the second convolutional layer is provided with 32 3×3 convolutional kernels, the third convolutional layer is provided with 64 3×3 convolutional kernels, the fourth convolutional layer is provided with 128 3×3 convolutional kernels, and the convolutional stride is 1;

[0042] The Max_Pooling layer is provided with 2×2 convolutional kernels, and the pooling stride is 2;

[0043] The Flatten layer is used to flatten the seismic record to form one-dimensional data.

[0044] Preferably, in Step 3, it specifically includes the following steps:

[0045] Step 3.1, according to the training set and the validation set, respectively select the training data and label data formed by the same single-frequency point seismic record in the training set and the validation set, and preset the accuracy value of the loss function of the neural network model;

[0046] Step 3.2, use the selected training data and label data as the input to the convolutional neural network model, use the selected training data as the sample data, and use the label data as the true value of the sample data;

[0047] Step 3.3, use the convolutional neural network model to calculate the predicted value of the sample data according to the input sample data;

[0048] Step 3.4: Compare the predicted values of the sample data with the true values of the input sample data, and calculate the loss function of the convolutional neural network model as shown in Equation (8):

[0049]

[0050] where MSE(·) is the loss function of the convolutional neural network model, y i is the true value of the i-th sample data, and y i ′ is the predicted value of the i-th sample data, and q is the number of input sample data;

[0051] Step 3.5: When the value of the loss function of the convolutional neural network model is less than the preset precision value or all the training data in the training set have been selected, go to Step 3.7; otherwise, go to Step 3.6;

[0052] Step 3.6: Adjust the weights of the convolutional neural network model, and re-determine the biases of each convolutional layer of the convolutional neural network model as shown in Equation (9):

[0053]

[0054] where a [s] is the output value of the s-th convolutional layer after being activated by the activation function, z [s] is the node value of the s-th convolutional layer, W [s] is the weight of the s-th convolutional layer, b [s] is the bias of the s-th convolutional layer, and a [s-1] is the output value of the (s - 1)-th convolutional layer after being activated by the activation function;

[0055] Re-select the training data and label data formed by the same single-frequency point seismic record in the training set and the validation set respectively, return to Step 3.2, and continue to train the convolutional neural network model;

[0056] Step 3.7: Obtain the trained convolutional neural network model.

[0057] Preferably, in Step 3.6, the activation function is the Relu function as shown in Equation (10):

[0058] Relu(x) = max(0, x) (10)

[0059] where Relu(·) is the Relu function, and max(0, x) is to take the maximum value between 0 and x.

[0060] The present invention has the following beneficial effects:

[0061] The present invention proposes a frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN. Based on the non-linear relationship between high-frequency seismic data and low-frequency seismic data in the same single-frequency seismic data, multiple sample data are generated. The sample data are input into a convolutional neural network model for training to perform low-frequency extension of the convolutional neural network model. The trained convolutional neural network model is used to predict the low-frequency seismic wavefield data according to the high-frequency seismic wavefield data of each single frequency point in the frequency-domain seismic wavefield data, and the high-frequency seismic wavefield data and the low-frequency seismic wavefield data are combined to construct wide-band seismic data from low frequency to high frequency.

[0062] The present invention uses wide-band seismic data for multi-scale full waveform inversion, effectively avoiding the problem that multi-scale full waveform inversion is prone to falling into local minima. The trained convolutional neural network model is used for low-frequency extension, increasing the low-frequency seismic wavefield data in the seismic data, achieving accurate acquisition of the inversion velocity model, and improving the accuracy of multi-scale full waveform velocity inversion. Description of the Drawings

[0063] Figure 1 It is a flow chart of the frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN of the present invention.

[0064] Figure 2 It is a schematic structural diagram of the convolutional neural network model in this embodiment.

[0065] Figure 3 It is the descending curve of the loss function in this embodiment.

[0066] Figure 4 It is the prediction result of the trained convolutional neural network model of the present invention for the 0.5 Hz seismic wavefield data. In the figure, (a) is the true value of the low-frequency seismic wavefield data corresponding to the 0.5 Hz seismic wavefield data, (b) is the predicted value of the low-frequency seismic wavefield data corresponding to the 0.5 Hz seismic wavefield data, and (c) is the difference between the predicted value and the true value of the low-frequency seismic wavefield data corresponding to the 0.5 Hz seismic wavefield data.

[0067] Figure 5 It is the prediction result of the trained convolutional neural network model of the present invention for the 1 Hz seismic wavefield data. In the figure, (a) is the true value of the low-frequency seismic wavefield data corresponding to the 1 Hz seismic wavefield data, (b) is the predicted value of the low-frequency seismic wavefield data corresponding to the 1 Hz seismic wavefield data, and (c) is the difference between the predicted value and the true value of the low-frequency seismic wavefield data corresponding to the 1 Hz seismic wavefield data.

[0068] Figure 6The prediction results of the convolutional neural network model trained by the present invention for 2 Hz seismic wavefield data. In the figure, (a) is the true value of the low-frequency seismic wavefield data corresponding to the 2 Hz seismic wavefield data, (b) is the predicted value of the low-frequency seismic wavefield data corresponding to the 2 Hz seismic wavefield data, and (c) is the difference between the predicted value and the true value of the low-frequency seismic wavefield data corresponding to the 2 Hz seismic wavefield data.

[0069] Figure 7 The prediction results of the convolutional neural network model trained by the present invention for 2.5 Hz seismic wavefield data. In the figure, (a) is the true value of the low-frequency seismic wavefield data corresponding to the 2.5 Hz seismic wavefield data, (b) is the predicted value of the low-frequency seismic wavefield data corresponding to the 2.5 Hz seismic wavefield data, and (c) is the difference between the predicted value and the true value of the low-frequency seismic wavefield data corresponding to the 2.5 Hz seismic wavefield data.

[0070] Figure 8 The prediction results of the convolutional neural network model trained by the present invention for 3 Hz seismic wavefield data. In the figure, (a) is the true value of the low-frequency seismic wavefield data corresponding to the 3 Hz seismic wavefield data, (b) is the predicted value of the low-frequency seismic wavefield data corresponding to the 3 Hz seismic wavefield data, and (c) is the difference between the predicted value and the true value of the low-frequency seismic wavefield data corresponding to the 3 Hz seismic wavefield data.

[0071] Figure 9 The prediction results of the convolutional neural network model trained by the present invention for 4 Hz seismic wavefield data. In the figure, (a) is the true value of the low-frequency seismic wavefield data corresponding to the 4 Hz seismic wavefield data, (b) is the predicted value of the low-frequency seismic wavefield data corresponding to the 4 Hz seismic wavefield data, and (c) is the difference between the predicted value and the true value of the low-frequency seismic wavefield data corresponding to the 4 Hz seismic wavefield data.

[0072] Figure 10 The prediction results of the convolutional neural network model trained by the present invention for 4.5 Hz seismic wavefield data. In the figure, (a) is the true value of the low-frequency seismic wavefield data corresponding to the 4.5 Hz seismic wavefield data, (b) is the predicted value of the low-frequency seismic wavefield data corresponding to the 4.5 Hz seismic wavefield data, and (c) is the difference between the predicted value and the true value of the low-frequency seismic wavefield data corresponding to the 4.5 Hz seismic wavefield data.

[0073] Figure 11 The frequency composition schematic diagram of two groups of data used in the multi-scale full waveform inversion in this embodiment.

[0074] Figure 12This is the multi-scale full waveform inversion result in this embodiment. In the figure, (a) is the actual Marmousi model; (b) is the multi-scale full waveform inversion result obtained by inverting from 5 Hz using the forward simulation data; (c) is the multi-scale full waveform inversion result obtained by inverting from 0.5 Hz using the forward simulation data; (d) is the multi-scale full waveform inversion result obtained by inverting from 0.5 Hz using the predicted low-frequency data; (e) is the difference between the multi-scale full waveform inversion results obtained by inverting from 0.5 Hz using the forward simulation data and by inverting from 0.5 Hz using the predicted low-frequency data; (f) is the difference between the multi-scale full waveform inversion results obtained by inverting from 5 Hz using the forward simulation data and by inverting from 0.5 Hz using the predicted low-frequency data. Detailed implementation manners

[0075] The following further describes the detailed implementation manners of the present invention with reference to the accompanying drawings:

[0076] A frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN proposed by the present invention, as Figure 1 shown, specifically includes the following steps:

[0077] Step 1, establish a layer velocity model for the work area according to the prior geological information of the area. Based on the discrete cosine transform, add random perturbations to the layer velocity model of the work area to generate multiple velocity models. Respectively use each velocity model to perform forward simulation of seismic records in the frequency domain to generate frequency-domain seismic forward records, and obtain seismic records at multiple single frequency points. Generate sample data by segmenting the seismic data at each single frequency point, specifically including the following steps:

[0078] Step 1.1, establish a layer velocity model for the work area according to the prior geological information of the area. In this embodiment, the layer velocity model of the work area adopts the Marmousi model.

[0079] Step 1.2, based on the discrete cosine transform, add random perturbations to the Marmousi model to obtain a velocity model with random perturbations. The specific process is as follows:

[0080] Step 1.2.1, perform a discrete cosine transform on the layer velocity model of the work area to obtain a DCT matrix, as shown in formula (1):

[0081]

[0082] Among them,

[0083]

[0084] In the formula, DCT kl is the layer velocity model of the work area after discrete cosine transform; c k 、c lare all normalization constants; M is the number of longitudinal grids in the discrete grid of the formation velocity model in the work area, N is the number of transverse grids in the discrete grid of the formation velocity model in the work area; m is the serial number of the longitudinal grid in the discrete grid of the formation velocity model in the work area, 0 ≤ m ≤ M - 1 and m is an integer, n is the serial number of the transverse grid in the discrete grid of the formation velocity model in the work area, 0 ≤ n ≤ N - 1 and n is an integer; k and l are both coefficients, 0 ≤ k ≤ M - 1 and k is an integer, 0 ≤ l ≤ N - 1 and l is an integer.

[0085] Step 1.2.2: To make the energies of the DCT matrices within the same range, perform normalization processing on the DCT matrices, and generate the Φ matrix using the normalized DCT matrices. The Φ matrix includes M×N vectorized DCT matrices, as shown in formula (3):

[0086]

[0087] In the formula, DCT M×N (:) is the M×Nth vectorized DCT matrix.

[0088] After discretizing the Φ matrix, the calculation formula of the DCT coefficient matrix is as shown in formula (4):

[0089]

[0090] In the formula, m is a model parameter, is the DCT coefficient matrix.

[0091] Perform SVD decomposition on the eigenvalues of the DCT matrices in the Φ matrix, and calculate the DCT coefficient matrix as shown in formula (5):

[0092]

[0093] In the formula, λ is a diagonal matrix composed of the eigenvalues of the Φ matrix, u is the first singular value decomposition matrix, v is the second singular value decomposition matrix, and T is the transpose matrix.

[0094] Step 1.2.3: Add a random number obeying the normal distribution to the DCT coefficient matrix to generate a random perturbation, and generate the DCT coefficient matrix with random perturbation as shown in formula (6):

[0095]

[0096] In the formula, is the random perturbation.

[0097] Step 1.2.4: The DCT coefficient matrix with random perturbation Multiply with the Φ matrix to obtain a velocity model with added random perturbations, as shown in Equation (7):

[0098]

[0099] where m p is the velocity model with added random perturbations.

[0100] Step 1.3: To increase the number of sample data and improve the convergence and universality of the deep learning model, in this embodiment, Step 1.2 is repeated multiple times to generate six Marmousi models with different amounts of random perturbations.

[0101] Step 1.4: Use each Marmousi model with added random perturbations to perform forward simulation of frequency-domain seismic records to generate frequency-domain seismic forward records. The observation systems used for forward simulation of each Marmousi model with added random perturbations are the same; there are 340 shots excited in the observation system, 680 geophones are used for reception per shot, and seismic records of 100 frequencies are forward simulated. The frequency distribution of the seismic records is from 0.1 Hz to 10 Hz, with an interval of 0.1 Hz.

[0102] Thus, multiple single-frequency seismic records are obtained. According to the frequency, the single-frequency seismic records are divided into a high-frequency part (5 - 10 Hz) and a low-frequency part (0.1 - 5 Hz). After equally spacing and dividing the high-frequency part and the low-frequency part of the single-frequency seismic records into 20 parts, use the high-frequency part of each single-frequency seismic record after segmentation as training data to form a training set, and use the low-frequency part of each single-frequency seismic record after segmentation as label data to form a validation set.

[0103] In this embodiment, each shot of seismic records is evenly divided into 20 parts along the geophone component. The size of each part of the data is 50×34×2 (where the third dimension 2 represents the real part and the imaginary part of the frequency-domain data). Thus, there are a total of 6×340×20 = 40800 sample data, and each sample data is a matrix of 50×34×2.

[0104] Step 2: Build a convolutional neural network model based on the TensorFlow and Keras deep learning platforms.

[0105] In this embodiment, the convolutional neural network (CNN) is as Figure 2 shown, including a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a Flatten layer, a first fully connected layer, and a second fully connected layer; a BN layer (i.e., batch normalization layer) and a Max_Pooling layer (i.e., max pooling layer) are provided after each of the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer.

[0106] There are 16 3×3 convolutional kernels in the first convolutional layer, 32 3×3 convolutional kernels in the second convolutional layer, 64 3×3 convolutional kernels in the third convolutional layer, and 128 3×3 convolutional kernels in the fourth convolutional layer. The convolutional stride is 1; a 2×2 convolutional kernel is set in the Max_Pooling layer, and the pooling stride is 2;

[0107] The Flatten layer is used to flatten the seismic record to form one-dimensional data, and the last fully connected layer is the target data for training.

[0108] In this embodiment, the sample data input to the convolutional neural network model is a 50×34×2 matrix. When the sample data is processed by the first convolutional layer, the size of the sample data becomes 25×17×16. After being processed by the second convolutional layer, the size of the sample data becomes 12×8×32. After being processed by the third convolutional layer, the size of the sample data becomes 6×4×64. After being processed by the fourth convolutional layer, the size of the sample data becomes 3×2×128. Then, through the Flatten layer, the sample data is flattened from three-dimensional data to one-dimensional data. After being processed by the first fully connected layer, it enters the second fully connected layer to output the target data.

[0109] Step 3: Input the training set and the validation set into the convolutional neural network model, and adjust the weights and biases of the convolutional neural network model by training the convolutional neural network model to obtain the trained convolutional neural network model, which specifically includes the following steps:

[0110] Step 3.1: According to the training set and the validation set, select the training data and label data formed by the same single-frequency point seismic record in the training set and the validation set respectively, and preset the accuracy value of the loss function of the neural network model;

[0111] Step 3.2: Input the selected training data and label data into the convolutional neural network model, use the selected training data as the sample data, and use the label data as the true value of the sample data.

[0112] Step 3.3: Use the convolutional neural network model to calculate the predicted value of the sample data according to the input sample data.

[0113] Step 3.4: Compare the predicted value of the sample data with the true value of the input sample data, and calculate the loss function of the convolutional neural network model as shown in formula (8):

[0114]

[0115] In the formula, MSE(·) is the loss function of the convolutional neural network model, y i is the true value of the i-th sample data, y i′ is the predicted value of the i-th sample data, and q is the number of input sample data.

[0116] Step 3.5, when the loss function value of the convolutional neural network model is less than the preset precision value or all the training data in the training set have been selected, go to Step 3.7; otherwise, go to Step 3.6.

[0117] Step 3.6, adjust the weights of the convolutional neural network model, and re-determine the biases of each convolutional layer of the convolutional neural network model, as shown in formula (9):

[0118] a [s] =σ(z [s] )=σ(W [s] a [s-1] +b [s] ) (9)

[0119] In the formula, a [s] is the output value of the s-th convolutional layer after being activated by the activation function. In this embodiment, the Relu function is used as the activation function; z [s] is the node value of the s-th convolutional layer, W [s] is the weight of the s-th convolutional layer, b [s] is the bias of the s-th convolutional layer, and a [s-1] is the output value of the (s - 1)-th convolutional layer after being activated by the activation function.

[0120] Re-select the training data and label data formed by the same single-frequency point seismic record in the training set and the validation set respectively, return to Step 3.2, and continue to train the convolutional neural network model.

[0121] Step 3.7, obtain the trained convolutional neural network model.

[0122] To prevent overfitting, in this embodiment, the early-stopping mechanism is used in combination with the sample data in the training set during the training process of the convolutional neural network model, that is, when the loss function value is less than 1×e -4 or all the sample data in the training set have been selected, stop the training of the convolutional neural network model to obtain the trained convolutional neural network model. Figure 3 is the descending curve of the loss function in this embodiment. It can be seen from Figure 3 that the training result of the convolutional neural network model is good and can fit the selected data set.

[0123] Step 4, use the trained convolutional neural network model for low-frequency prediction. Input the high-frequency seismic wavefield data in the frequency-domain seismic wavefield data into the trained convolutional neural network model, and use the trained convolutional neural network model to predict the low-frequency seismic wavefield data of each single-frequency point in the frequency-domain seismic wavefield data.

[0124] In this embodiment, the trained convolutional neural network model is used to predict the low-frequency seismic wavefield data at each single frequency point in the frequency-domain seismic wavefield data. The input data is the frequency-domain seismic wavefield data from 5 Hz to 10 Hz at an interval of 0.1 Hz, and the output data is the low-frequency seismic wavefield data at a single frequency point. The trained convolutional neural network model is used for prediction for each single frequency point respectively, and the prediction results are as Figures 4 to 10 shown. It can be seen that the error of the relatively low frequency predicted by the method of the present invention is smaller than that of the relatively high frequency. The reason is that the influence of the shallow complex structural features on the overall objective function is greater than that of the deep structural features on the overall objective function.

[0125] Step 5, in order to better verify the effect of low-frequency extension applied to multi-scale full waveform inversion, two groups of data are constructed in this embodiment. One group of data is the forward broadband data from low frequency to high frequency, and the other group of data is the broadband data composed of the combination of the low-frequency data predicted by the high-frequency data using the trained convolutional neural network model and the high-frequency data.

[0126] Frequency-domain multi-scale full waveform inversion is performed on the two groups of data respectively. During the inversion process, a multi-scale strategy of progressive frequency groups is adopted, and 13 progressive frequency groups are used. The frequency points of each frequency group are as Figure 11 shown.

[0127] Frequency-domain multi-scale full waveform inversion has certain advantages compared with time-domain multi-scale full waveform inversion. Since the forward modeling of seismic data can be performed at a finite number of single frequency points in the frequency domain, its calculation speed is much faster than that of time-domain forward modeling. By performing full waveform inversion from low frequency to high frequency on single frequency point data, the risk of "cycle skipping" can be effectively reduced, and the correctness of the inversion can be improved. At the same time, the interval between frequency points during the frequency-domain inversion process is not a constant, but changes with the increase of frequency, as shown in formula (11):

[0128]

[0129] where,

[0130]

[0131] In the formula, f n+1 is the frequency of the (n + 1)-th inversion data, f n is the frequency of the n-th inversion data, and α min is the ratio of the half-offset h max to the depth z.

[0132] The progressive frequency-group multi-scale strategy can increase the stability of inversion. Both sets of data contain 13 frequency groups, and the maximum frequency of each frequency group gradually increases, while the minimum frequency starts from 0.5 Hz. The open-source software TOY2DAC for full waveform inversion in the frequency domain is used for full waveform inversion. The velocity model obtained after the inversion of each frequency group is used as the initial model for the inversion of the next frequency group, and the iteration is carried out until the inversion of the last frequency group is completed, obtaining the multi-scale full waveform inversion result, as Figure 12 shown.

[0133] According to the multi-scale full waveform inversion result, due to the lack of low-frequency data, some errors occurred in the inversion result obtained using the original Marmousi model, and there was high-frequency interference in the deep part, indicating that the objective function fell into a local minimum during the inversion process and the correct velocity was not obtained. However, the broadband data composed of the low-frequency data predicted by the method of the present invention and the high-frequency data is used for inversion starting from the low frequency of 0.5 Hz, avoiding the objective function falling into a local minimum and obtaining the correct inversion result.

[0134] By comparing the inversion results of the two sets of data, it is fully verified that the convolutional neural network model trained by the method of the present invention can accurately predict low-frequency data, effectively avoid the inversion falling into a local minimum, obtain an accurate velocity model, and improve the accuracy of the multi-scale full waveform inversion result.

[0135] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.

Claims

1. A frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN, characterized in that, Specifically, it includes the following steps: Step 1: Establish a layer velocity model for the work area based on the prior geological information of the area. Add random perturbations to the layer velocity model of the work area based on the discrete cosine transform to generate multiple velocity models. Use each velocity model to perform forward simulation of seismic records in the frequency domain to generate forward seismic records in the frequency domain, obtain seismic records at multiple single-frequency points, and generate sample data by segmenting the seismic data at each single-frequency point; Step 2: Build a convolutional neural network model based on a deep learning platform; Step 3: Input the training set and the validation set into the convolutional neural network model. Adjust the weights and biases of the convolutional neural network model by training the convolutional neural network model to obtain the trained convolutional neural network model; Step 4: Use the trained convolutional neural network model for low-frequency prediction. Input the high-frequency seismic wavefield data in the frequency-domain seismic wavefield data into the trained convolutional neural network model, and use the trained convolutional neural network model to predict the low-frequency seismic wavefield data at each single-frequency point in the frequency-domain seismic wavefield data; Step 5: Use the low-frequency seismic wavefield data at each single-frequency point in the predicted frequency-domain seismic wavefield data combined with the high-frequency seismic wavefield data to construct broadband seismic data from low frequency to high frequency, and perform multi-scale full waveform inversion using the broadband seismic data to obtain the multi-scale full waveform inversion result based on low-frequency extension in the frequency domain.

2. The multi-scale full waveform inversion method for low-frequency extension in the frequency domain based on CNN according to claim 1, wherein In the said Step 1, it specifically includes the following steps: Step 1.1: Establish a layer velocity model for the work area based on the prior geological information of the area; Step 1.2: Add random perturbations to the layer velocity model of the work area based on the discrete cosine transform to obtain a velocity model with random perturbations added; Step 1.3: Repeat Step 1.2 multiple times to generate multiple velocity models with random perturbations added; Step 1.4: Use each velocity model to perform forward simulation of seismic records in the frequency domain to generate forward seismic records in the frequency domain, obtain seismic records at multiple single-frequency points. Divide the single-frequency point seismic records into high-frequency and low-frequency parts according to frequency. After equally spacing the high-frequency and low-frequency parts of the single-frequency point seismic records, use the high-frequency parts of the segmented single-frequency point seismic records as training data to form a training set, and use the low-frequency parts of the segmented single-frequency point seismic records as label data to form a validation set.

3. The frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN according to claim 2, characterized in that, In the said Step 1.2, it specifically includes the following steps: Step 1.2.1: Perform discrete cosine transform on the layer velocity model of the work area to obtain a DCT matrix, as shown in formula (1): Wherein, In the formula, DCT kl is the layer velocity model of the work area after discrete cosine transform; c k , c l are both normalization constants; M is the number of longitudinal grids in the discrete grid of the layer velocity model of the work area, N is the number of transverse grids in the discrete grid of the layer velocity model of the work area; m is the serial number of the longitudinal grid in the discrete grid of the layer velocity model of the work area, 0 ≤ m ≤ M - 1 and m is an integer, n is the serial number of the transverse grid in the discrete grid of the layer velocity model of the work area, 0 ≤ n ≤ N - 1 and n is an integer; k, l are both coefficients, 0 ≤ k ≤ M - 1 and k is an integer, 0 ≤ l ≤ N - 1 and l is an integer; Step 1.2.2: Perform normalization processing on the DCT matrix, and use the normalized DCT matrix to generate a Φ matrix. The Φ matrix includes M×N vectorized DCT matrices, as shown in formula (3): Φ = [DCT1(:) DCT2(:)... DCT M×N (:)] (3) Wherein, DCT M×N (:) is the M×N vectorized DCT matrix; The DCT coefficient matrix is obtained after discretizing the Φ matrix The calculation formula is shown in Formula (4) as follows: where m is a model parameter, is the DCT coefficient matrix; Perform SVD decomposition on the eigenvalues of the DCT matrix in the Φ matrix, and calculate to obtain the DCT coefficient matrix As shown in formula (5): In the formula, λ is a diagonal matrix composed of the eigenvalues of the Φ matrix, u is the first singular value decomposition matrix, v is the second singular value decomposition matrix, and T is the transpose matrix; Step 1.2.3, add a random number that follows a normal distribution to the DCT coefficient matrix to generate random perturbations and produce a DCT coefficient matrix with random perturbations added, as shown in formula (6): In the formula, is the random perturbation; Step 1.2.4, multiply the DCT coefficient matrix with added random perturbations by the Φ matrix to obtain a velocity model with added random perturbations, as shown in Equation (7): where m p is the velocity model with added random perturbations.

4. The multi-scale full waveform inversion method for low-frequency extension in the frequency domain based on CNN according to claim 2, characterized in that In the said Step 1.4, the frequency of the high-frequency part of the single-frequency point seismic wave is 5 - 10 Hz, and the frequency of the low-frequency part is 0.1 - 5 Hz.

5. The multi-scale full waveform inversion method for low-frequency extension in the frequency domain based on CNN according to claim 3, characterized in that In the said Step 2, the convolutional neural network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a Flatten layer, a first fully connected layer, and a second fully connected layer; A BN layer and a Max_Pooling layer are arranged after each of the first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer; There are 16 3×3 convolutional kernels in the first convolutional layer, 32 3×3 convolutional kernels in the second convolutional layer, 64 3×3 convolutional kernels in the third convolutional layer, and 128 3×3 convolutional kernels in the fourth convolutional layer, and the convolutional stride is 1; The Max_Pooling layer is provided with 2×2 convolutional kernels, and the pooling stride is 2; The Flatten layer is used to flatten the seismic record to form one-dimensional data.

6. The multi-scale full waveform inversion method for frequency domain low frequency extension based on CNN according to claim 5, characterized in that In step 3, it specifically includes the following steps: Step 3.1, according to the training set and the validation set, select the training data and label data formed by the same single-frequency point seismic record in the training set and the validation set respectively, and preset the accuracy value of the loss function of the neural network model; Step 3.2, take the selected training data and label data as the input to the convolutional neural network model, take the selected training data as the sample data, and take the label data as the true value of the sample data; Step 3.3, use the convolutional neural network model to calculate the predicted value of the sample data according to the input sample data; Step 3.4, compare the predicted value of the sample data with the true value of the input sample data, and calculate the loss function of the convolutional neural network model, as shown in formula (8): where MSE(·) is the loss function of the convolutional neural network model, y i is the true value of the i-th sample data, and y i ′ is the predicted value of the i-th sample data, and q is the number of input sample data; Step 3.5, when the loss function value of the convolutional neural network model is less than the preset accuracy value or all the training data in the training set have been selected, enter step 3.7; otherwise, enter step 3.6; Step 3.6, adjust the weights of the convolutional neural network model, and re-determine the biases of each convolutional layer of the convolutional neural network model, as shown in formula (9): a [s] = σ(z [s] ) = σ(W [s] a [s-1] + b [s] ) (9) where a [s] is the output value after activation by the activation function of the s-th convolutional layer, z [s] is the node value of the s-th convolutional layer, W [s] is the weight of the s-th convolutional layer, b [s] is the bias of the s-th convolutional layer, and a [s-1] is the output value after activation by the activation function of the (s - 1)-th convolutional layer; Re-select the training data and label data formed by the same single-frequency point seismic record in the training set and the validation set respectively, return to step 3.2, and continue to train the convolutional neural network model; Step 3.7, obtain the trained convolutional neural network model.

7. The frequency-domain low-frequency extension multi-scale full waveform inversion method based on CNN according to claim 6, characterized in that In step 3.6, the activation function is the Relu function, as shown in formula (10): Relu(x) = max(0, x) (10) In the formula, Relu(·) is the Relu function, and max(0, x) is to take the maximum value between 0 and x.

Citation Information

Patent Citations

  • Multi-scale seismic full-waveform inversion method based on local adaptive convexification method

    CN107422379A

  • Resolution controllable envelope generating operator-based multi-scale full-waveform inversion method

    CN107450102A