Convolutional neural network-based dual-phase medium pre-stack inversion system and method

CN122592468APending Publication Date: 2026-08-18LIAONING UNIVERSITY OF PETROLEUM AND CHEMICAL TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610753565.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

此种实现方式中,卷积神经网络的特征提取与参数回归过程独立于双相介质的物理波动机制,物理规律往往仅作为损失函数中的外部惩罚项参与网络优化,网络内部拓扑结构未包含双相介质固体骨架与孔隙流体耦合作用的物理约束

Benefits of technology

1.本发明构建物理机制嵌入的双分支卷积神经网络反演架构,将比奥双相介质波动方程构建为可微分正演算子层嵌入网络内部,使网络特征提取与参数回归过程受双相介质物理规律的内在约束,克服纯数据驱动模型违背波传播物理机制的缺陷,提升反演参数在地层流体替换场景下的物理一致性,缓解双相介质多参数强耦合导致的反演多解性问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122592468A_ABST
    Figure CN122592468A_ABST
Patent Text Reader

Abstract

This invention relates to the field of convolutional neural network technology, specifically to a pre-stack inversion system and method for two-phase media based on convolutional neural networks. A two-branch convolutional neural network inversion architecture with embedded physical mechanisms is constructed. The data feature extraction branch extracts the amplitude variation characteristics of pre-stack seismic gathers with offset. The initial predicted values ​​output by the parameter decoder are input into the physical forward modeling constraint branch. This branch constructs a differentiable forward modeling operator layer based on the Biot two-phase media wave equation, transforming the initial predicted values ​​into the theoretical seismic response under the coupling effect of a solid skeleton and pore fluid. The output features and the theoretical seismic response are fused at the feature level. The fused features are iteratively corrected by the parameter decoder to output the final two-phase media parameters. The differentiable forward modeling operator layer transforms the partial derivatives of the equation into convolution operations embedded in the graph computation structure. This subjectes the network to inherent physical constraints, improves the physical consistency of the inversion parameters, and alleviates the problem of multiple solutions caused by strong coupling of multiple parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of convolutional neural network technology, specifically to a biphase medium pre-stack inversion system and method based on convolutional neural networks. Background Technology

[0002] Existing pre-stack inversion techniques for two-phase media typically employ purely data-driven convolutional neural networks (CNNs) to construct the mapping relationship between seismic data and reservoir parameters. Pre-stack seismic gathers are used as network inputs, while solid bulk modulus, fluid bulk modulus, skeleton shear modulus, and porosity are used as network outputs. An end-to-end black-box mapping model is constructed through labeled sample training. In this approach, the feature extraction and parameter regression processes of the CNN are independent of the physical fluctuation mechanisms of the two-phase medium. Physical laws often only participate in network optimization as external penalty terms in the loss function, and the internal topology of the network does not include the physical constraints of the coupling between the solid skeleton and pore fluids in the two-phase medium.

[0003] The aforementioned pure data-driven black-box mapping implementation method results in the convolutional neural network lacking the inherent constraints of the physical laws of the two-phase medium during feature extraction and parameter regression. The mapping relationship learned by the network is prone to violate the physical mechanism of wave propagation. When faced with new data that is inconsistent with the distribution of the training data, the generalization ability is poor. Furthermore, the strong coupling between the solid parameters and fluid parameters of the two-phase medium leads to serious inversion ambiguity. Summary of the Invention

[0004] The purpose of this invention is to provide a two-phase medium pre-stack inversion system and method based on convolutional neural networks, which can effectively solve the problems in the background art mentioned above.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The dual-phase medium pre-stack inversion system based on convolutional neural networks includes a dual-branch convolutional neural network inversion architecture with embedded physical mechanisms. The architecture consists of a data feature extraction branch, a physical forward modeling constraint branch, and a parameter decoder. The data feature extraction branch receives pre-stack seismic gathers and extracts the amplitude variation features with offset using a multidimensional convolution kernel. The initial predicted values ​​of solid bulk modulus, fluid bulk modulus, skeleton shear modulus, and porosity output by the parameter decoder are input into the physical forward modeling constraint branch. The physical forward modeling constraint branch constructs a differentiable forward modeling operator layer based on the Biot two-phase medium wave equation, and transforms the initial predicted values ​​into the theoretical seismic response under the coupling action of solid skeleton and pore fluid. The output features of the data feature extraction branch are fused with the theoretical seismic response reconstructed by the physical forward modeling constraint branch at the feature level. The fused features are iteratively corrected by the parameter decoder to output the final two-phase medium parameters. The differentiable forward calculus layer transforms the partial derivative operations of the Biot equation into convolution operations of the kernel weights and a specific stride, and embeds them into the graph computation structure of the convolutional neural network.

[0006] Preferably, the data feature extraction branch includes a multi-scale residual feature extraction structure, which uses parallel dilated convolution kernels with different dilation rates to process the pre-stack seismic gathers and extract amplitude variation features at different offset scales. The output of the parallel dilated convolution kernel is concatenated along the channel dimension and then input to the channel attention mechanism layer, which assigns adaptive weights to features at each scale. The weighted multi-scale features are added element-wise to the basic features extracted by the pre-stack seismic gather through one-dimensional convolutional shorting to obtain the output features of the data feature extraction branch.

[0007] Preferably, the differentiable forward calculus sublayer converts the time and spatial partial derivatives in the Biot equation into a two-dimensional convolution operation with a specific stride and kernel parameters. The weights of the two-dimensional convolution operation are initialized and fixed according to the difference coefficients in the finite difference scheme of the wave equation. The wave propagation process under the coupling effect of solid skeleton and pore fluid is simulated by the multi-layer cascaded two-dimensional convolution operation, and a learnable damping coefficient parameter characterizing the attenuation characteristics of the two-phase medium is set between the cascaded two-dimensional convolution operations. The theoretical seismic response is obtained by performing a differentiable convolution operation on the output of the two-dimensional convolution operation and the source wavelet.

[0008] Preferably, the process of feature-level difference fusion between the output features of the data feature extraction branch and the theoretical seismic response reconstructed by the physical forward modeling constraint branch in the intermediate layer includes: mapping the theoretical seismic response through a feature extraction convolutional layer into physical implicit features with the same dimension as the output features; Calculate the absolute difference between the output feature and the physical implicit feature in the channel dimension to obtain the physical inconsistency feature map; The physical inconsistency feature map and the output feature are concatenated along the channel dimension, and cross-channel information interaction and dimensionality reduction are performed through a 1×1 convolutional layer to output the fused feature under physical constraints.

[0009] Preferably, the parameter decoder includes a skeleton parameter decoding channel and a fluid parameter decoding channel, wherein the skeleton parameter decoding channel and the fluid parameter decoding channel share a feature extraction coding layer but are each provided with an independent fully connected regression layer; The feature extraction encoding layer performs dimensionality reduction and flattening processing on the fused features; the skeleton parameter decoding channel outputs the solid bulk modulus and skeleton shear modulus; and the fluid parameter decoding channel outputs the fluid bulk modulus and porosity. A parameter physical boundary constraint layer is set at the end of the parameter decoder, and the final output two-phase medium parameters are truncated and mapped by using the Sigmoid activation function combined with the two-phase medium rock physical boundary value.

[0010] Preferably, the system employs a composite loss function driven by both physics and data during the training phase. The composite loss function is composed of a weighted average of a data reconstruction loss term, a parameter regression loss term, and a physical equation residual loss term. The data reconstruction loss term measures the waveform difference between the theoretical seismic response and the actual pre-stack seismic record; The parameter regression loss term measures the mean square error between the final two-phase medium parameters and the actual logging parameters. The physical equation residual loss term substitutes the final two-phase medium parameters into the Biot two-phase medium wave equation to calculate the equation residual. The weights of the composite loss function are dynamically adjusted during training based on the norm of each loss gradient.

[0011] Preferably, the parallel dilated convolution kernels with different dilation rates include a first dilated convolution kernel, a second dilated convolution kernel, and a third dilated convolution kernel with sequentially increasing dilation rates. The first dilated convolution kernel extracts near-offset reflection features, and the third dilated convolution kernel extracts far-offset fluid-sensitive features. The channel attention mechanism layer includes a global average pooling layer and a fully connected layer. The global average pooling layer compresses features at each scale into channel description vectors, and the fully connected layer outputs the normalized weight coefficients of each channel. The fully connected layer introduces a masking mechanism to randomly zero out the weight coefficients of some far-offset fluid-sensitive features during training.

[0012] Preferably, the learnable damping coefficient parameters in the differentiable forward operand layer are subject to physical prior projection constraints during the gradient backpropagation update process; The physical prior projection constraint maps the learnable damping coefficient parameter after each gradient update to the viscous coupling coefficient range set by Biot theory. If the updated learnable damping coefficient parameter exceeds the viscous coupling coefficient range, then the learnable damping coefficient parameter is truncated to the nearest boundary value; By embedding the physical prior projection constraint in the cascaded structure of the two-dimensional convolution operation, the search space for learnable parameters is limited.

[0013] Preferably, the physical inconsistency feature map is processed by an adaptive smoothing mask before being input into the 1×1 convolutional layer; The adaptive smoothing masking process generates a spatial weight matrix based on the local gradient variance of the physical inconsistency feature map, and reduces the weight value of the spatial weight matrix in regions where the local gradient variance is greater than a preset threshold. The physical inconsistency feature map after the adaptive smoothing masking is concatenated with the output feature along the channel dimension. The 1×1 convolutional layer performs cross-channel information interaction and dimensionality reduction with a learnable fusion scaling factor, and outputs the fused feature under physical constraints.

[0014] A two-phase medium pre-stack inversion method based on convolutional neural networks includes: acquiring pre-stack seismic gathers, extracting feature branches from the input data of the pre-stack seismic gathers, and extracting amplitude variation features with offset using a multidimensional convolutional kernel. The parameter decoder outputs the initial predicted values ​​of solid bulk modulus, fluid bulk modulus, skeleton shear modulus and porosity. The initial predicted values ​​are input into the physical forward modeling constraint branch. Based on the Biot two-phase medium wave equation, a differentiable forward modeling operator layer is constructed. The initial predicted values ​​are then transformed into the theoretical seismic response under the coupling action of solid skeleton and pore fluid. The output features of the data feature extraction branch are fused with the theoretical seismic response reconstructed by the physical forward modeling constraint branch at the feature level. The fused features are iteratively corrected by the parameter decoder to output the final two-phase medium parameters. The differentiable forward calculus layer transforms the partial derivative operations of the Biot equation into convolution operations of the kernel weights and a specific stride, and embeds them into the graph computation structure of the convolutional neural network.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention constructs a dual-branch convolutional neural network inversion architecture with embedded physical mechanisms. The Biot two-phase medium wave equation is constructed as a differentiable forward modeling operator layer and embedded inside the network. This makes the network feature extraction and parameter regression process subject to the inherent constraints of the physical laws of the two-phase medium, overcomes the defect of pure data-driven models that violate the physical mechanism of wave propagation, improves the physical consistency of inversion parameters in the context of formation fluid replacement, and alleviates the problem of multiple solutions in inversion caused by strong coupling of multiple parameters in two-phase medium.

[0016] 2. This invention generates a physical inconsistency feature map by calculating the absolute difference between the output features of the data feature extraction branch and the theoretical seismic response reconstructed by the physical forward modeling constraint branch in the channel dimension. This map is then concatenated with the output features to reduce dimensionality, achieving refined fusion of physical constraints at the feature level and improving the parameter identification accuracy of reservoirs with complex pore structures. Furthermore, it utilizes parallel void convolution kernels with different expansion rates to process pre-stack seismic gathers, extracting amplitude variation features at different offset scales and optimizing the network's feature extraction efficiency for geological targets at different scales. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the overall workflow of the pre-stack inversion system for biphase media based on convolutional neural networks according to the present invention. Figure 2 This is a flowchart of the multi-scale residual feature extraction branch of the data feature extraction method in this invention; Figure 3 This is a flowchart of the wavefield simulation and convolution transformation of the differentiable forward calculus operator layer of the present invention; Figure 4 This is a flowchart of the feature-level difference fusion and adaptive smoothing masking process of the present invention; Figure 5 This is a flowchart of the dual-channel decoding and physical boundary constraint process of the parameter decoder of the present invention; Figure 6 This is a flowchart of the composite loss function calculation and dynamic weight adjustment during the system training phase of this invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] Please refer to Figure 1 This embodiment provides a pre-stack inversion system and method for two-phase media based on convolutional neural networks. The pre-stack inversion system for two-phase media based on convolutional neural networks includes a two-branch convolutional neural network inversion architecture with embedded physical mechanisms. This architecture consists of a data feature extraction branch, a physical forward modeling constraint branch, and a parameter decoder. The data feature extraction branch receives pre-stack seismic gathers, which are seismic record matrices containing different offsets and sampling points at different times, with dimensions of [missing information]. ,in Indicates the offset track number. This indicates the number of time sampling points. The data feature extraction branch extracts the amplitude variation features with offset using a multidimensional convolution kernel. The size of the multidimensional convolution kernel is... ,in The kernel size is the offset dimension. The size of the convolution kernel is the time dimension.

[0020] The parameter decoder receives the output features from the data feature extraction branch and outputs the solid bulk modulus. Fluid bulk modulus Skeleton shear modulus and porosity The initial predicted values ​​are input into the physical forward modeling constraint branch. This branch constructs a differentiable forward modeling operator layer based on the Biot two-phase medium wave equation, transforming the initial predicted values ​​into the theoretical seismic response under the coupling effect of the solid skeleton and the porous fluid. The Biot two-phase medium wave equation describes the propagation law of elastic waves in a saturated fluid porous medium, considering the inertial and viscous coupling effects between the solid skeleton and the porous fluid.

[0021] Differentiable forward calculus layers transform the partial derivative operations of the Biot equation into convolution operations of kernel weights and a specific stride, embedding them into the graph computation structure of a convolutional neural network. Specifically, the temporal and spatial partial derivatives in the Biot equation are discretized using a finite difference scheme. The discretized difference coefficients serve as the weights of the convolution kernel, and the stride of the difference scheme serves as the stride of the convolution operation. This transformation allows the solution process of the wave equation to be efficiently parallelized using the graph computation framework of a convolutional neural network, while maintaining the mathematical consistency of the physical equations.

[0022] The output features of the data feature extraction branch are fused with the theoretical seismic response reconstructed by the physical forward modeling constraint branch at the feature level. The feature-level difference fusion process first maps the theoretical seismic response to a feature vector of the same dimension as the output features of the data feature extraction branch, then calculates the difference features between the two, and fuses these difference features with the original data features. The fused features are iteratively corrected by the parameter decoder to output the final two-phase medium parameters. The iterative correction process is achieved through multiple forward and backward propagations. In each iteration, the parameter decoder updates the predicted values ​​of the two-phase medium parameters based on the fused features, the physical forward modeling constraint branch recalculates the theoretical seismic response based on the updated parameter prediction values, and the data feature extraction branch and the physical forward modeling constraint branch perform feature-level difference fusion again until a preset iteration termination condition is met.

[0023] The preprocessing of pre-stack seismic gathers includes gather normalization, denoising, and static correction. Gather normalization maps the amplitude value of each seismic gather to the interval [-1, 1]. The normalization formula is as follows: in Indicates the first The offset of the first channel The original amplitude values ​​at each time sampling point and They represent the first The maximum and minimum amplitude values ​​of each offset channel. This represents the normalized amplitude value. Denoising is performed using wavelet transform thresholding, decomposing the seismic trace into wavelet coefficients of different scales. A soft thresholding function is applied to the high-frequency wavelet coefficients, followed by wavelet reconstruction to obtain the denoised seismic trace. Static correction eliminates the influence of variations in surface elevation and near-surface velocity on seismic wave propagation time, aligning the seismic trace's time axis to a unified reference plane.

[0024] The input to the data feature extraction branch is the preprocessed pre-stack seismic gather, with dimensions of [dimension value missing]. The last dimension represents the number of channels. The data feature extraction branch contains multiple convolutional layers, each consisting of convolution operations, batch normalization operations, and activation function operations. The output feature map dimension of the convolution operation is: in and These represent the height and width of the input feature map, respectively. and These represent the height and width of the output feature map, respectively. For fill size, The kernel size is [size]. The step size is [value]. Batch normalization normalizes the features of each channel using the following formula: in As input features, This is the batch average. For batch variance, To prevent small constants with a denominator of zero, These are the normalized features. The activation function used is the ReLU function, whose expression is: The ReLU function keeps the input unchanged when the input is positive and outputs zero when the input is negative, which can effectively alleviate the gradient vanishing problem and improve the training speed of the network.

[0025] The input to the parameter decoder is the output feature of the data feature extraction branch, and its dimension is... ,in and These represent the offset dimension and the time dimension after multiple convolution operations, respectively. Here is the number of channels. The parameter decoder first compresses the feature map into a one-dimensional vector using a global average pooling layer. The formula for the global average pooling operation is: in For the first Location in the feature map of each channel eigenvalues, and These represent the height and width of the feature map, respectively. For the first The result of global average pooling for each channel. The compressed one-dimensional vector has a dimension of... Then, feature transformation is performed through multiple fully connected layers, ultimately outputting the initial predicted values ​​for four biphase medium parameters. The output formula of the fully connected layer is: in For the input vector, This is the weight matrix. For bias vectors, This is the output vector.

[0026] The differentiable forward modeling operator layer in the physical forward modeling constraint branch is constructed based on the two-dimensional form of the Biot two-phase medium wave equation. The displacement form of the two-dimensional Biot two-phase medium wave equation is: in and Solid skeleton in direction and Displacement components in the direction, and These represent the pore fluid relative to the solid framework at... direction and Displacement components in the direction, The total density of the two-phase medium. For solid particle density, For fluid density, The effective mass density of the fluid. For curvature, The viscosity coefficient of the fluid. For penetration rate, The volumetric strain is that of the solid framework. For changes in fluid content, For compressibility modulus, For Biot coefficient, The bulk modulus of the skeleton. This refers to the fluid modulus.

[0027] Differentiable forward calculus layers transform the above partial differential equations into convolution operations. For time partial derivatives... The discretization is performed using a second-order central difference scheme, and the discretization formula is as follows: in The time step is [value]. The convolution kernel corresponding to this difference scheme is [kernel type]. The convolution stride is 1. For spatial partial derivatives... and The discretization is performed using a first-order central difference scheme, and the discretization formula is as follows: in The spatial stride is [value]. The convolution kernel corresponding to this difference scheme is [kernel name]. The convolution stride is 1. For the Laplace operator... The discretization is performed using a second-order central difference scheme, and the discretization formula is as follows: Where the assumption The convolution kernel corresponding to this difference scheme is: The convolution stride is 1.

[0028] Through the above transformation, the solution process of Biot's two-phase medium wave equation is converted into a series of cascaded two-dimensional convolution operations. The weights of each convolution operation are determined by the corresponding difference coefficients and remain fixed during network training, not participating in gradient updates. This design ensures that the differentiable forward calculus sublayer strictly follows the physical laws of Biot's two-phase medium wave equation, while enabling efficient computation using the graph computation framework of convolutional neural networks.

[0029] The generation process of theoretical seismic response includes two steps: wavefield simulation and source convolution. The wavefield simulation process simulates the propagation of seismic waves in a two-phase medium through cascaded two-dimensional convolution operations, calculating the wavefield value at each time step starting from the source location. The source wavelet uses the Ricker wavelet, whose expression is: in The dominant frequency of the wavelet is given. The reflection coefficient sequence obtained from wavefield simulation is convolved with the Ricker wavelet to obtain the theoretical seismic response. The formula for the convolution operation is: in For the reflection coefficient sequence, For the source wavelet, This represents the theoretical seismic response. The convolution operation is also transformed into a convolution operation, with the source wavelet as the convolution kernel and the reflection coefficient sequence as the input feature map.

[0030] The feature-level difference fusion process first inputs the theoretical seismic response output from the physical forward modeling constraint branch into a feature extraction convolutional layer. This convolutional layer has the same structure as the last convolutional layer in the data feature extraction branch, ensuring that the output physical implicit features have the same dimension as the output features of the data feature extraction branch. The output dimension of the feature extraction convolutional layer is... The output feature dimension is consistent with that of the data feature extraction branch.

[0031] The output features of the computational data feature extraction branch With physical implicit features The absolute difference along the channel dimension yields the physical inconsistency feature map. The calculation formula is: in and These are the spatial coordinates of the feature map. This is the channel index. The physical inconsistency feature map reflects the degree of difference between data characteristics and physical forward modeling results. The larger the difference, the higher the inconsistency between the current parameter prediction value and the physical law, and the more significant the correction is required.

[0032] Physical inconsistency feature map Output features of the data feature extraction branch By stitching along the channel dimension, a stitched feature map is obtained. Its dimensions are The concatenated feature map is input into a 1×1 convolutional layer for cross-channel information interaction and dimensionality reduction. The output channel number of the 1×1 convolutional layer is... Output fusion features under physical constraints The formula for operating a 1×1 convolutional layer is: in The weights are those of a 1×1 convolution kernel. This is the bias term. A 1×1 convolutional layer can achieve cross-channel information fusion and dimensionality transformation while maintaining spatial resolution.

[0033] Fusion features The data is input to the parameter decoder for iterative correction. The parameter decoder updates the predicted values ​​of the two-phase medium parameters based on the fusion features. The updated parameter predictions are then input back into the physical forward modeling constraint branch to calculate the new theoretical seismic response, and then feature-level difference fusion is performed again. This iterative process is repeated. Second-rate, This is the preset number of iterations. In each iteration, the weights of the parameter decoder are shared, ensuring that the number of network parameters does not increase significantly with the number of iterations.

[0034] The iteration termination condition can be set to reaching a preset number of iterations, or the difference between the predicted parameter values ​​of two adjacent iterations being less than a preset threshold. The formula for calculating the difference between the predicted parameter values ​​of two adjacent iterations is: in , , and The first Predicted values ​​of solid bulk modulus, fluid bulk modulus, skeletal shear modulus, and porosity from the next iteration. , , and The first The predicted value of the corresponding parameter in the next iteration. When the iteration process terminates, the final predicted values ​​of the two-phase medium parameters are output.

[0035] Table 1 shows the parameter settings for each layer of the dual-branch convolutional neural network inversion architecture in this embodiment.

[0036] Table 1. Parameter settings for each layer of the dual-branch convolutional neural network inversion architecture. In Table 1, the data feature extraction branch contains four convolutional layers, progressively reducing the spatial resolution of the feature map and increasing the number of channels to extract amplitude variation features with offset at different levels. The physical forward modeling constraint branch first transforms the parameter predictions into theoretical seismic responses through a differentiable forward modeling operator layer, and then maps the theoretical seismic responses into physical implicit features of the same dimension as the data features through a feature extraction convolutional layer. The feature fusion layer concatenates the data features and physical features and then performs dimensionality reduction and fusion through a 1×1 convolutional layer. The parameter decoder transforms the fused features into predicted values ​​of four biphasic medium parameters through a global average pooling layer and two fully connected layers.

[0037] This embodiment constructs a dual-branch convolutional neural network inversion architecture with embedded physical mechanisms. The Biot two-phase medium wave equation is transformed into a differentiable forward modeling operator layer embedded within the network, subjecting the network's feature extraction and parameter regression processes to the inherent constraints of the two-phase medium's physical laws. Through a feature-level difference fusion mechanism, the physical forward modeling results are finely integrated with data features, achieving an organic combination of data-driven and physics-driven approaches. The iterative correction process further optimizes the parameter prediction values, improving the physical consistency of the inversion results.

[0038] In a preferred embodiment, reference Figure 2 The data feature extraction branch includes a multi-scale residual feature extraction structure. This structure uses parallel dilated convolution kernels with different dilation rates to process pre-stack seismic gathers and extract amplitude variation features at different offset scales. The parallel dilated convolution kernels include a first, second, and third dilated convolution kernel with sequentially increasing dilation rates. The first dilated convolution kernel has a dilation rate of 1, corresponding to standard convolution, and extracts near-offset reflection features; the second dilated convolution kernel has a dilation rate of 2, extracting mid-offset reflection features; and the third dilated convolution kernel has a dilation rate of 4, extracting far-offset fluid-sensitive features.

[0039] Dilated convolution expands the receptive field by inserting zero values ​​between kernel elements without increasing the number of kernel parameters or computational cost. The effective kernel size for dilated convolution is calculated using the following formula: in This is the original kernel size. For expansion rate, This is the effective kernel size. For a kernel of size ... The convolution kernel, when its dilation rate is 1, has an effective kernel size of... When the dilation rate is 2, the effective kernel size is When the dilation rate is 4, the effective kernel size is Hollow convolutions with different expansion rates can capture geological features at different scales. Hollow convolutions with smaller expansion rates capture local fine features, while hollow convolutions with larger expansion rates capture global macroscopic features.

[0040] The outputs of the parallel dilated convolution kernels are concatenated along the channel dimension and then input into the channel attention mechanism layer. This layer assigns adaptive weights to features at each scale. The channel attention mechanism layer consists of a global average pooling layer and a fully connected layer. The global average pooling layer compresses features at each scale into channel description vectors, and the fully connected layer outputs the normalized weight coefficients for each channel. The calculation process of the channel attention mechanism is as follows: First, the input feature map Perform global average pooling to obtain the channel description vector. , of which The formula for calculating each element is: Then, the channel description vector The input is fed into the first fully connected layer, where feature transformation is performed to obtain the intermediate vector. ,in The compression ratio is calculated using the following formula: in This is the weight matrix of the first fully connected layer. For bias vectors, This is the Sigmoid activation function.

[0041] Next, the intermediate vector The input is fed into the second fully connected layer to restore the channel dimension, resulting in the channel weight vector. The calculation formula is: in This is the weight matrix for the second fully connected layer. This is the bias vector.

[0042] Finally, the channel weight vector Compared with the original input feature map Perform channel-by-channel multiplication to obtain the weighted feature map. The calculation formula is: Channel attention mechanisms can adaptively learn the importance of features in different channels, assigning higher weights to channels containing important information and lower weights to channels containing noisy or redundant information, thereby improving the efficiency and accuracy of feature extraction.

[0043] A masking mechanism is introduced into the fully connected layer to randomly zero out the weight coefficients of some far-offset fluid-sensitive features during training. The implementation process of the masking mechanism is as follows: In each training iteration, a mask is generated that is consistent with the channel weight vector. Binary mask vectors of the same dimension The elements corresponding to the far-offset fluid-sensitive feature channels are expressed with probability. The mask vector is set to zero, while the elements of other channels remain 1. Then, the mask vector... With channel weight vector Perform element-wise multiplication to obtain the masked channel weight vector. The calculation formula is: in This indicates element-wise multiplication. The masking mechanism can prevent the network from over-relying on long-offset fluid-sensitive features and improve the network's generalization ability.

[0044] The weighted multi-scale features are element-wise added to the basic features extracted from pre-stack seismic gathers via one-dimensional convolutional shortening, yielding the output features of the data feature extraction branch. The one-dimensional convolutional shortening path contains a one-dimensional convolutional layer with a kernel size of [missing value]. The step size is 1, and the padding is... This ensures that the temporal dimension of the output feature map is the same as that of the multi-scale features. The one-dimensional convolutional shortening path can extract the basic temporal features of pre-stack seismic gathers, and fuse them with the multi-scale features extracted by the multi-scale residual feature extraction structure to enrich the expressive power of the features.

[0045] refer to Figure 5 The parameter decoder includes a skeleton parameter decoding channel and a fluid parameter decoding channel. These two channels share a feature extraction encoding layer but each has an independent fully connected regression layer. The feature extraction encoding layer performs dimensionality reduction and flattening on the fused features, outputting a shared feature vector. The skeleton parameter decoding channel contains two fully connected layers, outputting the solid bulk modulus and skeleton shear modulus; the fluid parameter decoding channel contains two fully connected layers, outputting the fluid bulk modulus and porosity. This dual-channel decoding structure can utilize the correlation between solid and fluid parameters while maintaining their independence, mitigating the inversion ambiguity problem caused by strong coupling of multiple parameters.

[0046] A parametric physical boundary constraint layer is set at the end of the parametric decoder, and the final output two-phase medium parameters are truncated and mapped using the Sigmoid activation function combined with the rock physical boundary values ​​of the two-phase medium. The expression of the Sigmoid activation function is: The Sigmoid activation function maps input values ​​to the (0,1) interval. The formula for calculating the parametric physical boundary constraint layer is: in The output value of the fully connected layer. and These are the minimum and maximum physical boundary values ​​for the corresponding parameters. These are the parameter output values ​​after boundary constraints.

[0047] Table 2 shows the range of physical boundary values ​​for each parameter of the two-phase medium.

[0048] Table 2 Physical boundary value ranges of various parameters for two-phase media In Table 2, the solid bulk modulus ranges from 30 GPa to 80 GPa, corresponding to the bulk modulus of common rocks and minerals; the fluid bulk modulus ranges from 1.0 GPa to 3.0 GPa, corresponding to the bulk modulus of common reservoir fluids such as water and oil; the skeleton shear modulus ranges from 5 GPa to 30 GPa, corresponding to the shear modulus of rock skeletons with different degrees of compaction; and the porosity ranges from 0.01 to 0.40, corresponding to the porosity range from dense rocks to high-porosity reservoirs. The parametric physical boundary constraint layer ensures that the parameter values ​​output by the network are always within a reasonable physical range, avoiding prediction results that do not conform to physical laws.

[0049] refer to Figure 4 The process of feature-level difference fusion between the output features of the data feature extraction branch and the theoretical seismic response reconstructed by the physical forward modeling constraint branch in the intermediate layer includes: mapping the theoretical seismic response to physical implicit features with the same dimension as the output features through the feature extraction convolutional layer; calculating the absolute difference between the output features and the physical implicit features along the channel dimension to obtain a physical inconsistency feature map; the physical inconsistency feature map undergoes adaptive smoothing masking before being input to the 1×1 convolutional layer; the physical inconsistency feature map after adaptive smoothing masking is concatenated with the output features along the channel dimension, and cross-channel information interaction and dimensionality reduction are performed through the 1×1 convolutional layer to output the fused features under physical constraints.

[0050] The adaptive smoothing masking process generates a spatial weight matrix based on the local gradient variance of the physically inconsistent feature map. In regions where the local gradient variance exceeds a preset threshold, the weight values ​​of the spatial weight matrix are reduced. The calculation process for the local gradient variance is as follows: For the physically inconsistent feature map... Calculate each spatial location gradient magnitude The calculation formula is: in and The physical inconsistency feature maps are respectively in direction and Gradient of direction.

[0051] Then, calculate each spatial location. Local gradient variance The local window size is The calculation formula is: in The average gradient magnitude within the local window is calculated using the following formula: Generate a spatial weight matrix based on the local gradient variance. The calculation formula is: in This is a preset gradient variance threshold. When the local gradient variance is greater than the threshold, the weight values ​​of the spatial weight matrix are inversely proportional to the local gradient variance, thereby reducing the weight of physical inconsistencies in high-gradient regions and minimizing the impact of noise and outliers on the fusion process.

[0052] Spatial weight matrix Physical inconsistency feature map Element-wise multiplication yields the physical inconsistency feature map after adaptive smoothing masking. The calculation formula is: Physical inconsistency feature map after adaptive smoothing masking Output features of the data feature extraction branch By stitching along the channel dimension, a stitched feature map is obtained. Its dimensions are The concatenated feature map is input into a 1×1 convolutional layer for cross-channel information interaction and dimensionality reduction. The output channel number of the 1×1 convolutional layer is... Output fusion features under physical constraints The formula for operating a 1×1 convolutional layer is: in The weights are those of a 1×1 convolution kernel. This is the bias term. A 1×1 convolutional layer can learn the fusion ratio between data features and physically inconsistent features, achieving adaptive feature fusion.

[0053] This embodiment refines the multi-scale residual feature extraction structure of the data feature extraction branch, employing parallel dilated convolution kernels with different dilation rates to extract multi-scale features and adaptively allocating feature weights through a channel attention mechanism. The parameter decoder uses a dual-channel decoding structure to process skeleton parameters and fluid parameters separately, alleviating the problem of strong coupling among multiple parameters. The feature-level difference fusion process introduces adaptive smoothing masking to reduce the impact of noise and outliers, thereby improving the quality of the fused features.

[0054] In a preferred embodiment, reference Figure 3The differentiable forward calculus layer transforms the temporal and spatial partial derivatives in the Biot equation into two-dimensional convolution operations with specific strides and kernel parameters. The weights of the two-dimensional convolution operations are initialized and fixed based on the difference coefficients in the finite difference scheme of the wave equation. The wave propagation process under the coupling between the solid skeleton and the porous fluid is simulated through multi-layer cascaded two-dimensional convolution operations, with learnable damping coefficient parameters characterizing the attenuation properties of the two-phase medium set between the cascaded two-dimensional convolution operations.

[0055] The structure of the differentiable forward calculus layer includes wavefield update layers at multiple time steps. Each wavefield update layer contains multiple two-dimensional convolution operations, corresponding to different partial derivative terms in the Biot equation. The weights of each two-dimensional convolution operation are determined by the corresponding finite difference coefficients and remain fixed during network training. Learnable damping coefficient parameters... It is set after the wave field update layer at each time step to simulate the wave attenuation characteristics in a two-phase medium.

[0056] The formula for calculating the wave field update process is as follows: in and The first Displacement wave field of solid skeleton and relative displacement wave field of fluid at each time step and The first The corresponding wave field at each time step and The first The corresponding wave field at each time step For time step, , , , For the convolution kernel corresponding to the partial derivative terms, This represents the convolution operation. For the first Learnable damping coefficient parameters for each time step.

[0057] The learnable damping coefficient parameter is subject to a physical prior projection constraint during gradient backpropagation. This constraint maps the learnable damping coefficient parameter after each gradient update to a viscous coupling coefficient range defined by Biot theory. The range of the viscous coupling coefficient range is... ,in and These are the minimum and maximum values ​​of the viscous coupling coefficient in Biot theory, respectively.

[0058] The formula for calculating the gradient update process is: in The damping coefficient parameters before the update. For learning rate, The gradient of the loss function with respect to the damping coefficient parameter. This refers to the updated damping coefficient parameters.

[0059] The formula for calculating the physical prior projection constraints is: in This represents the damping coefficient parameter after projection constraints. If the updated damping coefficient parameter exceeds the viscous coupling coefficient range, it is truncated to the nearest boundary value. By embedding physical prior projection constraints in the cascaded structure of the two-dimensional convolution operation, the search space of learnable parameters is limited, ensuring that the learnable parameters always conform to the physical laws of Biot theory.

[0060] The theoretical seismic response is obtained by performing a differentiable convolution operation between the output of a two-dimensional convolution operation and the source wavelet. The formula for the convolution operation is: in The reflection coefficient sequence is obtained from wave field simulation. For the source wavelet, This represents the theoretical earthquake response. The source wavelet can be a fixed Ricker wavelet or a learnable wavelet. When the source wavelet is a learnable wavelet, its parameters are updated during network training via gradient backpropagation, which better matches the wavelet characteristics of actual earthquake data.

[0061] refer to Figure 6 During the training phase, the system employs a composite loss function driven by both physics and data. This composite loss function is a weighted average of data reconstruction loss terms, parametric regression loss terms, and physical equation residual loss terms. The expression for the composite loss function is: in For data reconstruction loss terms, For parametric regression loss term, This is the residual loss term in the physical equation. , , These are the corresponding weighting coefficients.

[0062] The data reconstruction loss term measures the waveform difference between the theoretical seismic response and the actual pre-stack seismic record. The mean square error loss function is used, and the calculation formula is as follows: in The theoretical seismic response output by the physical forward modeling constraint branch is in the 1st... The offset of the first channel The values ​​at each time sampling point This represents the value of the actual pre-stack earthquake record at the corresponding location.

[0063] The parameter regression loss term measures the mean square error between the final two-phase medium parameters and the actual logging parameters. The calculation formula is as follows: in , , , These are the predicted values ​​of solid bulk modulus, fluid bulk modulus, skeleton shear modulus, and porosity output by the parameter decoder, respectively. , , , These are the corresponding actual logging parameter values.

[0064] The residual loss term of the physical equation is calculated by substituting the final two-phase medium parameters into Biot's two-phase medium wave equation. The calculation formula is as follows: in , , , The four equations of Biot's two-phase medium wave equation are located at... The residuals are calculated as follows: Substitute the two-phase medium parameters output by the parameter decoder into the Biot equation, calculate the difference between the left and right sides of the equation, and the absolute value of the difference is the equation residual. The smaller the equation residual, the more the predicted parameter values ​​conform to the physical laws of the Biot two-phase medium wave equation.

[0065] The weights of the composite loss function are dynamically adjusted during training based on the norm of each loss gradient. The formula for dynamic adjustment is: in For the first The L2 norm of the gradient of the loss function. These are the corresponding weight coefficients. This dynamic weight adjustment mechanism can balance the contributions of various loss functions to network training, preventing any one loss function from dominating the training process and ensuring that the network learns both data features and physical laws simultaneously.

[0066] Table 3 shows the dynamic changes of the weights of the composite loss function at different training stages.

[0067] Table 3. Dynamic changes of the weights of the composite loss function at different training stages. Table 3 shows that in the initial training stage, the data reconstruction loss has a high weight, and the network mainly learns the basic mapping relationship between seismic data and reservoir parameters. In the middle training stage, the weight of the parameter regression loss gradually increases, and the network begins to focus on the accuracy of parameter predictions. In the later training stage, the weight of the physical equation residual loss increases significantly, and the network focuses on optimizing the physical consistency of parameter prediction values. This dynamic weight adjustment strategy enables the network to gradually transition from data-driven to physics-driven during training, improving the reliability of the inversion results.

[0068] The network training process uses the Adam optimizer with an initial learning rate of 0.001. The learning rate decay strategy is to reduce it to 0.9 times its original value every 1000 iterations. The batch size is set to 32, and the ratio of training set to validation set is 8:2. During training, the validation loss is calculated on the validation set every 100 iterations. Training is terminated early when the validation loss no longer decreases for 500 consecutive iterations.

[0069] This embodiment refines the structure of the differentiable forward calculus operator layer, introduces learnable damping coefficient parameters, and applies physical prior projection constraints, thereby improving the accuracy of wavefield simulation. The composite loss function employs a dynamic weight adjustment mechanism, balancing the losses from data reconstruction, parameter regression, and physical constraints, ensuring the network simultaneously learns data features and physical laws. The training process utilizes an adaptive optimizer and a learning rate decay strategy, improving the network's training efficiency and stability.

Claims

1. A two-phase medium pre-stack inversion system based on convolutional neural networks, characterized in that, The architecture includes a dual-branch convolutional neural network inversion architecture with embedded physical mechanisms, which consists of a data feature extraction branch, a physical forward modeling constraint branch, and a parameter decoder. The data feature extraction branch receives pre-stack seismic gathers and extracts the amplitude variation features with offset using a multidimensional convolution kernel. The initial predicted values ​​of solid bulk modulus, fluid bulk modulus, skeleton shear modulus, and porosity output by the parameter decoder are input into the physical forward modeling constraint branch. The physical forward modeling constraint branch constructs a differentiable forward modeling operator layer based on the Biot two-phase medium wave equation, and transforms the initial predicted values ​​into the theoretical seismic response under the coupling action of solid skeleton and pore fluid. The output features of the data feature extraction branch are fused with the theoretical seismic response reconstructed by the physical forward modeling constraint branch at the feature level. The fused features are iteratively corrected by the parameter decoder to output the final two-phase medium parameters. The differentiable forward calculus layer transforms the partial derivative operations of the Biot equation into convolution operations of the kernel weights and a specific stride, and embeds them into the graph computation structure of the convolutional neural network.

2. The biphase medium pre-stack inversion system based on convolutional neural networks according to claim 1, characterized in that, The data feature extraction branch includes a multi-scale residual feature extraction structure. The multi-scale residual feature extraction structure uses parallel dilated convolution kernels with different expansion rates to process the pre-stack seismic gathers and extract amplitude variation features at different offset scales. The output of the parallel dilated convolution kernel is concatenated along the channel dimension and then input to the channel attention mechanism layer, which assigns adaptive weights to features at each scale. The weighted multi-scale features are added element-wise to the basic features extracted by the pre-stack seismic gather through one-dimensional convolutional shorting to obtain the output features of the data feature extraction branch.

3. The pre-stack inversion system for biphase media based on convolutional neural networks according to claim 1, characterized in that, The differentiable forward calculus sublayer converts the time and spatial partial derivatives in the Biot equation into two-dimensional convolution operations with specific strides and kernel parameters. The weights of the two-dimensional convolution operations are initialized and fixed according to the difference coefficients in the finite difference scheme of the wave equation. The wave propagation process under the coupling effect of solid skeleton and pore fluid is simulated by the multi-layer cascaded two-dimensional convolution operation, and a learnable damping coefficient parameter characterizing the attenuation characteristics of the two-phase medium is set between the cascaded two-dimensional convolution operations. The theoretical seismic response is obtained by performing a differentiable convolution operation on the output of the two-dimensional convolution operation and the source wavelet.

4. The biphase medium pre-stack inversion system based on convolutional neural networks according to claim 1, characterized in that, The process of feature-level difference fusion between the output features of the data feature extraction branch and the theoretical seismic response reconstructed by the physical forward modeling constraint branch in the intermediate layer includes: mapping the theoretical seismic response through the feature extraction convolutional layer into physical implicit features with the same dimension as the output features; Calculate the absolute difference between the output feature and the physical implicit feature in the channel dimension to obtain the physical inconsistency feature map; The physical inconsistency feature map and the output feature are concatenated along the channel dimension, and cross-channel information interaction and dimensionality reduction are performed through a 1×1 convolutional layer to output the fused feature under physical constraints.

5. The pre-stack inversion system for biphase media based on convolutional neural networks according to claim 1, characterized in that, The parameter decoder includes a skeleton parameter decoding channel and a fluid parameter decoding channel. The skeleton parameter decoding channel and the fluid parameter decoding channel share a feature extraction coding layer but each is set with an independent fully connected regression layer. The feature extraction encoding layer performs dimensionality reduction and flattening processing on the fused features; the skeleton parameter decoding channel outputs the solid bulk modulus and skeleton shear modulus; and the fluid parameter decoding channel outputs the fluid bulk modulus and porosity. A parameter physical boundary constraint layer is set at the end of the parameter decoder, and the final output two-phase medium parameters are truncated and mapped by using the Sigmoid activation function combined with the two-phase medium rock physical boundary value.

6. The pre-stack inversion system for biphase media based on convolutional neural networks according to claim 1, characterized in that, The system employs a composite loss function driven by both physics and data during the training phase. This composite loss function is composed of a weighted average of a data reconstruction loss term, a parameter regression loss term, and a physical equation residual loss term. The data reconstruction loss term measures the waveform difference between the theoretical seismic response and the actual pre-stack seismic record; The parameter regression loss term measures the mean square error between the final two-phase medium parameters and the actual logging parameters. The physical equation residual loss term substitutes the final two-phase medium parameters into the Biot two-phase medium wave equation to calculate the equation residual. The weights of the composite loss function are dynamically adjusted during training based on the norm of each loss gradient.

7. The pre-stack inversion system for biphase media based on convolutional neural networks according to claim 2, characterized in that, The parallel dilated convolution kernels with different expansion rates include a first dilated convolution kernel, a second dilated convolution kernel, and a third dilated convolution kernel with sequentially increasing expansion rates. The first dilated convolution kernel extracts near-offset reflection features, and the third dilated convolution kernel extracts far-offset fluid-sensitive features. The channel attention mechanism layer includes a global average pooling layer and a fully connected layer. The global average pooling layer compresses features at each scale into channel description vectors, and the fully connected layer outputs the normalized weight coefficients of each channel. The fully connected layer introduces a masking mechanism to randomly zero out the weight coefficients of some far-offset fluid-sensitive features during training.

8. The pre-stack inversion system for biphase media based on convolutional neural networks according to claim 3, characterized in that, The learnable damping coefficient parameters in the differentiable forward operand layer are subject to physical prior projection constraints during the gradient backpropagation update process. The physical prior projection constraint maps the learnable damping coefficient parameter after each gradient update to the viscous coupling coefficient range set by Biot theory. If the updated learnable damping coefficient parameter exceeds the viscous coupling coefficient range, then the learnable damping coefficient parameter is truncated to the nearest boundary value; By embedding the physical prior projection constraint in the cascaded structure of the two-dimensional convolution operation, the search space for learnable parameters is limited.

9. The pre-stack inversion system for biphase media based on convolutional neural networks according to claim 4, characterized in that, The physical inconsistency feature map is processed by an adaptive smoothing mask before being input into the 1×1 convolutional layer; The adaptive smoothing masking process generates a spatial weight matrix based on the local gradient variance of the physical inconsistency feature map, and reduces the weight value of the spatial weight matrix in regions where the local gradient variance is greater than a preset threshold. The physical inconsistency feature map after the adaptive smoothing masking is concatenated with the output feature along the channel dimension. The 1×1 convolutional layer performs cross-channel information interaction and dimensionality reduction with a learnable fusion scaling factor, and outputs the fused feature under physical constraints.

10. A two-phase medium pre-stack inversion method based on convolutional neural networks, characterized in that, include: Obtain pre-stack seismic gathers, extract feature branches from the input data of the pre-stack seismic gathers, and extract amplitude variation features with offset using multidimensional convolution kernels; The parameter decoder outputs the initial predicted values ​​of solid bulk modulus, fluid bulk modulus, skeleton shear modulus and porosity. The initial predicted values ​​are input into the physical forward modeling constraint branch. Based on the Biot two-phase medium wave equation, a differentiable forward modeling operator layer is constructed. The initial predicted values ​​are then transformed into the theoretical seismic response under the coupling action of solid skeleton and pore fluid. The output features of the data feature extraction branch are fused with the theoretical seismic response reconstructed by the physical forward modeling constraint branch at the feature level. The fused features are iteratively corrected by the parameter decoder to output the final two-phase medium parameters. The differentiable forward calculus layer transforms the partial derivative operations of the Biot equation into convolution operations of the kernel weights and a specific stride, and embeds them into the graph computation structure of the convolutional neural network.