Porous medium three-dimensional image reconstruction method based on multi-scale feature fusion
By designing an extended feature pyramid network (EFPN-StyleGAN), combining encoder, generator and discriminator, accurate mapping and reconstruction from two-dimensional images to three-dimensional images is achieved, solving the problem of high cost and complexity of reconstruction of three-dimensional structures of porous media in the prior art, and achieving low-cost and efficient three-dimensional image reconstruction.
Patent Information
- Application Number
- CN202410103917.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to accurately reconstruct the three-dimensional structure of a porous medium from a two-dimensional image, especially when a large number of three-dimensional images of porous medium are required for macroscopic characteristics analysis in industrial applications. The existing methods are costly and complex in operation.
Using a porous media three-dimensional image reconstruction method based on multi-scale feature fusion, the extended feature pyramid network (EFPN-StyleGAN) is designed, combined with encoder, generator and discriminator, and the multi-scale feature pyramid and generative adversarial network are used to achieve accurate mapping and reconstruction of two-dimensional images to three-dimensional images.
It realizes the accurate reconstruction of the three-dimensional structure of the porous medium at low cost and high efficiency, maintains the texture morphology and statistical characteristics of the image, and improves the accuracy and stability of the reconstruction.
Smart Images

Figure BDA0004681058240000031 
Figure BDA0004681058240000033 
Figure BDA0004681058240000071
Abstract
Description
Technical Field
[0001] The present invention relates to a three-dimensional reconstruction method based on a two-dimensional image, and particularly to a three-dimensional image reconstruction method of a porous medium based on multi-scale feature fusion, belonging to the technical field of three-dimensional image processing. Background Art
[0002] In nature and people's lives, porous media widely exist and play an important role especially in many industrial application fields. Typical porous media include composite materials, soil, battery materials, rocks, etc. In order to understand the physical properties and transport properties of porous media, such as porosity, pore structure, conductivity, wettability, permeability, etc., it is necessary to accurately obtain the internal structure of the porous medium. However, a two-dimensional image cannot completely represent the internal structure of the porous medium, so it is necessary to reconstruct the corresponding three-dimensional image from a two-dimensional image.
[0003] The methods for three-dimensional imaging of porous media mainly include: (1) three-dimensional imaging technology, (2) three-dimensional reconstruction methods based on images. X-ray computed tomography (CT) and focused ion beam scanning electron microscope (FIB-SEM) and other technologies can directly obtain the three-dimensional image of the porous medium. However, for these methods to obtain their three-dimensional CT images, not only the instrument operation is complex and the price is expensive, especially in the industrial application field, a large number of three-dimensional images of the porous medium are required for the analysis and research of its macroscopic properties. However, the imaging methods for obtaining two-dimensional images are simple and the cost is low. Therefore, the reconstruction of its three-dimensional structure from two-dimensional images has attracted the attention of researchers.
[0004] In recent years, deep learning generative adversarial networks have been widely used in the reconstruction of porous media, bringing new research methods to image generation. These methods all extract the features of the two-dimensional image and then map the features of the two-dimensional image to the reconstructed three-dimensional image. Since the porous medium image has pores of different sizes in a two-dimensional image, accurately extracting the multi-scale features of the two-dimensional image of the porous medium can achieve the purpose of accurately reproducing the texture morphology and statistical features of the two-dimensional image in the reconstructed three-dimensional image. Therefore, based on the obtained two-dimensional image, designing a two-dimensional image feature extraction network, a generative network that maps the features of the two-dimensional image to the features of the three-dimensional image, and a loss function are the key problems to be solved for accurately and stably reconstructing the three-dimensional image. Summary of the Invention
[0005] The purpose of the present invention is to provide an accurate and stable three-dimensional image reconstruction method of a porous medium based on multi-scale feature fusion to solve the above problems.
[0006] The present invention achieves the above object through the following technical solutions:
[0007] A three-dimensional image reconstruction method for porous media based on multi-scale feature fusion, comprising the following steps:
[0008] (1) Data acquisition, data preprocessing, and construction of a two-dimensional image dataset;
[0009] (2) Design a network structure (EFPN-StyleGAN) for reconstructing three-dimensional images from two-dimensional images based on an extended feature pyramid (Feature Pyramid Network, FPN);
[0010] (3) Design the loss function of the EFPN-StyleGAN network The purpose is to enable the three-dimensional image generated by the generator to learn the morphology and multi-scale distribution characteristics of the real image, and improve the accuracy of the discriminator in distinguishing the generated image;
[0011] (4) Design the mode density loss function of the EFPN-StyleGAN network The purpose is to constrain the mode distribution of the generated three-dimensional image;
[0012] (5) Based on the above-constructed dataset, network structure, and loss function, train the neural network until the generated three-dimensional image meets the target requirements, thereby obtaining the EFPN-StyleGAN model for image three-dimensional reconstruction;
[0013] (6) Based on the described EFPN-StyleGAN model, complete the three-dimensional reconstruction of the two-dimensional image.
[0014] In the above solution, in the network structure EFPN-StyleGAN described in step (2), the backbone network design is based on the EFPN encoder E. The EFPN-StyleGAN network structure includes an encoder E based on an extended feature pyramid (Extended FPN, EFPN), a generator G, and three discriminators D xy 、D yz 、D zx, based on the fact that the main function of the EFPN encoder E is to extract features of different resolutions and scales of a two-dimensional image, the backbone network of the encoder selects ResNet50. Based on the design requirements of reconstructing a three-dimensional image, the backbone network ResNet50 is improved. In the first stage, first, a 4×4 convolutional kernel is used on the backbone network ResNet50 to extract features from the original image to obtain the first-layer feature map C0, with a size of (66, 66). The pooling layer after the first convolutional layer in the original ResNet50 network is discarded, and the obtained high-resolution feature of the first layer is embedded into the bottom layer of the FPN as an extended feature. In this way, the problem of small pores occupying few pixels is solved, and the high-resolution features are well utilized. This feature contains the semantic and position information of small pores and the texture and edge information of large pores; in the second stage, an adaptive average pooling is used to obtain a feature map with a size of (34, 34), and then three residual modules are connected. After passing through these three residual modules, the size of the feature map remains unchanged, but the network more accurately learns the features of this layer size, denoted as C1, and so on, to obtain C i (i = 2, 3, 4); C i (i = 0, 1, 2, 3, 4) are the multi-scale features obtained from the bottom-up structure of the backbone network, that is, the input of the FPN; and a new high-resolution extended feature map C0 introduced from the output of the convolutional layer Conv0 of the backbone network, which contains more spatial and content information and plays a key role in capturing and positioning small pores in subsequent reconstruction.
[0015] In the above solution, in the network structure EFPN-StyleGAN described in step (2), based on the feature fusion network design of the EFPN encoder E, considering the different sensitivities of feature maps of different scales and different layers in the EFPN encoder E, a feature fusion network is designed. These five features C i(i = 0, 1, 2, 3, 4) is sent as input to the EFPN encoder. The EFPN encoder network consists of two branches, a vertical branch and a horizontal branch. The vertical path upsamples the high-semantic information porous medium high-level feature map with low spatial resolution to obtain a high-resolution feature map. The horizontal path fuses the EFPN feature map with the corresponding low-level features to obtain the required multi-scale fused feature map. The EFPN encoder uses upsampling and summation operations to fuse different features in the image. Taking the fusion of feature maps C1 and P2 as an example, first, the number of channels of feature map C1 is reduced through a 1×1 convolution to obtain C1'. To make the size of feature map P2 consistent with C1', P2 is upsampled. Finally, C1' and the upsampled P2 are added to obtain the fused feature map. To eliminate the aliasing effect generated by upsampling, the fused feature map P2 is processed using a 3×3 convolution to obtain feature map F1'. Drawing on the squeezing and excitation idea and attention mechanism in SENet, global average pooling is used to convert feature map F1' from c×h×w to c×1×1, thus expanding the receptive field of the feature map from local to global to obtain the final feature map F1, thereby capturing richer global context information and breaking the limitation that convolution only extracts local information. The final feature map F obtained from the entire process i (i = 0, 1, 2, 3, 4) is represented by the mathematical relationship:
[0016]
[0017] where C i represents the feature map of the i-th lateral connection, P i represents the feature map upsampled from P i+1 to the size of C i , S up is the upsampling operation. The upsampling method is bilinear interpolation. The size of the upsampled feature map is the size of the image it is to be fused with. f13×3 is a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1. Ave represents global average pooling, aiming to convert the feature map from c×h×w to c×1×1, represents the fusion operation. Finally, the multi-scale feature F i (i = 0, 1, 2, 3, 4) obtained by the EFPN. Correspondingly, since C0 is the feature map of a high-resolution image, the fused F0 contains not only the spatial and content feature information of the high-resolution image but also strong semantic feature information.
[0018] In the above solution, in the network structure EFPN-StyleGAN described in step (2), for the map2style design based on the EFPN encoder E, in order to inject meaningful features into the generator, a style encoder of map2style is used, which consists of a fully connected layer (FC) and a ReLU (rectified linear unit nonlinearity) to form a 5-layer perceptron. The input is the multi-scale image feature F extracted. i (i = 0, 1, 2, 3, 4), and the output is the style feature for the generator, denoted as w i (i = 0, 1, 2, 3, 4).
[0019] In the above solution, the purpose of the generator G in step (2) is to map the multi-resolution and multi-scale feature information extracted by the EFPN encoder to the reconstructed three-dimensional image layer by layer. The multi-scale image features of the porous medium extracted by the encoder all have rich semantic information. In particular, the high-resolution features in the shallow layer have texture and edge information features. Through the learnable affine transformation "Α", the multi-scale feature information w i (i = 0, 1, 2, 3, 4) is converted into y i =(y is , y ib )(i = 0, 1, 2, 3, 4). The purpose is to control the adaptive instance normalization after each layer of convolution in the generation network. The dimension of y i is twice the number of feature maps on that layer. The mathematical representation of the AdaIN operation in this network is:
[0020]
[0021] where each feature map x i is normalized respectively, and then the style y of this scale is used iThe corresponding components scale and bias the feature map, and through this operation, the features of the input image are added to the reconstructed three-dimensional structure. During the reconstruction process, the AdaIN operation is used to map the image features of different resolutions and scales to the reconstructed three-dimensional image layer by layer. First, a constant with a size of 4×4×4 is initialized, and the first-layer image with a size of 6×6×6 is obtained through three-dimensional transposed convolution. Then, InstanceNorm is performed, and through the AdaIN operation, the extracted high-level semantic features of the image are added to determine the reconstruction target. Finally, the ReLU activation function is used to reduce information loss. Since the reconstruction target has been determined by the high-level semantic information in the first layer, the second layer is reconstructed based on the first layer. In this way, on the basis of retaining the image feature information of the first layer, the reconstruction of the second-layer image is completed, and the image becomes 10×10×10. The continuity between layers of the porous medium is naturally inherited in this operation, making the generated three-dimensional structure have better connectivity; and so on, for the third layer to the fifth layer, the operations are the same. Especially when generating a high-resolution image in the fifth layer, finer texture and edge detail features of the image are added, achieving a better purpose of modifying the morphology of the porous medium; in the last layer, only three-dimensional transposed convolution and the hyperbolic tangent function Tanh are used to facilitate restricting the output value to the range of -1 to 1.
[0022] In the above solution, the three discriminators D xy 、D yz 、D zx For homogeneous and isotropic porous media, the three-dimensional structure has texture and statistical characteristics similar to those of a two-dimensional image in the xy, yz, and zx orthogonal cross-sections. The purpose of designing the three discriminators D xy 、D yz 、D zx is to judge the degree of similarity between the cross-sections of the generated three-dimensional image x G in the three orthogonal directions and the real two-dimensional image.
[0023] In the above solution, the network structure EFPN-StyleGAN described in step (2) includes an encoder E based on an Extended Feature Pyramid (EFPN), a generator G, and three discriminators D xy 、D yz 、D zx , ① The main function of the encoder E is to extract multi-scale features of the two-dimensional image x; ② The main function of the generator G is to map the multi-resolution and multi-scale features of the extracted two-dimensional image to the three-dimensional image x G step by step; ③ The three discriminators D xy 、D yz 、Dzx The main function is to identify the generated three-dimensional image x G The similarity degree between the cross-sections of the three orthogonal directions and the real two-dimensional image.
[0024] In the above scheme, the loss function of the network structure EFPN-StyleGAN described in step (3) The purpose is to enable the three-dimensional image generated by the generator to learn the morphology and multi-scale distribution characteristics of the real image, and improve the ability of the discriminator to distinguish true from false. In order to prevent gradient explosion, gradient disappearance of the generative adversarial network and improve the training speed, and achieve a good reconstruction effect, a gradient penalty-based method is used to constrain the discriminator, and the Lipschitz constraint is achieved by directly controlling the gradient norm of the discriminator output. The discriminator D xy Realize the similarity degree between the slices in the xy direction and the input image, D xy The objective function of is:
[0025]
[0026] Among them, x represents an input two-dimensional image, and "." represents the operation of taking the xy direction slice, S xy Represents the xy direction cross-section operation, ε xy ∈U[0,1], Represents a sample randomly selected in the xy direction, λ xy Is the penalty coefficient.
[0027] Thus, the discriminator D yz And D zx The objective functions of are respectively:
[0028]
[0029]
[0030] Among them, S yz 、S zx Represents the cross-section operations in the yz and zx directions, ε yz ,ε zx ∈U[0,1], λ yz And λ zx Are respectively the penalty coefficients.
[0031] Therefore, the adversarial loss between the generator and the three discriminators is:
[0032]
[0033] Among them, the weights of the constraints between the generator and the three discriminators are the same.
[0034] In the above solution, the mode density loss function of the EFPN-StyleGAN mode in step (4) The purpose is to make the reconstructed three-dimensional structure have a pattern distribution information similar to the input two-dimensional image. Therefore, the reconstructed three-dimensional structure is constrained from the perspective of the mode density function.
[0035] It is defined as:
[0036]
[0037] where x pattern represents the pattern distribution of the two-dimensional image, and (G(E(x))) pattern represents the pattern distribution of the generated three-dimensional image.
[0038] In the above solution, the total loss function of the network EFPN-StyleGAN in step (5) It is defined as:
[0039]
[0040] where λ Pattern represents the weight of the loss .
[0041] This invention is funded by the National Natural Science Foundation of China, "Research on 3D Reconstruction Technology of Core 2D Images Based on Hyperdimension (62071315)". BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a structural block diagram of a three-dimensional image reconstruction method for porous media based on multi-scale feature fusion according to the present invention;
[0043] Figure 2 is the backbone network structure of the encoder E in the present invention;
[0044] Figure 3 is the feature fusion structure of the encoder E in the present invention;
[0045] Figure 4 is the network structure diagram of the generator G in the present invention;
[0046] Figure 5 is the network structure diagrams of the three discriminators D xy , D yz , D zx in the present invention;
[0047] Figure 6 are schematic diagrams of the two-point correlation function, the linear path, and the two-point cluster function;
[0048] Figure 7 Visual comparison of the 3D reconstruction results of core images in the present invention;
[0049] Figure 8 Schematic diagram of the comparison of the middle five layers of the 3D reconstruction results of core images in the present invention;
[0050] Figure 9 Statistical parameter comparison of the 3D reconstruction results of core images in the present invention. Specific embodiments
[0051] The following uses specific embodiments and drawings to further illustrate the present invention. However, the described embodiments are only a specific and detailed description of the implementation method of the present invention, and should not be construed as any limitation to the protected content of the present invention. The provided drawings and the following described embodiments are to enable the present invention to be more completely and accurately understood by those skilled in the art.
[0052] Figure 1 In a three-dimensional image reconstruction method for porous media based on multi-scale feature fusion, it can be specifically divided into the following steps:
[0053] (1) Data acquisition, data preprocessing, and construction of a two-dimensional image dataset;
[0054] (2) Design a network structure (EFPN-StyleGAN) for reconstructing three-dimensional images from two-dimensional images based on an extended Feature Pyramid Network (FPN);
[0055] (3) Design the loss function of the EFPN-StyleGAN network The purpose is to enable the three-dimensional image generated by the generator to learn the morphology and multi-scale distribution characteristics of the real image, and improve the accuracy of the discriminator in distinguishing the generated image;
[0056] (4) Design the mode density loss function of the EFPN-StyleGAN network The purpose is to constrain the mode distribution of the generated three-dimensional image;
[0057] (5) Based on the above constructed dataset, network structure, and loss function, train the neural network until the generated three-dimensional image meets the target requirements, thereby obtaining the image three-dimensional reconstruction model EFPN-StyleGAN;
[0058] (6) Based on the described EFPN-StyleGAN model, complete the three-dimensional reconstruction of two-dimensional images.
[0059] Specifically, in step (1), considering the processing speed and video memory size of the current computer, in the present invention, as an implementation example, for the obtained original core CT images, binaryzation processing is performed, and different two-dimensional images are randomly intercepted along one direction, obtaining more than 8,000 two-dimensional image samples with a size of 128×128.
[0060] In step (2), the present invention constructs Figure 1 the structural block diagram of a three-dimensional image reconstruction method for porous media based on multi-scale feature fusion shown in the figure. This network structure includes an encoder E, a generator G, and three discriminators D xy 、D yz 、D zx , ① The main function of the encoder E is to extract multi-scale features of the two-dimensional image x; ② The main function of the generator G is to map the extracted multi-resolution and multi-scale features of the two-dimensional image to the three-dimensional image x G step by step; ③ The main functions of the three discriminators D xy 、D yz 、D zx are to identify the similarity between the cross-sections of the generated three-dimensional image x G in three orthogonal directions and the real two-dimensional images.
[0061] In step (2), the backbone network structure of the encoder E designed by the present invention is as shown in Figure 2 the figure. The main function of the encoder E is to extract features of different resolutions and scales of the two-dimensional image. The backbone network of the encoder selects ResNet50, and based on the design requirements of the reconstructed three-dimensional image, the backbone network ResNet50 is improved. Input a two-dimensional image with a size of 128×128. In the first stage, first use a 4×4 convolutional kernel on the backbone network ResNet50 to extract features from the original image to obtain the first-layer feature map C0, with a size of (66,66). Discard the pooling layer after the first layer of convolution in the original ResNet50 network, and embed the obtained high-resolution feature of the first layer into the bottom layer of the FPN as an extended feature, thus solving the problem that small pores occupy few pixels and making good use of high-resolution features at the same time. This feature contains semantic and position information of small pores as well as texture and edge information of large pores. In the second stage, use an adaptive average pooling to obtain a feature map with a size of (34,34), and then connect three residual modules. After passing through these three residual modules, the size of the feature map remains unchanged, but the network more accurately learns the features of this layer size, denoted as C1, and so on, to obtain C i (i = 2, 3, 4). C i(i = 0, 1, 2, 3, 4) is to obtain multi-scale features from the bottom-up structure of the backbone network, which is the input of FPN. And a new high-resolution extended feature map C0 introduced from the output of the convolution Conv0 of the backbone network, which contains more spatial and content information and plays a key role in capturing and locating small pores in subsequent reconstruction.
[0062] In the step (2), the encoder E feature fusion structure designed by the present invention is as Figure 3 shown. Figure 3 (a) is a schematic diagram of the feature fusion process. Considering the different sensitivities of feature maps of different scales and different layers, a feature fusion architecture is designed. These five features C i (i = 0, 1, 2, 3, 4) are sent as inputs to the EFPN encoder. The EFPN encoder network includes two branches: vertical and horizontal. The vertical path upsamples the high-level feature map of the porous medium with low spatial resolution and high semantic information to obtain a high-resolution feature map; the horizontal path fuses the EFPN feature map with the corresponding underlying features to obtain the required multi-scale fusion feature map. The EFPN encoder uses upsampling and summation operations to fuse different features in the image. Taking the fusion of the feature map C1 and P2 as an example, first, the number of channels of the feature map C1 is reduced by 1×1 convolution to obtain C1'. To make the feature map P2 the same size as C1', P2 is upsampled, and finally, C1' and the upsampled P2 are added to obtain the fused feature map. To eliminate the aliasing effect caused by upsampling, the fused feature map P2 is processed by a 3×3 convolution to obtain the feature map F1'. Drawing on the squeezing and excitation idea and attention mechanism in SENet, global average pooling is used to convert the feature map F1' from c×h×w to c×1×1, so that the receptive field of the feature map is extended from local to global, obtaining the final feature map F1. Thus, richer global context information is captured, breaking the limitation that convolution only extracts local information. The final feature map F i (i = 0, 1, 2, 3, 4) is expressed by the mathematical relationship as:
[0063]
[0064] Among them, it represents C i the i-th lateral connection feature map, P i represents the feature map upsampled from P i+1 to the size of C i , S up is the upsampling operation, and the upsampling method is bilinear interpolation. The size of the upsampled feature map is the size of the image it is to be fused with. f13×3 is a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1. Ave represents global average pooling, and the purpose is to convert the feature map from c×h×w to c×1×1. Indicates the fusion operation. Finally, the multi-scale feature F obtained by EFPN i (i = 0, 1, 2, 3, 4). Correspondingly, since C0 is the feature map of the high-resolution image, the fused F0 contains not only the spatial and content feature information of the high-resolution image, but also strong semantic feature information.
[0065] To inject meaningful features into the generator, the style encoder of map2style is used. Figure 3 (b) is the style structure diagram of map2style. It consists of a fully connected layer (FC) and ReLU (rectified linear unit nonlinearity) to form a 5-layer perceptron. The input is the extracted multi-scale image feature F i (i = 0, 1, 2, 3, 4), and the output is the style feature for the generator, denoted as w i (i = 0, 1, 2, 3, 4).
[0066] In step (2), the generator G network structure designed by the present invention is as Figure 4 shown. The purpose of the generator G is to map the multi-resolution and multi-scale feature information extracted by the EFPN encoder layer by layer into the reconstructed three-dimensional image. The multi-scale image features of the porous medium extracted by the encoder all have rich semantic information. In particular, the high-resolution features in the shallow layer have texture and edge information features. Through the learnable affine transformation "Α", the multi-scale feature information w i (i = 0, 1, 2, 3, 4) is converted to y i =(y is , y ib )(i = 0, 1, 2, 3, 4). The purpose is to control the adaptive instance normalization after each layer of convolution in the generation network. The dimension of y i is twice the number of feature maps on that layer. The mathematical representation of the AdaIN operation in this network is:[[]]
[0067]
[0068] where each feature map x i is normalized separately, and then the style y of that scale is used iThe corresponding components scale and bias the feature map, and through this operation, the features of the input image are added to the reconstructed three-dimensional structure. During the reconstruction process, the AdaIN operation is used to map the image features of different resolutions and scales to the reconstructed three-dimensional image layer by layer. First, a constant of size 4×4×4 is initialized, and the first-layer image of size 6×6×6 is obtained through three-dimensional transposed convolution. Then, InstanceNorm is performed, and through the AdaIN operation, the extracted high-level semantic features of the image are added to determine the reconstruction target. Finally, the ReLU activation function is used to reduce information loss. Since the reconstruction target has been determined by the high-level semantic information in the first layer, the second layer is reconstructed based on the first layer. In this way, on the basis of retaining the image feature information of the first layer, the reconstruction of the second-layer image is completed, and the image becomes 10×10×10. The continuity between layers of the porous medium is naturally inherited in this operation, making the generated three-dimensional structure have better connectivity. By analogy, the operations for the third to fifth layers are the same. Especially when generating a high-resolution image in the fifth layer, finer texture and edge detail features of the image are added, achieving a better purpose of modifying the morphology of the porous medium. In the last layer, only three-dimensional transposed convolution and the hyperbolic tangent function Tanh are used to facilitate restricting the output value to the range of -1 to 1.
[0069] In step (2), the three discriminators D xy , D yz , D zx designed in the present invention have a network structure as Figure 5 shown. For homogeneous and isotropic porous media, the three-dimensional structure has textures and statistical characteristics similar to those of two-dimensional images in the xy, yz, and zx orthogonal planes. The purpose of designing the three discriminators D xy , D yz , D zx is to judge the degree of similarity between the generated three-dimensional image x G in the three orthogonal planes and the real two-dimensional image. Therefore, the network structures of the three discriminators are the same. For a three-dimensional image with a size of 128×128×128 to be reconstructed, one image is extracted every two pixels along the xy, yz, and zx orthogonal directions, and a total of 192 two-dimensional sectional images are obtained. During the training of the discriminator, for each generated three-dimensional structure, the three discriminators respectively distinguish the difference between each sectional image and the two-dimensional image. The input of the discriminator is a true or false image with a size of 1×128×128, and the discriminator outputs a discrimination probability value.
[0070] In step (3), the loss function of the network structure EFPN-StyleGAN designed in the present invention The purpose is to enable the three-dimensional images generated by the generator to learn the morphology and multi-scale distribution characteristics of real images, and improve the ability of the discriminator to distinguish true from false. To prevent gradient explosion, gradient disappearance in the generative adversarial network and improve the training speed, achieving a good reconstruction effect, a gradient penalty-based method is adopted to constrain the discriminator, and the Lipschitz constraint is realized by directly controlling the gradient norm of the discriminator output.
[0071] Discriminator D xy To achieve the similarity between the xy-direction slices of the discriminator D and the input image, xy the objective function of D is:
[0072]
[0073] where x represents an input two-dimensional image, "." represents the operation of taking xy-direction slices, S xy represents the xy-direction slicing operation, ε xy ∈U[0,1], represents a sample randomly selected along the xy direction, λ xy is the penalty coefficient.
[0074] Thus, the objective functions of discriminators D yz and D zx are respectively:
[0075]
[0076]
[0077] where S yz 、S zx represent the yz- and zx-direction slicing operations, ε yz ,ε zx ∈U[0,1], λ yz and λ zx are respectively the penalty coefficients.
[0078] Therefore, the adversarial loss between the generator and the three discriminators is:
[0079]
[0080] Here, the constraint weights between the generator and the three discriminators are the same.
[0081] In step (4), the mode density loss function of the EFPN-StyleGAN network structure of the present design invention The purpose is to make the reconstructed three-dimensional structure have a pattern distribution information similar to that of the input two-dimensional image. Therefore, the reconstructed three-dimensional structure is constrained from the perspective of the pattern density function. It is defined as:
[0082]
[0083] where the pattern template size is set to 3.
[0084] In the step (5), the total loss function of the network structure EFPN-StyleGAN of the present design invention is defined as:
[0085]
[0086] where λ Pattern represents the weight of the loss , and here it is taken as 1.0×10 5 .
[0087] Several other important network structure parameters: the batch size is set to 16, the Adam optimizer is used, the learning rates for the generator and discriminator are set to 0.0001, and the penalty coefficients λ xy , λ yz and λ zx are all set to 10, and their parameters are continuously iteratively updated. During the training process, the discriminator is trained five times and then the generator and encoder are trained once.
[0088] In the step (6), the generator G model in EFPN-StyleGAN is used to reconstruct the input two-dimensional image x to obtain the three-dimensional structure x G :
[0089] x G = G(E(x)) (16)
[0090] Specifically, in order to verify the effectiveness of the method of the present invention, relevant experiments are carried out in the present invention.
[0091] Figure 6 The schematic diagrams of common evaluation index functions are given, including the two-point correlation function, the linear path, and the two-point cluster function.
[0092] Such as Figure 7As shown, where (a), (b), and (e) respectively represent a homogeneous two-dimensional image of the input, the corresponding target three-dimensional image, and the orthogonal sectional view of the target three-dimensional image. To compare the quality of the reconstruction results, a comparison was made with the method in the super-dimensional reconstruction in the literature [Xia Z, Teng Q, Wu X, et al. Three-dimensional reconstruction of porous media using super-dimension-based adjacent block-matching algorithm [J]. Physical Review E, 2021, 104(4): 045308.]. (c) and (d) are the three-dimensional images reconstructed by the method in the literature [Xia Z, Teng Q, Wu X, et al. Three-dimensional reconstruction of porous media using super-dimension-based adjacent block-matching algorithm [J]. Physical Review E, 2021, 104(4): 045308.] and our method respectively, and (f) and (g) are the orthogonal sectional views of the three-dimensional images reconstructed by the method in the literature [Xia Z, Teng Q, Wu X, et al. Three-dimensional reconstruction of porous media using super-dimension-based adjacent block-matching algorithm [J]. Physical Review E, 2021, 104(4): 045308.] and our method respectively. Visually, the three-dimensional structure reconstructed by our method is more similar to the real target three-dimensional structure, and the connectivity and morphological distribution characteristics of the target three-dimensional structure are well reconstructed. At the same time, the variability of pores of different sizes in the three-dimensional structure is also maintained, indicating that the method we proposed can well maintain homogeneity and pore connectivity.
[0093] Figure 8 For the target three-dimensional image, the three-dimensional image reconstructed by the method in the literature [Xia Z, Teng Q, Wu X, et al. Three-dimensional reconstruction of porous media using super-dimension-based adjacent block-matching algorithm [J]. Physical Review E, 2021, 104(4): 045308.] and the three-dimensional image reconstructed by our method, five images were taken from the middle for comparison of the sectional images. It can be seen that the pore morphology and distribution of the sectional view of the three-dimensional image reconstructed by our method are closer to the sectional view of the target three-dimensional image, further verifying the advantages of our method.
[0094] In addition to the visual comparison, we also conducted a quantitative parameter comparison, including S2, L, C2, and the local porosity distribution, as Figure 9 shown. To verify the stability of the method, a two-dimensional image of the input was reconstructed 10 times respectively by the method in the literature [Xia Z, Teng Q, Wu X, et al. Three-dimensional reconstruction of porous media using super-dimension-based adjacent block-matching algorithm [J]. Physical Review E, 2021, 104(4): 045308.] and our method. The target value, the 10 reconstruction results, the average value of the reconstruction results, and the average value of the reconstruction results of the method in the literature [Xia Z, Teng Q, Wu X, et al. Three-dimensional reconstruction of porous media using super-dimension-based adjacent block-matching algorithm [J]. Physical Review E, 2021, 104(4): 045308.] were compared. It can be seen that the four statistical description functions of the images reconstructed by our method well match the statistical description functions of the real images, indicating that the method of the present invention can accurately and stably reproduce the target structure.
[0095] In addition, for the reconstruction of an image with a size of 128×128×128, the present invention only requires 0.002 s on the GPU. Although the training time of the method in the literature [Xia Z, Teng Q, Wu X, et al. Three-dimensional reconstruction of porous media using super-dimension-based adjacent block-matching algorithm [J]. Physical Review E, 2021, 104(4): 045308.] is very long, the reconstruction speed has been greatly improved. Therefore, there is a great improvement for the generation of a large number of three-dimensional images.
[0096] The method proposed by Xia et al., see the reference "Xia Z, Teng Q, Wu X, et al. Three-dimensional reconstruction of porous media using super-dimension-based adjacent block-matching algorithm [J]. Physical Review E, 2021, 104(4): 045308.".
[0097] In summary, the present invention is an effective method for generating a three-dimensional image of a porous medium based on a two-dimensional image. The invention can serve and be applied to fields such as petroleum geology, soil, materials, etc., and has great value especially in practical applications such as oil and gas exploration, exploitation, and materials design.
[0098] The above embodiments are only preferred embodiments of the present invention and do not limit the technical solutions described in the present invention. Any technical solutions that can be achieved on the basis of the above embodiments without creative labor shall be regarded as falling within the protection scope of the present invention.
Claims
1. A three-dimensional image reconstruction method for porous media based on multi-scale feature fusion, characterized in that: It includes the following steps: (1) Data acquisition, data preprocessing, and construction of a two-dimensional image dataset; (2) Design a network structure (EFPN-StyleGAN) for reconstructing three-dimensional images from two-dimensional images based on an extended Feature Pyramid Network (FPN); (3) Design the loss function of the EFPN-StyleGAN network The purpose is to enable the three-dimensional images generated by the generator to learn the morphology and multi-scale distribution features of real images, and improve the accuracy of the discriminator in distinguishing the generated images; (4) Design the mode density loss function of the EFPN-StyleGAN network The purpose is to constrain the mode distribution of the generated 3D images; (5) Based on the constructed dataset, network structure, and loss function above, train the neural network until the generated three-dimensional image meets the target requirements, thereby obtaining the EFPN-StyleGAN model for image three-dimensional reconstruction; (6) Based on the EFPN-StyleGAN model described above, complete the three-dimensional reconstruction of two-dimensional images.
2. The three-dimensional image reconstruction method of porous media based on multi-scale feature fusion according to claim 1, characterized in that In the backbone network design of the EFPN-StyleGAN in step (2), the EFPN-StyleGAN network structure includes an encoder E based on an Extended Feature Pyramid Network (EFPN), a generator G, and three discriminators D xy , D yz , D zx . The main function of the EFPN encoder E is to extract features of different resolutions and scales of two-dimensional images. The backbone network of the encoder selects ResNet50. Based on the design requirements of reconstructing three-dimensional images, the backbone network ResNet50 is improved. In the first stage, first, a 4×4 convolutional kernel is used on the backbone network ResNet50 to extract features from the original image to obtain the first-layer feature map C0, with a size of (66, 66). The pooling layer after the first layer of convolution in the original ResNet50 network is discarded, and the obtained high-resolution feature of the first layer is embedded into the bottom layer of the FPN as an extended feature. In this way, the problem of small pores occupying small pixels is solved, and the high-resolution features are well utilized. This feature contains The semantic and location information of small pores and the texture and edge information of large pores; in the second stage, an adaptive average pooling is used to obtain a feature map with a size of (34, 34), and then three residual modules are connected. After passing through these three residual modules, the size of the feature map remains unchanged, but the network more accurately learns the features of this layer size, denoted as C1, and so on, to obtain C i (i = 2, 3, 4); C i (i = 0, 1, 2, 3, 4) are to obtain multi-scale features from the bottom-up structure of the backbone network, that is, the input of FPN; and a new high-resolution extended feature map C0 introduced from the output of the convolution Conv0 of the backbone network, which contains more spatial and content information and plays a key role in capturing and locating small pores in subsequent reconstruction.
3. The three-dimensional image reconstruction method of porous media based on multi-scale feature fusion according to claim 1, wherein In the feature fusion network design based on the EFPN encoder E in the network structure EFPN-StyleGAN described in step (2), considering the different sensitivities of feature maps of different scales and different layers in the EFPN encoder E, a feature fusion architecture is designed, and these five features C i (i = 0, 1, 2, 3, 4) are sent as inputs to the EFPN encoder. The EFPN encoder network includes two branches: vertical and horizontal. The vertical path obtains high-resolution feature maps by upsampling the high-level feature maps of the porous medium with low spatial resolution and high semantic information. The horizontal path fuses the EFPN feature maps with the corresponding low-level features to obtain the required multi-scale fusion feature maps. The EFPN encoder uses upsampling and summation operations to fuse different features in the image. Taking the fusion of feature maps C1 and P2 as an example, first, the number of channels of feature map C1 is reduced by 1×1 convolution to obtain C1'. To make the size of feature map P2 consistent with C1', P2 is upsampled. Finally, C1' and the upsampled P2 are added together to obtain the fused feature map. To eliminate the aliasing effect caused by upsampling, the fused feature map P2 is processed by 3×3 convolution to obtain feature map F1'. Drawing on the squeezing and excitation idea and attention mechanism in SENet, global average pooling is used to convert feature map F1' from c×h×w to c×1×1, so that the receptive field of the feature map is extended from local to global, obtaining the final feature map F1, thereby capturing richer global context information and breaking the limitation that convolution only extracts local information. The final feature map F i (i = 0, 1, 2, 3, 4) is expressed by the mathematical relationship as: Among them, C i represents the feature map of the i-th lateral connection, and P i represents the feature map upsampled from P i+1 to the size of C i , S up is the upsampling operation. The upsampling method is bilinear interpolation. The size of the feature map after upsampling is the size of the image it is to fuse. f1 3×3 is a convolutional layer with a convolutional kernel size of 3×3 and a stride of 1. Ave represents global average pooling, aiming to convert the feature map from c×h×w to c×1×1. represents the fusion operation. Finally, the multi-scale feature F i obtained by EFPN (i = 0, 1, 2, 3, 4). Correspondingly, since C0 is the feature map of the high-resolution image, the fused F0 contains not only the spatial and content feature information of the high-resolution image, but also strong semantic feature information.
4. The three-dimensional image reconstruction method of porous media based on multi-scale feature fusion according to claim 1, characterized in that In the map2style design based on the EFPN encoder E in the network structure EFPN-StyleGAN described in step (2), in order to inject meaningful features into the generator, the style encoding method of map2style is used. It consists of a fully connected layer (FC) and a ReLU (rectified linear unit nonlinearity) to form a 5-layer perceptron. The input is the multi-scale image feature F extracted. i (i = 0, 1, 2, 3, 4), and the output is the style feature for the generator, denoted as w i (i = 0, 1, 2, 3, 4).
5. The three-dimensional image reconstruction method of porous media based on multi-scale feature fusion according to claim 1, characterized in that The generator G in the network structure EFPN-StyleGAN described in step (2). The purpose of the generator is to map the multi-resolution and multi-scale feature information extracted by the EFPN encoder to the reconstructed 3D image step by step. The multi-scale image features of the porous medium extracted by the encoder all have rich semantic information. Especially, the high-resolution features in the shallow layer have texture and edge information features. The multi-scale feature information w i (i = 0, 1, 2, 3, 4) is converted to The purpose is to control the adaptive instance normalization after each convolution in the generation network, and the dimension of y i is twice the number of feature maps on that layer. The mathematical representation of the AdaIN operation in this network is: where each feature map x i is normalized separately and then scaled and offset using the corresponding component of the scale style y i In this operation, the features of the input image are added to the reconstructed three-dimensional structure. During the reconstruction process, the AdaIN operation is used to map the image features of different resolutions and scales to the reconstructed three-dimensional image layer by layer. First, a constant with a size of 4×4×4 is initialized, and the first layer of the image with a size of 6×6×6 is obtained through three-dimensional transposed convolution. Then, InstanceNorm is performed, and through the AdaIN operation, the extracted high-level semantic features of the image are added to determine the reconstruction target. Finally, the ReLU activation function is used to reduce information loss. Since the reconstruction target has been determined by the high-level semantic information in the first layer, the second layer is reconstructed based on the first layer. In this way, the reconstruction of the second layer of the image is completed while retaining the image feature information of the first layer, and the image becomes 10×10×10. The continuity between layers of the porous medium is naturally inherited in this operation, making the generated three-dimensional structure have better connectivity. By analogy, the operations for the third to fifth layers are the same. Especially when generating a high-resolution image in the fifth layer, finer texture and edge detail features of the image are added, which serves the purpose of better modifying the morphology of the porous medium. In the last layer, only three-dimensional transposed convolution and the hyperbolic tangent function Tanh are used to limit the output value to the range of -1 to 1.
6. The three-dimensional image reconstruction method of porous media based on multi-scale feature fusion according to claim 1, characterized in that The three discriminators D in the network structure EFPN-StyleGAN described in step (2) xy , D yz , D zx , for homogeneous and isotropic porous media, the three-dimensional structure has textures and statistical characteristics similar to two-dimensional images in the cross-sections in the three orthogonal directions of xy, yz, and zx. The role of designing the three discriminators D xy , D yz , D zx is to judge the degree of similarity between the cross-sections of the generated three-dimensional image x G in the three orthogonal directions and the real two-dimensional images.
7. The three-dimensional image reconstruction method of porous media based on multi-scale feature fusion according to claim 1, characterized in that Loss function of the EFPN-StyleGAN network described in step (3) The purpose is to enable the three-dimensional images generated by the generator to learn the morphology and multi-scale distribution features of real images, and improve the ability of the discriminator to distinguish true from false. To prevent gradient explosion, gradient disappearance of the generative adversarial network and improve the training speed, and achieve good reconstruction effects, a gradient penalty-based method is used to constrain the discriminator, and the Lipschitz constraint is realized by directly controlling the gradient norm of the discriminator output. Discriminator D xy To realize the similarity between the xy-direction slices of the discriminator and the input image, D xy The objective function of is as follows: where \(x\) represents a two - dimensional input image, "." represents the operation of taking slices in the \(x - y\) direction, \(S\) xy represents the slicing operation in the \(x - y\) direction, \(\epsilon\) xy \(\in U[0,1]\), represents a sample randomly selected in the \(x - y\) direction, \(\lambda\) xy is the penalty coefficient. Thus, the objective functions of the discriminators \(D\) yz and \(D\) zx are respectively: Among which S yz and S zx represent the cutting plane operations in the yz and zx directions. ε yz , ε zx ∈ U[0,1], λ yz and λ zx are respectively penalty coefficients. Therefore, the adversarial loss between the generator and the three discriminators is as follows: Among them, the weights of the constraints between the generator and the three discriminators are the same.
8. The three-dimensional image reconstruction method of porous media based on multi-scale feature fusion according to claim 1, characterized in that The mode density loss function of the EFPN-StyleGAN network described in step (4) The purpose is to make the reconstructed three-dimensional structure have similar mode distribution information to the input two-dimensional image. Therefore, the reconstructed three-dimensional structure is constrained from the perspective of the mode density function. It is defined as: where x pattern represents the pattern distribution of the two-dimensional image, and (G(E(x))) pattern represents the pattern distribution of the generated three-dimensional image.
9. The three-dimensional image reconstruction method of porous media based on multi-scale feature fusion according to claim 1, characterized in that Total loss function of the EFPN-StyleGAN network described in step (5) It is defined as: where λ Pattern represents the loss weight.