A rock CT two-dimensional image super-resolution reconstruction method based on multi-scale structure feature aggregation
By using a multi-scale structural feature aggregation method, the problems of insufficient structural detail recovery, limited multi-scale modeling capability, and insufficient delineation of salient regions in super-resolution reconstruction of rock CT images are solved, achieving high-precision reconstruction of rock CT images and improving image quality and analysis accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN UNIV OF SCI & TECH
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-09
AI Technical Summary
Existing super-resolution reconstruction methods for rock CT images suffer from insufficient image quality and analytical accuracy due to unclear restoration of rock pore and fracture boundaries, inadequate multi-scale structural modeling capabilities, insufficient delineation of salient regions, and insufficient structural consistency constraints.
A method based on multi-scale structural feature aggregation is adopted, which realizes high-precision reconstruction of low-resolution rock CT images through a structural scale perception module (SSA), a multi-stage feature extraction backbone, a structural inference bottleneck module (SRB), and a decoding module. This includes shallow feature extraction, spatial weighted modulation, multi-scale feature fusion, and global structural inference.
It significantly improves the edge clarity and detail reproduction of rock microstructures, ensures the topological continuity and physical authenticity of cross-scale structures, and enhances the accuracy and reliability of digital core analysis.
Smart Images

Figure CN122175786A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing and deep learning technology, specifically relating to rock CT image enhancement and digital core analysis, and particularly to a super-resolution reconstruction method for two-dimensional rock CT images based on multi-scale structural feature aggregation. This method utilizes the structure perception and structural relationship modeling mechanisms of deep learning networks to achieve high-precision restoration and reconstruction of microscopic pores and fractures in low-resolution rock CT images. Background Technology
[0002] With the increasing demand for detailed characterization of unconventional oil and gas, geothermal energy, and underground reservoirs, digital core analysis and quantitative characterization of pore structure in rocks have been widely applied in rock physics research and reservoir evaluation. Computed tomography (CT), as a non-destructive imaging technique, can acquire three-dimensional information about the internal microstructure of rocks, providing crucial data support for the calculation and prediction of key parameters such as porosity, connectivity, and permeability.
[0003] However, in actual acquisition processes, due to limitations such as CT equipment resolution, scan time, radiation dose, field of view, and sample size, obtaining high-resolution rock CT images is often costly and inefficient. Furthermore, high-resolution scanning increases data volume, thereby raising storage and computational costs. Therefore, super-resolution (SR) technology, which reconstructs low-resolution rock CT images into high-resolution images, has become an important approach to improve the observability and analytical accuracy of rock microstructures.
[0004] Existing super-resolution reconstruction methods mainly include traditional interpolation methods and learning-based methods. Traditional interpolation methods, such as bilinear interpolation and bicubic interpolation, are simple to implement, but often result in blurred edges and loss of detail, making it difficult to effectively recover fine structures such as pores and cracks. In recent years, deep learning-based super-resolution methods have learned the mapping relationship from low resolution to high resolution through convolutional neural networks or attention mechanisms, achieving certain results in the field of natural images and are gradually being applied to CT image enhancement and reconstruction tasks.
[0005] However, rock CT images exhibit typical multi-scale complex structural features: pores, microfractures, and their boundary morphology are small and unevenly distributed, while structural continuity and connectivity have a significant impact on subsequent quantitative analysis. Current techniques for super-resolution reconstruction of rock CT images still suffer from one or more of the following limitations: 1. Insufficient ability to restore structural details: The boundaries of rock pores and fractures are usually thin and have low contrast. Existing methods often lack targeted structural detail enhancement strategies in the process of high-frequency detail reconstruction, which can easily lead to over-smoothing or the introduction of artifacts, resulting in unclear structural boundaries and affecting the accuracy of pore geometry and connectivity determination. 2. Limited multi-scale structural modeling capabilities: Pores and cracks exhibit cross-scale distribution characteristics, and there is a coupling relationship between small-scale details and large-scale structures; some methods are insufficient in multi-scale feature fusion, making it difficult to simultaneously take into account local detail enhancement and global structural consistency, which may lead to problems such as scale inconsistency or structural fracture. 3. Insufficient characterization of salient structural regions: Regions of great significance for analysis in rock CT images are usually concentrated in salient structural regions such as pore edges, fracture tips, and thin-walled connecting channels; existing models lack explicit perception or guidance mechanisms for structural edges, and are easily interfered with by large background or uniform regions during feature extraction, resulting in insufficient response and recovery capabilities for key structural regions. 4. Insufficient constraints on structural consistency and usability: In addition to visual quality, rock CT super-resolution results should also ensure the authenticity and continuity of physical structure. Some methods lack effective modeling or constraints on structural topological relationships or cross-scale geometric correlations, which may lead to phenomena such as non-physical discontinuity of pores and boundary drift in the reconstruction results, thereby reducing the reliability of digital core analysis.
[0006] Therefore, there is an urgent need in this field for a new super-resolution reconstruction technology for rock CT images that can address the complex multi-scale structural characteristics of rock micropores and fissures. This technology should explicitly enhance structurally significant regions during feature extraction and fusion, and introduce a structural relationship modeling mechanism in the cross-scale aggregation stage. This would enable shallow structural information to effectively participate in deep structural reasoning and fusion, thereby improving cross-scale structural consistency and reducing structural ambiguity and disconnection. This would achieve high-precision structural restoration and reconstruction of low-resolution rock CT images, and improve the accuracy and robustness of subsequent quantitative characterization and physical property analysis of pore structures. Summary of the Invention
[0007] To overcome the problems of insufficient structural detail recovery, limited multi-scale structural modeling capability, insufficient characterization of salient structural regions, and insufficient structural consistency constraints in the super-resolution reconstruction of two-dimensional rock CT images, this invention proposes a super-resolution reconstruction method for two-dimensional rock CT images based on multi-scale structural feature aggregation. The technical solution of this invention to solve the above-mentioned technical problems is as follows: This invention provides a method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation, comprising: S100: Acquire the low-resolution CT 2D image of the rock to be reconstructed and preprocess it to obtain the network input; S200: The network input is fed into the super-resolution reconstruction network, and shallow features are extracted within the super-resolution reconstruction network to obtain shallow features; S300: Input the shallow features into the structure scale sensing module SSA, and perform spatial weighted modulation on the shallow features to obtain structure-aware features; S400: Input the structure-aware features into the multi-stage feature extraction backbone to extract at least two stages of multi-scale stage features; S500: Input the structure-aware features and the multi-scale stage features into the structure reasoning bottleneck module SRB to perform global structure reasoning and obtain cross-scale structure-consistent aggregated features. S600: The aggregated features with consistent cross-scale structure are input into the decoding module for decoding to obtain decoded features. The decoded features are then fused with shallow features and input into the reconstruction module to output a high-resolution two-dimensional rock CT image.
[0008] Furthermore, the preprocessing in step S100 includes normalization processing.
[0009] Further, in step S300, obtaining structure-aware features using the Structure Scale Awareness (SSA) module includes: S310: Average the shallow features in the channel dimension to obtain a single-channel structure map; S320: Calculate the horizontal and vertical gradient responses of the single-channel structure diagram using the Sobel operator and obtain the structure strength diagram; S330: The structural strength map is mapped to a spatial weight map via a convolutional network and then normalized using Sigmoid; S340: Spatially weight the shallow features based on the spatial weights in the form of residuals to obtain the structure-aware features.
[0010] Furthermore, the multi-stage feature extraction backbone includes at least two cascaded stage units, each stage unit consisting of multiple NSTBs and patch-merging modules; wherein, the NSTB sequence is used to output stage features, and the patch-merging is used to downsample the stage features to achieve scale transformation.
[0011] Furthermore, the NSTB employs a scaled cosine attention mechanism and a post-normalization structure; the patch-merging achieves downsampling by merging adjacent tiles to reduce spatial resolution and adjusts the channel dimensions.
[0012] Furthermore, the Structural Reasoning Bottleneck Module (SRB) is configured to perform node-level relationship modeling, and to achieve global structural reasoning by performing alignment and multi-scale fusion operations on the structurally aware features and the multi-scale stage features.
[0013] Furthermore, the global structural reasoning using the Structural Reasoning Bottleneck Module (SRB) in step S500 includes: S510: Project the features, pool the projected features to obtain a node grid representation, and perform average fusion and / or weighted fusion on the multi-scale features. S520: Map the multi-scale stage features of at least two stages to node representations; S530: Construct an adaptive adjacency matrix based on node similarity, wherein the node similarity is obtained by dot product operation or matrix multiplication; S540: Perform graph convolutional inference based on the adaptive adjacency matrix to update node features; S550: Map the updated node features back to the feature space dimension to obtain the projected features; S560: Using the highest resolution feature among the multi-scale stage features as the query vector, and the projection feature as the key vector, perform cross-scale attention interaction, and after residual connection and normalization processing, obtain the aggregated feature with consistent cross-scale structure.
[0014] Further, the decoding and image reconstruction in step S600 include: S610: Align and fuse the aggregated features with consistent cross-scale structure and the first-stage features in the multi-scale stage features to obtain intermediate fused features; S620: The intermediate fusion features are then superimposed with the shallow features through global residual connections after NSTB sequence decoding and layer normalization. S630: Input the superimposed features into the reconstruction module for pixel domain reconstruction and output the high-resolution rock CT two-dimensional image.
[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) Significantly improves the edge clarity and detail restoration capability of rock microstructure. Addressing the issues of blurred pore edges and low texture contrast in rock CT images, this invention introduces a Structure Scale Awareness (SSA) module, utilizing local gradient information as prior guidance to spatially weighted modulate shallow features. This mechanism explicitly enhances the feature response of pore edges and particle contours, effectively compensating for the lack of high-frequency structural information characterization in the shallow feature extraction stage of traditional methods. This results in more accurate restoration of the geometric shape and boundary sharpness of micropores in the reconstruction results, reducing excessive image smoothing.
[0016] (2) Effectively ensures the topological continuity and physical realism of cross-scale structures. Addressing the problems of pore discontinuity, structural artifacts, and scale inconsistencies that easily occur in existing deep learning methods when processing complex rock structures, this invention designs a structural inference bottleneck module (SRB). This module overcomes the limitations of the receptive field of traditional convolutions, explicitly inferring the geometric relationships between features at different scales within the node space through a graph feature propagation mechanism, significantly enhancing the model's ability to model global contextual information and cross-scale structural continuity. This allows the reconstructed high-resolution image to maintain the topological connectivity of the global structure while preserving local details, avoiding non-physical structural breaks.
[0017] (3) Improved reliability of digital core quantitative analysis. By combining the above-mentioned multi-scale structural feature aggregation mechanism, this invention can reconstruct high-resolution images from low-resolution CT images with high quality without increasing additional scanning costs. This high-fidelity structural restoration not only improves the visual quality of the images, but also provides a more accurate and reliable data foundation for subsequent digital core analysis tasks such as porosity calculation, connectivity analysis, and seepage simulation. Attached Figure Description
[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a flowchart illustrating the rock CT two-dimensional image super-resolution reconstruction method provided in an embodiment of the present invention; Figure 2 A block diagram of the overall structure of the super-resolution reconstruction network provided in an embodiment of the present invention; Figure 3 This is a structural schematic diagram of the Structural Scale Sensing Module (SSA) provided in an embodiment of the present invention. Figure 4 This is a structural schematic diagram of the Structural Reasoning Bottleneck Module (SRB) provided in an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be further described below with reference to the accompanying drawings. It should be understood that the embodiments of the present invention are only used to explain the present invention and are not intended to limit the scope of protection of the present invention.
[0020] like Figure 1 As shown in the figure, this embodiment provides a method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation, which includes the following steps: S100: Acquire the low-resolution CT 2D image of the rock to be reconstructed and preprocess it to obtain the network input; S200: The network input is fed into the super-resolution reconstruction network, and shallow features are extracted within the super-resolution reconstruction network to obtain shallow features; S300: Input the shallow features into the structure scale sensing module SSA, and perform spatial weighted modulation on the shallow features to obtain structure-aware features; S400: Input the structure-aware features into the multi-stage feature extraction backbone to extract at least two stages of multi-scale stage features; S500: Input the structure-aware features and the multi-scale stage features into the structure reasoning bottleneck module SRB to perform global structure reasoning and obtain cross-scale structure-consistent aggregated features. S600: The aggregated features with consistent cross-scale structure are input into the decoding module for decoding to obtain decoded features. The decoded features are then fused with shallow features and input into the reconstruction module to output a high-resolution two-dimensional rock CT image.
[0021] In step S100, the preprocessing of the low-resolution rock CT 2D image is as follows: In one possible implementation, a low-resolution CT two-dimensional image of the rock to be reconstructed is first acquired. To reduce the impact of differences in scanning batches or grayscale dynamic range on network inference, [the following measures were taken]. Perform normalization processing to obtain the network input image. For example, pixel values can be linearly scaled to a range. Alternatively, the images can be standardized according to the statistics of the dataset to ensure that the input distribution is consistent during network training.
[0022] In step S200, shallow feature extraction is performed to obtain the following shallow features: like Figure 2 As shown, the network input image The input is fed into the shallow feature extraction layer of the super-resolution reconstruction network. This shallow feature extraction layer is a convolutional layer that performs a convolution operation on the input image, mapping the image data from pixel space to a high-dimensional feature space. This extracts shallow features containing low-frequency information such as texture and edges. Performing convolution operations yields shallow features. .
[0023] The processing steps of the S300 structural scale sensing module are as follows: Figure 3 As shown, the method includes the following steps: S310: SSA first requires aggregating feature channel information. Channel-dimensional average pooling is performed on shallow features to compress multi-channel features into a single-channel structure map. The calculation formula is as follows: in, For the first Channel characteristics, Characterizes the strength distribution of the basic structure.
[0024] S320: To explicitly capture the edges of microfractures and pores in rock images, the Sobel operator is used to calculate the horizontal and vertical gradient responses of the single-channel structure map, and the resulting structure strength map is synthesized. : in, This is a numerical stability term used to prevent numerical instability when the gradient is zero.
[0025] S330: The extracted gradient information is transformed into attention weights usable by the network. A lightweight mapping network consisting of two convolutional layers and ReLU activation is used to map the structure intensity map to a spatial weight map, and then normalized using the Sigmoid activation function. in, , For convolution operations, For the Sigmoid function, the Spatial weight distribution used to characterize salient regions of structure.
[0026] S340: SSA performs point-by-point spatial weighted modulation of the input features in the form of residuals based on the generated spatial weight map to obtain structure-aware features: in, This is the modulation intensity coefficient, used to control the weighted amplitude.
[0027] Step S400 involves inputting the structure-aware features into the multi-stage feature extraction backbone for processing. The method is as follows: like Figure 2 As shown, the backbone network adopts a hierarchical structure, and in each stage unit, multiple NSTBs are used to perform deep modeling of features. NSTB utilizes a scaled cosine attention mechanism instead of traditional dot product attention, which is more adaptable to the complexity of rock textures. Through cascading... Each NSTB sequence outputs features for each stage, thus obtaining a multi-scale stage feature set with at least two stages. The attention formula is as follows: in, Representing the query matrix and key matrix cosine similarity, For learnable scalars, This represents the relative positional deviation. Patch-merging is used to process adjacent stage units. By merging adjacent patches, the spatial resolution of the features is halved, while the channel dimension is expanded, achieving a feature pyramid structure. Its computational mapping is as follows: in, Output features for the current stage. These serve as input features for the next stage.
[0028] Step S500 inputs the structure-aware features and the multi-scale stage features into the structure inference bottleneck module SRB for global structure inference, obtaining cross-scale structure-consistent aggregated features, such as... Figure 4 As shown, the Structural Reasoning Bottleneck Module (SRB) achieves global structural reasoning through node-level relationship modeling. The method includes the following steps: S510: Project the features, pool the projected features to obtain a node mesh representation, and perform average fusion and / or weighted fusion on the multi-scale features. First, the features from the first... Features of each encoder stage Reconstructed into spatial feature maps and mapped to a unified feature dimension via convolution. To obtain the aligned features : Subsequently, adaptive average pooling is performed on each feature map layer to divide the feature map into... The grid region is used to construct node grid representations at various scales. : Finally, the node meshes at each scale are averaged and aggregated to form a global node representation that includes multi-scale information. : S520: Map the multi-scale stage features of at least two stages to node representations, and in order to incorporate shallow structural information, use structure-aware features. Mapped to the node space and associated with the global node representation. The nodes are then fused to obtain the final node representation used for reasoning. .
[0029] Specifically, for Perform projection and pooling operations And introduce fusion weights Execute weighted fusion: in, As learnable parameters or preset hyperparameters, this step ensures that the prior knowledge of rock edge and pore structure is explicitly included in the node features.
[0030] S530: Construct an adaptive adjacency matrix based on node similarity, where node similarity is obtained by dot product or matrix multiplication. To establish associations between nodes, the fused node mesh is first... Expanding in the spatial dimension yields a node sequence. : Subsequently, an adaptive adjacency matrix is constructed by calculating the dot product similarity between nodes. This describes the long-distance dependencies between different regions (nodes) in an image. S540: Perform graph convolutional inference based on the adaptive adjacency matrix to update node features, utilizing the constructed adjacency matrix. Information is propagated across the entire graph. This is achieved through a learnable linear weight matrix. Linear mapping is applied to the node features, and graph convolutional inference is performed in conjunction with the adjacency matrix to obtain the updated node features. : This step aggregates global contextual information, enhancing the structural consistency of features.
[0031] S550: Map the updated node features back to the feature space dimension to obtain projected features, and then update the node sequence after inference. Remapping back to the spatial dimensions of the original features yields the projected features. This process typically involves linear transformations and dimension reshaping operations: S560: Using the highest resolution feature among the multi-scale stage features as the query vector, and the projection feature as the key vector, perform cross-scale attention interaction, and after residual connection and normalization processing, obtain the aggregated feature with consistent cross-scale structure.
[0032] To inject the global structural information obtained from inference into the image reconstruction process, a cross-scale attention mechanism is adopted.
[0033] Based on the characteristics of the first stage in the backbone network As a query vector, Query, with structural reasoning features As keys and values, calculate attention and aggregate features: Finally, aggregate features With the original input Residual connections are performed, and after layer normalization, the final cross-scale structurally consistent aggregated features are output. : This feature is then fed into the decoding module for high-resolution image reconstruction.
[0034] Step S600 involves inputting the aggregated features with consistent cross-scale structure into the decoding module for decoding to obtain decoded features, and then fusing the decoded features with shallow features and inputting the fused features into the reconstruction module to output a high-resolution two-dimensional rock CT image. The method includes the following steps: S610: Aggregate the consistent features of the cross-scale structure The first-stage features in the multi-scale stage features are aligned and fused to obtain intermediate fused features.
[0035] S620: After the intermediate fused features are decoded by NSTB sequence and processed by layer normalization, they are connected to the shallow features through global residual connections. Superimpose layers to introduce shallow spatial details.
[0036] S630: The superimposed features are input into the reconstruction module, which performs spatial resolution enhancement mapping using convolutional layers and pixel-shuffle layers, and finally outputs a high-resolution 2D rock CT image. The pixel shuffling layer is used to reorganize information from the channel dimension to the spatial dimension, thereby achieving sub-pixel-level image reconstruction and restoring a clear rock microstructure.
[0037] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation, characterized in that, include: S100: Acquire the low-resolution CT 2D image of the rock to be reconstructed and preprocess it to obtain the network input; S200: The network input is fed into the super-resolution reconstruction network, and shallow features are extracted within the super-resolution reconstruction network to obtain shallow features; S300: Input the shallow features into the structure scale sensing module SSA, and perform spatial weighted modulation on the shallow features to obtain structure-aware features; S400: Input the structure-aware features into the multi-stage feature extraction backbone to extract at least two stages of multi-scale stage features; S500: Input the structure-aware features and the multi-scale stage features into the structure reasoning bottleneck module SRB to perform global structure reasoning and obtain cross-scale structure-consistent aggregated features. S600: The aggregated features with consistent cross-scale structure are input into the decoding module for decoding to obtain decoded features. The decoded features are then fused with shallow features and input into the reconstruction module to output a high-resolution two-dimensional rock CT image.
2. The method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation according to claim 1, characterized in that, The preprocessing in step S100 includes normalization.
3. The method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation according to claim 1, characterized in that, The step S300, which involves acquiring structural sensing features using the Structural Scale Awareness (SSA) module, includes: S310: Average the shallow features in the channel dimension to obtain a single-channel structure map; S320: Calculate the horizontal and vertical gradient responses of the single-channel structure diagram using the Sobel operator and obtain the structure strength diagram; S330: The structural strength map is mapped to a spatial weight map via a convolutional network and then normalized using Sigmoid; S340: Spatially weight the shallow features based on the spatial weight map in the form of residuals to obtain the structure-aware features.
4. The method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation according to claim 1, characterized in that, The multi-stage feature extraction backbone includes at least two cascaded stage units, each stage unit consisting of multiple NSTBs and patch-merging modules; wherein, the NSTB sequence is used to output stage features, and the patch-merging is used to downsample the stage features to achieve scale transformation.
5. The method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation according to claim 4, characterized in that, The NSTB employs a scaled cosine attention mechanism and a post-normalization structure; the patch-merging achieves downsampling by merging adjacent tiles to reduce spatial resolution and adjusts the channel dimensions.
6. The method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation according to claim 1, characterized in that, The Structural Reasoning Bottleneck Module (SRB) is configured to perform node-level relationship modeling. It achieves global structural reasoning by performing alignment and multi-scale fusion operations on the structurally aware features and the multi-scale stage features.
7. The method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation according to claim 6, characterized in that, The global structural reasoning using the Structural Reasoning Bottleneck Module (SRB) in step S500 includes: S510: Project the features, pool the projected features to obtain a node grid representation, and perform average fusion and / or weighted fusion on the multi-scale features. S520: Map the multi-scale stage features of at least two stages to node representations; S530: Construct an adaptive adjacency matrix based on node similarity, wherein the node similarity is obtained by dot product operation or matrix multiplication; S540: Perform graph convolutional inference based on the adaptive adjacency matrix to update node features; S550: Map the updated node features back to the feature space dimension to obtain the projected features; S560: Using the highest resolution feature among the multi-scale stage features as the query vector, and the projection feature as the key vector, perform cross-scale attention interaction, and after residual connection and normalization processing, obtain the aggregated feature with consistent cross-scale structure.
8. The method for super-resolution reconstruction of two-dimensional rock CT images based on multi-scale structural feature aggregation according to claim 1, characterized in that, The decoding and image reconstruction in step S600 include: S610: Align and fuse the aggregated features with consistent cross-scale structure and the first-stage features in the multi-scale stage features to obtain intermediate fused features; S620: The intermediate fusion features are then superimposed with the shallow features through global residual connections after NSTB sequence decoding and layer normalization. S630: Input the superimposed features into the reconstruction module for pixel domain reconstruction and output the high-resolution rock CT two-dimensional image.