Medical image super-resolution method based on feature aggregation
Through the feature aggregation method of the LGFANet model, the calculation complexity and resource requirements of the super-resolution model of medical images are solved, and efficient image reconstruction effect is achieved, especially in the details and depth feature processing of CT images.
Patent Information
- Application Number
- CN202510341016.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
Existing medical image super-resolution models have problems with computational complexity and resource requirements, and lack multi-level joint processing for image depth features and detail enhancement.
The LGFANet model is adopted, and the image reconstruction module is combined with the shallow and deep feature extraction module to realize local and global feature separation and reaggregation, reducing the amount of parameters and calculation complexity, while improving the reconstruction accuracy.
Shows higher reconstruction accuracy and lower computational cost on multiple benchmark datasets, suitable for efficient image reconstruction in medical environments.
Smart Images

Figure CN120259085A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and particularly to a medical image super-resolution method based on feature aggregation. Background Art
[0002] In the field of medical imaging, image super-resolution (SR) technology is usually used for post-processing of medical images. When existing SR models are applied in actual medical environments, they often face problems such as high computational complexity and large hardware resource requirements.
[0003] In recent years, researchers have proposed a variety of lightweight SR models to solve the above problems. For example, the FS RCNN model significantly improves the computational speed by removing bicubic interpolation preprocessing and using deconvolution operations for upsampling. The CARN model achieves feature reuse through a cascading mechanism and residual network design, maintaining excellent performance while reducing parameters and computational volume. The IDN and IMDN models utilize channel splitting and hierarchical distillation techniques to extract more feature information at different levels, further improving the performance of the models. The FALSR model introduces neural architecture search into the SR task and automatically generates an efficient model through multi-objective optimization and elastic search strategies, achieving performance comparable to state-of-the-art models under limited computational resources.
[0004] Although the existing technologies have solved the problems of computational complexity and resource requirements to a certain extent, most models focus on a single feature extraction strategy and lack multi-level joint processing for in-depth features and detail enhancement of images. There is still much room for improvement, especially in the refined processing of local and global information.
[0005] Therefore, there is an urgent need for a medical image super-resolution method to solve the problems of single feature extraction strategy and high computational cost during the super-resolution process. Summary of the Invention
[0006] In view of this, the present invention discloses a medical image super-resolution method based on feature aggregation to solve the above problems, including: processing a low-resolution medical image using LGFANet to obtain a high-resolution medical image with super-resolution reconstruction;
[0007] Further, LGFANet includes: a shallow feature extraction module for performing shallow feature extraction on a low-resolution medical image to obtain shallow features; a deep feature extraction module for further feature extraction of the shallow features to obtain deep features; and an image reconstruction module for processing the fused features of the shallow features and the deep features to obtain a high-resolution medical image.
[0008] The present invention adopts a lightweight architecture. Through the method of separating, extracting and then aggregating local and global features, each module works together to achieve the efficient reconstruction of complex CT medical images, reducing the number of model parameters and computational complexity. Compared with existing lightweight super-resolution models, the present invention solves the problem of single feature extraction strategy, shows higher reconstruction accuracy on multiple benchmark datasets, and has lower deployment and application costs in the medical environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is a schematic flowchart of LGFANet in an embodiment of the present invention;
[0010] Figure 2 It is a schematic flowchart of the local feature extraction module in an embodiment of the present invention;
[0011] Figure 3 It is a schematic flowchart of the global feature extraction module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] In order to make the purpose, technical solutions, features and advantages of the present invention clearer and more understandable, the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0013] Embodiment 1:
[0014] This embodiment includes a medical image super-resolution method based on feature aggregation, including: obtaining a low-resolution medical image to be processed; using LGFANet (Local and Global Feature Aggregation Network) to process the low-resolution medical image to obtain a high-resolution medical image with super-resolution reconstruction.
[0015] Specifically, LGFANet includes: a shallow feature extraction module for performing shallow feature extraction on the low-resolution medical image to obtain shallow features; a deep feature extraction module for further feature extraction of the shallow features to obtain deep features; an image reconstruction module for processing the fused features of the shallow features and the deep features to obtain a high-resolution medical image.
[0016] Further, the shallow feature extraction module includes a 3×3 convolutional layer; for the low-resolution (LR) image Use a 3×3 convolutional layer to extract shallow features Where H, W, C in and C respectively represent the height, width, number of input channels and number of output channels of the image. This stage can be expressed as:
[0017] F s = Conv 3×3 (ILR )
[0018] Among them, Conv 3×3 represents a 3×3 convolutional layer.
[0019] Furthermore, the shallow features are fed into the deep feature extraction module. The deep feature extraction module is composed of k cascaded Local and Global Feature Aggregation Blocks (LGFABs), a 3×3 convolutional layer, and a global residual connection module. Among them, the LGFAB is used to perform a more in-depth fusion of the local and global information of the input image. In this embodiment, k is preferably 8. The deep feature extraction module processes the shallow feature F s to obtain the deep feature which is expressed as:
[0020]
[0021] Among them, represents the k-th LGFAB.
[0022] Furthermore, as Figure 1 shown, the LGFAB includes a Local Feature Extraction Block (LFEB), a Global Feature Extraction Block (GFEB), and a Feature Fusion Block (FFB). The LGFAB realizes a more refined image representation by effectively extracting and fusing local and global features. Multiple LGFABs work in series. By repeatedly extracting and fusing local and global features, the reconstruction quality of the image is gradually improved.
[0023] Furthermore, the image features input to the LGFAB enter the LFEB and the GFEB respectively.
[0024] Specifically, the LFEB is composed of a Spatial Attention Block (SAB) and a depthwise separable convolution. Figure 2 is the flow schematic diagram of the SAB. The SAB performs average pooling and max pooling on the input feature X of the LGFAB module in the channel direction respectively. The feature maps obtained after pooling are and The formula is:
[0025] F avg = AvgPool(X)
[0026] F max = MaxPool(X)
[0027] Furthermore, for F avg and F max perform channel splicing, perform 7×7 convolution on the spliced features, and compress the number of channels to 1; apply the Sigmoid function to the result of the convolution for activation to generate a spatial attention map; use the spatial attention map to weight the input feature X and output the weighted feature The process of data processing by SAB is expressed by the formula as follows:
[0028] F SA =(f sig (Conv 7×7 (Concat(F avg ,F max ))))×X
[0029] where Concat, Conv 7×7 and F sig represent channel splicing, 7×7 convolution, and Sigmoid activation function respectively.
[0030] Furthermore, the depthwise separable convolution includes: sequentially performing depth convolution, 1×1 convolution, GELU activation, and 1×1 convolution on F SA to obtain the output F local of LFEB, and the formula is expressed as follows:
[0031] F local =Conv 1×1 (GELU(Conv 1×1 (DWConv(F SA ))))
[0032] where DWConv, Conv 1×1 and GELU represent depth convolution, 1×1 convolution, and GELU activation function respectively.
[0033] Furthermore, the image features input to LGFAB simultaneously pass through GFEB that is parallel to LFEB, and GFEB is used to capture the global context information in the image. Figure 3 is the process schematic diagram of GFEB. Specifically, GFEB linearly maps the input feature to the query matrix Q, the key matrix K, the value matrix V
[0034] and the intermediate matrix M. Among them, N = H×W represents the size of the feature map, and C represents the number of channels of the feature map. When obtaining M, it is necessary to perform max-pooling operation for downsampling, and compress its height and width by α times, and α is preferably 8 to reduce the computational complexity of subsequent matrix multiplication and Softmax, which is expressed as:
[0035] Q = ConvQ (X)
[0036] K = Conv K (X)
[0037] V = Conv V (X)
[0038] M = MaxPool(Conv M (X))
[0039] wherein, where n = H / α × W / α, representing the size of the feature map compressed by the pooling operation.
[0040] Furthermore, apply matrix multiplication to M and K, and normalize the result with the Softmax function to obtain the similarity matrix
[0041] S1 = Softmax(MK T )
[0042] Furthermore, apply matrix multiplication to S1 and V to obtain the intermediate global feature
[0043] X1 = S1V
[0044] Furthermore, apply matrix multiplication to M and Q, and normalize the result with the Softmax function to obtain the similarity matrix
[0045] S2 = Softmax(QM T )
[0046] Furthermore, apply matrix multiplication to S2 and X1, and perform a residual connection on the result and the input feature X to obtain the global feature
[0047] Y = S2X1 + X
[0048] Furthermore, the overall computational complexity Ω(GFEB) of GFEB can be expressed as:
[0049] Ω(GFEB) = 4NC 2 + 4nNC + 2nN
[0050] wherein, N represents the size of the feature map, C represents the number of channels of the feature map, and the complexity of GFEB is linearly related to N. In GFEB, by generating two different similarity matrices and performing two weighted operations successively, the model can more accurately model long-range dependencies, thereby enhancing the ability to capture global information. This design effectively reduces the computational overhead while ensuring the model reconstruction performance.
[0051] As Figure 1 shown, after being processed by LFEB and GFEB, the output features of the two modules will be further aggregated by the feature fusion module FFB.
[0052] Specifically, the local feature and the global feature sum is used as the input of FFB i.e., X in = F local + Y.
[0053] FFB uses a 1×1 convolutional layer and a GELU activation layer to preliminarily process X in to obtain the intermediate feature
[0054] X inter = GELU(Conv 1×1 (X in ))
[0055] Furthermore, X inter is successively subjected to depth convolution, 1×1 convolution, GELU activation, and 1×1 convolution, and residual connection is performed on the input of LGFAB to obtain the output fusion feature of LGFAB The whole process can be expressed as:
[0056] X out = Conv 1×1 (GELU(Conv 1×1 (DWConv(X inter )))+X inter )+X
[0057] Furthermore, the image features after deep feature extraction are fed into the image reconstruction module.
[0058] The image reconstruction module includes a 3×3 convolutional layer, a pixel mixing layer, and a 3×3 convolutional layer connected in sequence. Through the convolutional layer and sub-pixel convolution operation, a high-resolution medical image is generated.
[0059] Furthermore, the present invention adopts L1 loss (L1 Loss, absolute value loss). By calculating the sum of the absolute pixel differences between the predicted image and the real image, it pays attention to the differences at the detail level and has good adaptability in tasks with high requirements for details such as super-resolution reconstruction of medical CT images. Therefore, the present invention selects L1 loss as the objective function for optimizing network parameters, and the formula is as follows:
[0060]
[0061] Among them, IHR and I SR respectively represent the real high-resolution medical image and the high-resolution medical image generated by LGFANet.
[0062] Example 2:
[0063] This example includes a medical image super-resolution method based on feature aggregation, including: obtaining a low-resolution medical image to be processed; using LGFANet to process the low-resolution medical image to obtain a super-resolution reconstructed high-resolution medical image.
[0064] The difference from Example 1 is that the number k of cascades of the local and global feature aggregation modules of the LGFANet used in this example is 4.
[0065] Table 1. Comparison of the number of parameters and performance of LGFANet and various lightweight SR models in the prior art
[0066]
[0067]
[0068] Furthermore, in Table 1, Set5, Set14, BSD100, and Urban100 are 4 publicly available benchmark datasets in the field of super-resolution technology, and PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index Measure) are quantitative indicators for evaluating the performance of the model. The designed LGFANet of the present invention has a lower number of parameters compared to other models, and shows more excellent performance than other lightweight models at all magnification factors and datasets. The average PSNR of LGFANet on all datasets is improved by 0.09 dB compared to the LAPAR-A model.
[0069] Finally, it should be noted that the above only describes some embodiments of the present invention. For those skilled in the art, various changes, modifications, substitutions, and deformations can be conceived without departing from the principles and spirits of the present invention. The protection scope of the present invention is defined by the appended claims and their equivalents, and the above actions should all be covered within the protection scope of the present invention.
Claims
1. A medical image super-resolution method based on feature aggregation, characterized in that Including: Obtain a low-resolution medical image to be processed; use a pre-trained local and global feature aggregation network to process the low-resolution medical image to obtain a high-resolution medical image with super-resolution reconstruction. Among them, the local and global feature aggregation network includes: a shallow feature extraction module for performing shallow feature extraction on the low-resolution medical image to obtain shallow features; a deep feature extraction module for further feature extraction of the shallow features to obtain deep features; an image reconstruction module for processing the fused features of the shallow features and the deep features to obtain a high-resolution medical image.
2. The medical image super-resolution method based on feature aggregation according to claim 1, wherein The deep feature extraction module includes k consecutive local and global feature aggregation modules connected in series, a 3×3 convolutional layer, and a global residual connection module, which is expressed by the formula: Among them, F d represents the deep feature, F s represents the shallow feature, represents the k-th local and global feature aggregation module, Conv 3×3 () represents a 3×3 convolutional layer.
3. The method for super-resolution of medical images based on feature aggregation according to claim 2, wherein, The local and global feature aggregation module includes: a local feature extraction module, a global feature extraction module, and a feature fusion module. Among them, the local feature extraction module processes the input of the local and global feature aggregation module to obtain local features; the global feature extraction module processes the input of the local and global feature aggregation module to obtain global features; the feature fusion module processes the sum of the local features and the global features to obtain the output of the local and global feature aggregation module.
4. The medical image super-resolution method based on feature aggregation according to claim 3, wherein The processing of the local feature extraction module on the data includes: using a spatial attention module to preprocess the input data to obtain weighted features; performing a depthwise separable convolution operation on the weighted features to obtain local features; among them, the depthwise separable convolution operation includes: successively performing depth convolution, 1×1 convolution, GELU activation, and 1×1 convolution operations on the weighted features, which is expressed by the formula: F local = Conv 1×1 (GELU(Conv 1×1 (DWConv(F SA )))) Among them, F local represents the local feature, and F SA represents the weighted feature. DWConv, Conv 1×1 and GELU represent depth convolution, 1×1 convolution, and GELU activation function respectively.
5. The method for super-resolution of medical images based on feature aggregation according to claim 3, wherein The processing of the global feature extraction module on the data includes: Linearly mapping the input feature X of the global feature extraction module to a query matrix Q, a key matrix K, a value matrix V, and a mediation matrix M; Applying matrix multiplication to M and K, and normalizing the result with the Softmax function to obtain a similarity matrix S1; Applying matrix multiplication to S1 and V to obtain an intermediate global feature X1; Applying matrix multiplication to M and Q, and normalizing the result with the Softmax function to obtain a similarity matrix S2; Applying matrix multiplication to S2 and X1, and performing a residual connection on the result and the input feature X to obtain a global feature Y.
6. The method for super-resolution of medical images based on feature aggregation according to claim 3, wherein, The processing of the feature fusion module on the data includes: using a 1×1 convolutional layer and a GELU activation layer to preliminarily process the input feature of the feature fusion module to obtain an intermediate feature; successively performing depth convolution, 1×1 convolution, GELU activation, and 1×1 convolution on the intermediate feature, and performing a residual connection with the input of the local and global feature aggregation module to obtain the output of the feature fusion module, and the formula is: X out = Conv 1×1 (GELU(Conv 1×1 (DWConv(X inter )))+X inter )+X X inter = GELU(Conv 1×1 (X in )) Among them, X in represents the input feature of the feature fusion module, X out represents the output of the feature fusion module, X inter represents the intermediate feature, X represents the input of the local and global feature aggregation module, DWConv, Conv 1×1 and GELU respectively represent depth convolution, 1×1 convolution, and the GELU activation function.
7. The method for super-resolution of medical images based on feature aggregation according to claim 1, wherein LGFANet uses the absolute value loss as the loss function during the pre-training process.