Medical image super-resolution reconstruction method, system, equipment and medium
By constructing a deep learning model that includes channel window attention, combining local window attention and cross-channel weight adjustment, we can accurately capture the global and local information of medical images, solve the problems of image detail information loss and low computational efficiency in existing technologies, and improve the quality and efficiency of medical image super-resolution reconstruction.
Patent Information
- Application Number
- CN202510718812.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-10-03
AI Technical Summary
Existing medical image super-resolution methods find it difficult to balance the fusion of local details and global information, resulting in insufficient structural fidelity and diagnostic value of reconstructed images, as well as low computational efficiency, and are unable to meet the high-precision and real-time requirements of clinical practice.
A deep learning model with channel window attention (Cwin) is constructed, which combines the local window attention mechanism and the cross-channel weight adjustment strategy. The model parameters are optimized through multi-domain feature fusion and multi-task joint loss function to generate high-quality high-resolution medical images.
It significantly improves the reconstruction accuracy and stability of medical images, enhances the structural fidelity and diagnostic value of images, and meets the high-precision and real-time requirements of clinical diagnosis.
Smart Images

Figure CN120747262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image super-resolution reconstruction, and in particular to a medical image super-resolution reconstruction method, system, equipment and medium. Background Art
[0002] Medical imaging plays a vital role in modern medical diagnosis and treatment. The primary methods for acquiring medical images include X-rays, magnetic resonance imaging (MRI), computed tomography (CT), ultrasound, and positron emission tomography (PET). These images provide physicians with essential information for disease diagnosis and treatment planning. Therefore, higher resolution and higher quality images provide physicians with clearer and more detailed anatomical structures, aiding in accurate diagnosis of various diseases and improving patient outcomes.
[0003] However, existing medical image super-resolution methods still have some limitations. Medical images have complex structures and unique texture features. When processing such images, existing models often find it difficult to simultaneously take into account the fusion of local details and global information, resulting in deficiencies in the structural fidelity and diagnostic value of the reconstructed images. In addition, traditional super-resolution methods usually only focus on pixel-level reconstruction errors, ignoring the semantic information and diagnostic efficacy of medical images. They are prone to introducing artifacts or losing key details, thus affecting doctors' accurate diagnosis of diseases. These problems restrict the application of super-resolution technology in clinical practice. There is an urgent need for an innovative method that can take into account the fusion of global and local features and ensure the fidelity of medical structures. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0005] In a first aspect, the present invention provides a medical image super-resolution reconstruction method, comprising constructing a super-resolution reconstruction model based on a target module;
[0006] Generate high-resolution medical images from input low-resolution medical images through super-resolution reconstruction models;
[0007] Optimize the super-resolution reconstruction model based on the difference between high-resolution medical images and true resolution images.
[0008] As a preferred embodiment of the medical image super-resolution reconstruction method of the present invention, generating a high-resolution medical image includes:
[0009] Perform feature extraction on the input low-resolution medical image to generate an initial feature map;
[0010] Performing a first preset processing on the initial feature map to output a deep feature map;
[0011] Performing a second preset processing on the deep feature map to generate an enhanced feature map;
[0012] Image reconstruction is performed based on the enhanced feature map to generate high-resolution medical images.
[0013] As a preferred solution of the medical image super-resolution reconstruction method of the present invention, wherein: the initial feature map is subjected to a first preset processing to output a deep feature map, including:
[0014] Perform multi-level deep feature learning on the initial feature map;
[0015] Capture local image details through local window attention mechanism;
[0016] Combined with cross-channel weight adjustment, global information fusion is enhanced to output deep feature maps.
[0017] As a preferred solution of the medical image super-resolution reconstruction method of the present invention, wherein: performing a second preset processing on the deep feature map to generate an enhanced feature map includes:
[0018] Perform multi-domain feature fusion on deep feature maps;
[0019] The spatial domain and frequency domain analysis are combined to extract complementary features and generate enhanced feature maps.
[0020] As a preferred solution of the medical image super-resolution reconstruction method of the present invention, wherein: the local window attention mechanism includes,
[0021] Divide the initial feature map into multiple local windows and expand the receptive field through window shift operation;
[0022] Combined with the channel self-attention mechanism, the weights of different channels are adaptively adjusted to enhance cross-channel information interaction.
[0023] As a preferred embodiment of the medical image super-resolution reconstruction method of the present invention, wherein: optimizing the super-resolution reconstruction model includes:
[0024] Optimize super-resolution reconstruction model parameters through multi-task joint loss function;
[0025] The loss functions include pixel loss, segmentation-perceptual loss, and mean square error loss.
[0026] As a preferred solution of the medical image super-resolution reconstruction method of the present invention, wherein: multi-domain feature fusion includes,
[0027] Extract local texture features through convolution operation in the spatial domain;
[0028] The global structural features are extracted through spectral transformation in the frequency domain, and the dual-domain features are concatenated and fused through the convolutional layer.
[0029] In a second aspect, the present invention provides a medical image super-resolution reconstruction system, comprising: a construction module for constructing a super-resolution reconstruction model based on a target module;
[0030] A generation module, configured to generate a high-resolution medical image from an input low-resolution medical image through a super-resolution reconstruction model;
[0031] An optimization module is used to optimize the super-resolution reconstruction model based on the difference between the high-resolution medical image and the true resolution image.
[0032] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0033] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when the computer program is executed by a processor.
[0034] Compared with existing technologies, the present invention has the following advantages: by constructing a deep learning model that includes channel window attention (Cwin), it effectively integrates the local window attention mechanism with the cross-channel weight adjustment strategy to accurately capture the global and local information of medical images; it introduces multi-domain feature fusion technology, combines spatial and frequency domain analysis to extract complementary image features, and significantly enhances feature expression capabilities; and utilizes a multi-task joint loss function to combine pixel-level reconstruction errors with the semantic consistency constraints of medical prior knowledge, optimize model parameters, and improve the structural fidelity and diagnostic value of reconstructed images. This method can generate high-quality medical images and provide strong support for accurate disease diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0036] Figure 1 Schematic diagram of the process of medical image super-resolution reconstruction method.
[0037] Figure 2 Schematic diagram of the framework of the super-resolution reconstruction model, where (a) is the framework diagram of the CGCwin module, (b) is the framework diagram of the FEB module, and (c) is a schematic diagram of the segmentation perception loss.
[0038] Figure 3 Illustration of the channel window attention (Cwin) module.
[0039] Figure 4 Synthesized image for the first case randomly selected from the OASIS dataset.
[0040] Figure 5 Synthesized image for the second case randomly selected from the OASIS dataset.
[0041] Figure 6 Synthesized image for a randomly selected case from the ACDC dataset. DETAILED DESCRIPTION
[0042] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0043] Example 1, with reference to Figure 1 , which is the first embodiment of the present invention, provides a medical image super-resolution reconstruction method, comprising:
[0044] S100: Build a super-resolution reconstruction model based on the target module;
[0045] S200: generating a high-resolution medical image by using a super-resolution reconstruction model to input a low-resolution medical image;
[0046] S300: Optimizing super-resolution reconstruction models based on the difference between high-resolution medical images and true-resolution images.
[0047] It should be noted that in the field of medical image super-resolution reconstruction, although existing technical methods have made certain progress, they still face many challenges in practical applications. First, the acquisition process of medical images is limited by many factors, such as the resolution of the imaging equipment, scanning time, and the patient's physiological movement, resulting in the acquired medical images often having low resolution and unable to meet the high-precision requirements of clinical diagnosis. Secondly, when processing complex medical images, existing super-resolution reconstruction methods often find it difficult to balance the accuracy of the image's detailed information and overall structure, and are prone to artifacts, blurring, and other phenomena, which affect the image quality and diagnostic value. In addition, traditional super-resolution methods usually require long computing time and high computing resources, making it difficult to meet the real-time requirements in actual clinical applications. Therefore, the development of an efficient, accurate, and reliable medical image super-resolution reconstruction method is of great clinical significance.
[0048] Therefore, in response to the problems existing in the super-resolution reconstruction of medical images described above, through steps S100-S300, a super-resolution reconstruction model is constructed based on the target module, and the input low-resolution medical image is processed by the model, so that high-quality high-resolution medical images can be effectively generated. At the same time, based on the difference between the generated high-resolution medical image and the true resolution image, the super-resolution reconstruction model is optimized to further improve the reconstruction accuracy and stability of the model. The method of the present invention can effectively solve the problems existing in the prior art such as loss of image detail information, inaccurate overall structure, and low computational efficiency, and provides a new solution for super-resolution reconstruction of medical images, which helps to improve the accuracy and reliability of clinical diagnosis and provide more powerful support for the treatment and rehabilitation of patients.
[0049] Example 2, reference Figures 1 to 3 , which is an embodiment of the present invention, provides a medical image super-resolution reconstruction method based on the above embodiments.
[0050] In the embodiment of the present application, the super-resolution reconstruction model is constructed based on the target module in step S100, including the following steps A1-A2:
[0051] A1: The target modules include the feature extraction module (SFE module), the deep feature learning module (CGCwin module), the multi-domain fusion module (FEB module), and the image reconstruction module (IR module);
[0052] A2: Constructing a super-resolution reconstruction model includes sequentially connecting the extraction module (SFE module) - the deep feature learning module (CGCwin module) - the multi-domain fusion module (FEB module) - the image reconstruction module (IR module).
[0053] Specifically, in step A1, ① each CGCwin module consists of N Cwin blocks, where each Cwin block consists of a Cwin attention block, Layer Norm and MLP.
[0054] It should be noted that the MLP layer maps features from arbitrary size to g, which is then concatenated with the original channels. After N Cwin block operations, a 2D convolutional layer is used to restore the original channel size C = n.
[0055] ②FEB module: It consists of two main parts: the upper and lower parts. The lower part performs traditional spatial convolution operations, while the upper part uses fast Fourier convolution (FFC). The outputs of the two parts are concatenated and subjected to convolution operations to obtain the final result.
[0056] It should be noted that the FEB module is responsible for perceptual fusion of learned features.
[0057] In an optional embodiment, the super-resolution reconstruction model constructed in step S100 can also be constructed based on a generative adversarial network (GAN), that is, a generator and a discriminator are constructed, a loss function including an adversarial loss is defined, and the generator is trained with paired image data to produce realistic images.
[0058] In another optional embodiment, the super-resolution reconstruction model constructed in step S100 can also be constructed by combining Transformer and convolutional neural network (CNN), that is, utilizing the global feature extraction advantages of Transformer and local feature extraction advantages of CNN, designing a feature fusion mechanism, and using a composite loss function to optimize the model to achieve high-quality reconstruction of medical images.
[0059] In the embodiment of the present application, in step S200, the high-resolution medical image is generated by using the super-resolution reconstruction model for the input low-resolution medical image, including the following steps B1-B4:
[0060] B1: Extract features from the input low-resolution medical image and generate an initial feature map;
[0061] It should be noted that generating the initial feature map refers to the input image (Where H, W, and C represent the height, width, and number of channels of the image, respectively) After extraction by the feature extraction module (SFE module), an abstract feature with C=n channels is obtained.
[0062] B2: Perform the first preset processing on the initial feature map and output a deep feature map;
[0063] Furthermore, in step B2, the following steps B21-B23 are included:
[0064] B21: Multi-level deep feature learning on the initial feature map;
[0065] It can be understood that the initial feature map is subjected to deep feature learning through N CGCwin modules.
[0066] B22: Capturing local image details through local window attention mechanism;
[0067] It should be noted that the local window attention mechanism includes,
[0068] Divide the initial feature map into multiple local windows and expand the receptive field through window shift operation;
[0069] Combined with the channel self-attention mechanism, the weights of different channels are adaptively adjusted to enhance cross-channel information interaction.
[0070] B23: Combined with cross-channel weight adjustment to enhance global information fusion and output deep feature maps.
[0071] B3: Perform the second preset processing on the deep feature map to generate an enhanced feature map;
[0072] Furthermore, in step B3, the following steps B31-B32 are included:
[0073] B31: Multi-domain feature fusion of deep feature maps;
[0074] It should be noted that multi-domain feature fusion includes,
[0075] Extract local texture features through convolution operation in the spatial domain;
[0076] The global structural features are extracted through spectral transformation in the frequency domain, and the dual-domain features are concatenated and fused through the convolutional layer.
[0077] B32: Combine spatial domain and frequency domain analysis to extract complementary features and generate enhanced feature maps.
[0078] B4: Image reconstruction based on enhanced feature maps to generate high-resolution medical images.
[0079] It can be understood that the enhanced feature image is finally reconstructed into a ×4 super-resolution image through residual connection fusion
[0080] Specifically, in step B2, given the input It is divided into two branches. One of the branches first applies the LayerNorm operation to make the feature map more conducive to learning. Then, the Patch partitioning operation divides the feature map into M2 Small pieces and in Perform local window self-attention on (as shown in formula (1));
[0081] At the same time, another branch processes the input X through the LayerNorm block to obtain the feature map C. The GAP operation compresses the feature map into Then, channel self-attention with a variable kernel (as defined in Equation (2)) is applied to obtain Finally, applying the element-wise product operation yields X, X″ and Z c Add and input into LM module for perception learning, and we get
[0082] Similarly, the structure of the latter part is similar to that of the former part. The input feature X is processed by the shift window self-attention to obtain X′, and then processed by the reverse shift operation to obtain X″. The final output feature is
[0083] Furthermore, window self-attention means that when calculating window attention, there are M 2 We reshape the input feature X′ into It is also equal to K(key) and V(value). Attention Matrix It is calculated by interactively calculating the dot product of the query and the key. The specific formula for calculating window attention is as follows:
[0084]
[0085] Where d is the dimension of Q(query) or K(key). Since the relative position along each axis ranges from [-M+1, M-1], a smaller bias matrix is parameterized And the values in B are taken from B'.
[0086] Furthermore, channel attention is implemented through one-dimensional convolution to achieve cross-channel information interaction, which is specifically expressed as follows:
[0087] Z c =σ(C1d(G1,1(z),k))·z (2)
[0088] Where: z and Z c are the input and output, σ is the Sigmoid function, and C1d is a one-dimensional convolution with an adaptive kernel k. The size of the convolution kernel is adaptively determined by the function, where the kernel size k represents the range of local cross-channel interaction. Given the channel dimension C, the convolution kernel size k can be adaptively determined by formula (3):
[0089]
[0090] Where: |t| odd In this step, all experiments are set to 1 for γ and b.
[0091] In step B3, the input deep feature map is fed into two different domains: the spatial domain and the frequency domain. The spatial domain captures the spatial information of the image, while the frequency domain captures the contextual information in the spectrum. To further enhance the model's expressive power, residual connections and convolutional layers are inserted, surpassing the capabilities of a single convolutional layer. LeakyReLU operations are performed between convolutional layers.
[0092] In the frequency module, a 2D Fast Fourier Transform (FFT) is used to convert traditional spatial features to the frequency domain to extract global information. Subsequently, an inverse 2D FFT operation is applied to recover the frequency domain features. A convolution operation halves the number of channels. Finally, these features are concatenated and convolved with the spatial domain features to generate the output enhanced feature map.
[0093] In an optional embodiment, the high-resolution medical image generated in step S200 can also be generated through super-resolution based on a Laplacian pyramid network. Specifically, the image is decomposed into sub-band images at multiple scales. Super-resolution reconstruction is performed on each sub-band image, and then the sub-band images are synthesized to obtain the high-resolution medical image.
[0094] In another optional embodiment, the high-resolution medical image generated in step S200 can also be generated through super-resolution based on a diffusion model, that is, the low-resolution medical image is downsampled and denoised, and then the noise is gradually removed and upsampled through the diffusion model to finally generate a high-resolution image.
[0095] In the embodiment of the present application, in step S300, optimizing the super-resolution reconstruction model based on the difference between the high-resolution medical image and the true resolution image includes the following steps C1-C2:
[0096] C1: Optimize super-resolution reconstruction model parameters through a multi-task joint loss function;
[0097] C2: The loss function includes pixel loss, segmentation perception loss and mean square error loss.
[0098] Specifically, in step C2, the pixel loss (L Pixel ): used to measure super-resolution image I SR With high-resolution real image I HR The differences between them are shown in the following forms:
[0099] L Pixel=||I SR -I HR ||1 (4)
[0100] Segmentation-aware loss (L Seg ):Used to ensure super-resolution image I SR Segmentation features and high-resolution images in I HR The features in are similar. SR and I HR Input the pre-trained segmentation encoder ε to obtain the feature map I′ SR and I′ HR The calculation formula of segmentation perception loss is:
[0101] L Seg =||ε(I SR )-ε(I HR )||2=||I′ SR -I′ HR ||2 (5)
[0102] Mean square error loss (L Mse ): used to measure super-resolution image I SR With high-resolution real image I HR The pixel-by-pixel mean square error between , is expressed as:
[0103] L Mse =||I SR -I HR ||2 (6)
[0104] Furthermore, the total loss L Total The definition is as follows:
[0105] L Total =L Pixel +λ Seg L Seg +λ Mse L Mse (7)
[0106] Where: L Pixel is the pixel loss, L Seg is the segmentation-aware loss, L Mse is the mean square error loss, and λ Seg and λ Mse Is the weight parameter of the balance loss, here we take λ Seg is 0.1, λ Mse is 0.5.
[0107] In an optional embodiment, the super-resolution reconstruction model optimization in step S300 can also be optimized by employing a meta-learning algorithm. Specifically, by training the model on multiple related datasets, the initial distribution of model parameters is learned, enabling the model to quickly adapt and efficiently learn when faced with new super-resolution tasks. Under the supervision of a small amount of new data, the model parameters can be quickly adjusted to achieve super-resolution reconstruction for specific medical imaging devices or specific image types, thereby improving the model's generalization and adaptability.
[0108] In another optional embodiment, the super-resolution reconstruction model optimized in step S300 can also be optimized using a transfer learning-based optimization method. Specifically, a deep learning model pre-trained on a large, general-purpose image dataset is used as the initial model for transfer learning. The feature extraction layer of the pre-trained model is frozen or fine-tuned, and then a decoder and loss function suitable for the medical image super-resolution task are added. Through transfer learning, the model can utilize the rich features in general-purpose images, reducing its dependence on large-scale medical image data while improving the model's generalization ability and its ability to capture medical image details.
[0109] In summary, the medical image super-resolution reconstruction method proposed in this paper constructs a deep learning model that includes channel window attention (Cwin), effectively integrating the local window attention mechanism with the cross-channel weight adjustment strategy to accurately capture the global and local information of medical images. It also introduces multi-domain feature fusion technology, combining spatial and frequency domain analysis to extract complementary image features, significantly enhancing feature expression capabilities. Utilizing a multi-task joint loss function, it combines pixel-level reconstruction errors with the semantic consistency constraints of medical prior knowledge to optimize model parameters and improve the structural fidelity and diagnostic value of the reconstructed images. This method can generate high-quality medical images and provide strong support for accurate disease diagnosis.
[0110] Example 3 is the third embodiment of the present invention. In order to verify the advantages of the present invention, a quantitative evaluation was performed, and several standard indicators were calculated to compare the performance of the super-resolution (SR) method and Cwin-Net (this method). Peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), root mean square error (RMSE), and pixel domain visual information fidelity (VIFP) are used as quantitative indicators to evaluate the performance of SR methods. Higher 3D-PSNR, slice-PSNR, SSIM, and VIFP values indicate that the super-resolution image is very close to the reference image, while a lower RMSE value indicates better performance. It should be noted that when calculating these indicators, both this method and the comparison method require paired images.
[0111] In addition to comparing the PSNR / SSIM of super-resolution (SR) images, we also use the DICE score to evaluate the realism of the reconstructed images of Cwin-Net and the comparison methods. The DICE score reflects the similarity between the reconstructed image and the real data, thus more comprehensively evaluating the performance of the model in capturing details and structural fidelity.
[0112] Part I: Evaluation on the OASIS dataset.
[0113] To evaluate the effectiveness of the proposed method (Cwin-Net) in improving the performance of super-resolution (SR) algorithms, we compared our method with other SR methods, including ESRGAN, StarSRGAN, EDSR, RGT, and SwinIR. To ensure a fair comparison, the number of training iterations for all models was set to 100K, and the super-resolution magnification for all models was ×4.
[0114] ① Quantitative comparison - the results are shown in Table 1:
[0115] Table 1 SR results of this method and the comparison method
[0116] Method Slice-PSNR↑ SSIM↑ RMSE↓ VIFP↑ DICE↑ 3D-PSNR↑ HR \ \ \ \ 0.9411 \ Bicubic 30.8206 0.8572 7.7325 0.8116 0.9188 30.0454 ESRGAN 31.5953 0.9093 7.0276 0.8546 0.9180 30.8924 StarSRGAN 32.1637 0.9203 6.6059 0.8542 0.9192 31.4089 EDSR 35.5973 0.9533 4.4984 0.9515 0.9270 34.6780 RGT 35.1148 0.9488 4.6987 0.9296 0.9255 34.3660 SwinIR 36.5803 0.9590 4.0874 0.9559 0.9281 35.4313 Cwin-Net 36.8317 0.9596 3.9964 0.9579 0.9295 35.6113
[0117] As shown in Table 1, the best results are marked in bold. In the ×4 SR experiments conducted on the OASIS dataset, our proposed method (Cwin-Net) performed best. Compared with other methods, our method shows advantages in all five indicators, and the 3D reconstructed images generated by our model are 0.18dB higher than SwinIR in 3D-PSNR and 0.2514dB higher in slice-PSNR, showing significant improvement. In addition, our model also achieved a higher DICE score than SwinIR, further verifying its fidelity. However, it should be noted that the DICE scores of all SR images are still significantly lower than those of high-resolution (HR) images.
[0118] ②Perspective comparison:
[0119] Combined with attachment Figure 4 and Figure 5 As shown in Figure 2, two cases are randomly selected from the OASIS dataset and the synthetic images generated by various methods are compared.
[0120] Figure 4For Example 1, the results of different methods on random slices are visualized as follows: The first row shows the generated images from various methods, with yellow and cyan text indicating the best and suboptimal results, respectively. The second row compares the error plots for each method. The third row shows the segmentation results of the generated images, where red areas indicate segmentation errors and red text highlights the best segmentation result. Note that all segmentation results are expected to have scores lower than the ground truth segmentation.
[0121] Figure 5 In Case 2, visualizations of the random slices generated by different methods are shown below: The first row shows the generated images from different methods, with yellow and cyan text indicating the best and suboptimal results, respectively. The second row compares the error plots for each method. The third row shows the segmentation results of the generated images, where red areas indicate segmentation errors and red text highlights the best segmentation result.
[0122] It should be noted that in the second row of the error graph, the red arrows point to areas with significant differences, where darker areas indicate smaller deviations from the ground truth, reflecting more realistic generated details. In the third row, the red areas highlight segmentation errors, and fewer red areas indicate a higher degree of match with the true value. It is worth noting that even the high-resolution (HR) segmentation images contain some red areas because the Unet model used cannot completely match the true labels. However, if the segmentation results of other methods are close to the HR segmentation results, it indicates that they generate better details and higher quality results.
[0123] Observational results show that ESRGAN and StarSRGAN exhibit significant blurring, while EDSR introduces spurious texture details. Meanwhile, RGT and SwinIR fall short in recovering realistic details. Compared to other methods, our Cwin-Net excels in capturing the complex structures and subtle texture details in brain tissue, providing greater clarity while preserving the sharpness of gray and white matter boundaries and demonstrating greater realism in the process.
[0124] Part II: Evaluation on the ACDC dataset.
[0125] To demonstrate the effectiveness and robustness of our method in processing different categories of images, we conducted additional experiments using the ACDC dataset. The upscaling factor for super-resolution was kept at 4x, and the training steps were consistent. In addition, we evaluated not only PSNR and SSIM, but also the Dice score of the images.
[0126] ① Quantitative comparison - the results are shown in Table 2:
[0127] Table 2 SR results of this method and the comparison method
[0128]
[0129] As shown in Table 2, the best results are highlighted in bold. In ×4 super-resolution experiments on the ACDC dataset, we observed that our proposed method performed best. Compared to other methods, our method outperformed them across five metrics. However, the scores of the various methods were relatively close, which can be attributed to the challenging nature of single-image super-resolution (SR) for cardiac MRI. Motion artifacts caused by patient arrhythmias during scanning are individualized and difficult to remove. Significant improvements in PSNR and SSIM were difficult to achieve. Compared to SwinIR, our method improved PSNR by 0.0552 dB and SSIM by 0.0015. Furthermore, due to the limited training data (only 1462 slices), overfitting may occur. Dropout was used to minimize overfitting. As a result, despite similar DICE scores, our method still outperformed SwinIR by 0.0008. However, it should be noted that the DICE scores of all methods are expected to be lower than those of the high-resolution (HR) ground truth.
[0130] ②Perspective comparison:
[0131] Combined with attachment Figure 6 As shown in Figure 2, a case was randomly selected from the ACDC dataset and the synthetic images generated by various methods were compared. The results generated by ESRGAN show obvious artifacts, while StarSRGAN exhibits significant blurring. On the other hand, EDSR, RGT, and SwinIR fall short in recovering true details. Compared to other methods, our Cwin-Net performs well in capturing cardiac structures and subtle textures. We achieve higher scores in segmentation prediction while also demonstrating better realism.
[0132] Part III: Ablation experiment.
[0133] Two independent ablation experiments were conducted to evaluate the relative contributions of each proposed objective module. The first ablation experiment aimed to understand the impact of different modules on overall model performance. The second ablation experiment focused on the impact of different loss functions on model training. To ensure a fair comparison, all methods used the same infrastructure and training configuration. The model was trained for 100,000 iterations on the OASIS dataset. See Table 3 for details.
[0134] Table 3 Experimental results under different modules and various loss functions
[0135]
[0136] As shown in Table 3, when using the CGCwin architecture, the PSNR is improved by 0.2681dB compared to when not using it. The further implementation of the Cwin attention module improves the model by 0.0207dB. When all modules are enabled, the PSNR and SSIM scores reach the highest values. It should be noted that the models trained with different modules are all in L Pixe Training is performed with loss.
[0137] It is worth noting that models d, e, and f all contain the same modules. When adopting different loss functions, we start from model d and add segmentation-aware loss. As shown in Table 3, model e improves PSNR by 0.0414dB compared to training with only LPixel loss. Mse loss and assigned appropriate training loss weights, and finally achieved the best PSNR result of 36.8317dB.
[0138] Example 4. The above is an illustrative scheme of a medical image super-resolution reconstruction method. It should be noted that the technical scheme of this medical image super-resolution reconstruction system and the technical scheme of the medical image super-resolution reconstruction method described above are based on the same concept. For details not described in detail in the technical scheme of the medical image super-resolution reconstruction system in this embodiment, please refer to the description of the technical scheme of the medical image super-resolution reconstruction method described above.
[0139] This embodiment further provides a medical image super-resolution reconstruction system, comprising:
[0140] A construction module, used to construct a super-resolution reconstruction model based on the target module;
[0141] A generation module, configured to generate a high-resolution medical image from an input low-resolution medical image through a super-resolution reconstruction model;
[0142] An optimization module is used to optimize the super-resolution reconstruction model based on the difference between the high-resolution medical image and the true resolution image.
[0143] This embodiment also provides an electronic device suitable for super-resolution reconstruction of medical images, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the medical image super-resolution reconstruction method proposed in the above embodiment.
[0144] This embodiment further provides a storage medium storing a computer program, which, when executed by a processor, implements the medical image super-resolution reconstruction method proposed in the above embodiment.
[0145] The storage medium proposed in this embodiment and the method for realizing super-resolution reconstruction of medical images proposed in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0146] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general hardware, and of course can also be implemented by hardware. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0147] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A medical image super-resolution reconstruction method, characterized by: include, Build a super-resolution reconstruction model based on the target module; Generate a high-resolution medical image by using the super-resolution reconstruction model to input a low-resolution medical image; The super-resolution reconstruction model is optimized based on the difference between the high-resolution medical image and the true resolution image.
2. The medical image super-resolution reconstruction method according to claim 1, wherein: The generating of high-resolution medical images comprises, Perform feature extraction on the input low-resolution medical image to generate an initial feature map; Performing a first preset processing on the initial feature map to output a deep feature map; Performing a second preset processing on the deep feature map to generate an enhanced feature map; Image reconstruction is performed based on the enhanced feature map to generate a high-resolution medical image.
3. The medical image super-resolution reconstruction method according to claim 2, wherein: The first preset processing is performed on the initial feature map to output a deep feature map, including: Performing multi-level deep feature learning on the initial feature map; Capture local image details through local window attention mechanism; Combined with cross-channel weight adjustment, global information fusion is enhanced to output deep feature maps.
4. The medical image super-resolution reconstruction method according to claim 3, wherein: The performing a second preset processing on the deep feature map to generate an enhanced feature map includes: Performing multi-domain feature fusion on the deep feature map; The spatial domain and frequency domain analysis are combined to extract complementary features and generate enhanced feature maps.
5. The medical image super-resolution reconstruction method according to claim 4, wherein: The local window attention mechanism includes: Divide the initial feature map into multiple local windows and expand the receptive field through window shift operation; Combined with the channel self-attention mechanism, the weights of different channels are adaptively adjusted to enhance cross-channel information interaction.
6. The medical image super-resolution reconstruction method according to claim 5, wherein: The optimizing the super-resolution reconstruction model comprises: Optimizing the super-resolution reconstruction model parameters through a multi-task joint loss function; The loss functions include pixel loss, segmentation perception loss and mean square error loss.
7. A medical image super-resolution reconstruction method according to any one of claims 4 to 6, characterized in that: The multi-domain feature fusion, include, Extract local texture features through convolution operation in the spatial domain; The global structural features are extracted through spectral transformation in the frequency domain, and the dual-domain features are concatenated and fused through the convolutional layer.
8. A medical image super-resolution reconstruction system, applying the method according to any one of claims 1 to 7, characterized in that: include: A construction module, used to construct a super-resolution reconstruction model based on the target module; A generating module, configured to generate a high-resolution medical image from an input low-resolution medical image using the super-resolution reconstruction model; An optimization module is used to optimize the super-resolution reconstruction model based on the difference between the high-resolution medical image and the true resolution image.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.