3D Magnetic Resonance Super-Resolution Method Based on Cross-Modal and Cross-Scale Feature Fusion
By adopting a cross-modal and cross-scale feature fusion method in MRI image super-resolution reconstruction, combining attention mechanism and multi-level feature extraction, the problem of difficulty in making full use of MRI image prior information in the prior art is solved, and a higher quality MRI image super-resolution reconstruction is achieved.
Patent Information
- Application Number
- CN202310175391.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-02-28
AI Technical Summary
Existing super-resolution methods based on CNNs are difficult to fully mine the internal prior information of MRI images themselves and the external prior information of cross-modal MRI images, resulting in poor results in super-resolution reconstruction of MRI images.
The 3D magnetic resonance super-resolution method based on the fusion of cross-modal and cross-scale features is adopted. By constructing a cross-modal reference branch network and a main branch network, combining multiple residual channel attention blocks, cross-scale feature migration modules and attention mechanisms, the high-resolution features of the reference mode and target mode are captured and fused to improve the expression ability of the super-resolution network.
By fully utilizing the cross-modal and cross-scale features, the super-resolution reconstruction quality of MRI images is significantly improved, the high-resolution detail recovery ability of the image is enhanced, and the anatomical content details of the image and the recovery effect of the glioma part are improved.
Smart Images

Figure CN116029908B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of super-resolution in image processing, and particularly relates to a 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion. Background Art
[0002] Magnetic Resonance Imaging (MRI) is a multi-parameter non-invasive imaging technology. The quality of MRI images is affected by factors such as Signal to Noise Ratio (SNR), resolution, and scanning time. Usually, by increasing the slice thickness, a certain SNR requirement is ensured while reducing motion artifacts, but this will produce low-resolution MRI images and also limit the analysis accuracy.
[0003] Super Resolution (SR) is a technology that can break through hardware limitations and improve the spatial resolution of MRI images. SR is mainly divided into SR methods based on interpolation, SR methods based on reconstruction, and SR methods based on learning. SR methods based on interpolation are simple and efficient, but there will be obvious block effects, ringing effects, and jagged effects in the interpolated MRI images. SR methods based on reconstruction extract key information from low-resolution images and use prior knowledge to constrain the reconstruction process of high-resolution images. SR methods based on learning learn a certain mapping relationship between high-resolution images and low-resolution images from a large amount of training data and improve the resolution of the target low-resolution image according to the learned mapping relationship.
[0004] Convolutional Neural Networks (CNNs) have achieved success in natural image super-resolution. CNN-based SR methods can be divided into Single image super-resolution (SISR) methods and Reference-based image super-resolution (RefSR) methods. SISR reconstructs high-resolution MRI images from low-resolution MRI images by learning an end-to-end mapping function. The Super Nesolution Convolutional NeuralNetwork (SRCNN) is used to improve the super-resolution of 2D images. Subsequently, the SRCNN algorithm is extended from 2D to 3D and applied to the super-resolution task of 3D brain MRI images. Then, the idea of global and local residual skip connections is applied to the super-resolution reconstruction of 2D MRI slices, and a progressive residual network structure based on fixed skip connections is proposed.
[0005] To improve the expressive power of the SR network for MRI images, strategies such as multi-scale learning, attention mechanism, generative adversarial network, and multi-branch network are proposed. 3D channel and spatial feature attention are applied to improve the learning ability of the SR network to adjust features in different dimensions, enhancing valuable features while suppressing redundant information. A multi-level parallel convolutional and transposed convolutional network is used to utilize features from different branches. A deep channel division model is used for super-resolution of MRI image slices, dividing the feature maps into different branches to construct a multi-branch structure.
[0006] RefSR utilizes an additional high-resolution reference image to extract high-frequency information from the high-resolution reference image and reconstruct the corresponding high-resolution image from the low-resolution image. The high-resolution reference image and the low-resolution image are usually obtained from different angles of the same scene or different sequences of a video. MRI images are a multi-parameter imaging modality, and the imaging time required for different imaging parameters is different. The modality image with a short imaging time is used as the reference image to provide high-frequency information for the super-resolution of the modality image with a long imaging time.
[0007] Existing CNNs-based SR methods still lack in fully exploiting the internal prior information of MRI images themselves and the external prior information of cross-modal MRI images. Summary of the Invention
[0008] The present invention overcomes the above-mentioned deficiencies of the prior art and provides a 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion, which can fully exploit the external prior information of cross-modal MRI images and the internal prior information of MRI images for image super-resolution reconstruction.
[0009] The technical solution of the present invention is as follows: A 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion, characterized in that the operation is as follows: Input a low-resolution MRI image into a trained super-resolution network model based on cross-modal and cross-scale feature fusion to obtain the corresponding super-resolution MRI image;
[0010] The construction of the super-resolution network model based on cross-modal and cross-scale feature fusion includes the following steps:
[0011] S1: Construct a super-resolution network model based on cross-modal and cross-scale feature fusion;
[0012] The super-resolution network model includes a cross-modal reference branch network and a main branch network;
[0013] The cross-modal reference branch network includes a cross-modal reference branch network reference image gradient map extraction module and a cross-modal reference branch network reference image feature extraction module;
[0014] The main branch network includes a shallow feature extraction module of the main branch network, a deep feature extraction module of the main branch network, an upsampling and feature fusion module (UFF) of the main branch network, an image reconstruction module of the main branch network, and a high-resolution image output module of the main branch network;
[0015] S2: Obtain high-resolution MRI images of the target modality and the reference modality from the public dataset;
[0016] Among them, in the public dataset, T1W is used as the high-resolution MRI image of the target modality, and T2W and FLAIR are used as the high-resolution reference images of the reference modality;
[0017] S3: Perform simulated degradation preprocessing on the high-resolution MRI image data of the target modality to obtain the low-resolution MRI image data of the target modality;
[0018] S4: Input the high-resolution MRI image of the reference modality into the reference image gradient map extraction module of the cross-modal reference branch network to obtain the gradient map of the reference modality MRI image;
[0019] Construct a feature extraction module by combining a 3D convolutional layer and an activation function; Add the feature extraction module to the cross-modal reference branch network as the feature extraction module of the cross-modal reference branch network;
[0020] S5: Input the gradient map obtained in S4 into the feature extraction module of the reference branch network to capture the structural dependence and spatial relationship of the reference high-resolution image, and output the feature of the reference modality MRI image;
[0021] Add the feature extraction module to the main branch network as the shallow feature extraction module of the main branch network;
[0022] S6: Input the low-resolution MRI image data of the target modality obtained in S3 into the shallow feature extraction module of the main branch network to extract the shallow features of the low-resolution MRI image of the target modality;
[0023] S7: Input the shallow features extracted in S6 into the deep feature extraction module of the main branch network to obtain multi-level deep features;
[0024] Among them, stack multiple residual channel attention blocks (RCA) as the backbone, and flexibly embed the cross-scale feature transfer module (Plug-in Mutual-Projection Feature Enhancement, PMFE) between the RCA blocks to form the deep feature extraction module of the main branch network;
[0025] S8: Input the reference modal MRI image features output by S5, the target scale features output by the PMFE module, and the multi-level depth features obtained in S7 into the upsampling and feature fusion module of the main branch network, and adaptively adjust and fuse the features from different branches to obtain fused features;
[0026] S9: Input the fused features obtained in S8 and the target modal low-resolution MRI image obtained in S3 into the image reconstruction module of the main branch network to obtain the reconstructed high-resolution image;
[0027] S10: Set the loss function and perform iterative training on the super-resolution network model based on cross-modal and cross-scale feature fusion;
[0028] S11: Repeat S4 - S10 until the model converges to obtain the trained super-resolution network model based on cross-modal and cross-scale feature fusion.
[0029] Preferably, the output end of the reference image gradient map extraction module of the cross-modal reference branch network is connected to the reference image feature extraction module of the cross-modal reference branch network, and the output end of the reference image feature extraction module of the cross-modal reference branch network is respectively connected to the cross-scale feature transfer module and the upsampling and feature fusion module of the main branch network;
[0030] The output end of the shallow feature extraction module of the main branch network is connected to the deep feature extraction module of the main branch network. The output ends of the reference image feature extraction module of the cross-modal reference branch network, the deep feature extraction module of the main branch network, and the cross-scale feature transfer module are connected to the upsampling and feature fusion module of the main branch network. The output end of the upsampling and feature fusion module of the main branch network is connected to the image reconstruction module of the main branch network, and the output end of the image reconstruction module of the main branch network is connected to the high-resolution image output module of the main branch network.
[0031] Preferably, in step S3, the target modal high-resolution MRI image data is subjected to simulated degradation preprocessing to obtain the target modal low-resolution image data. The specific method is as follows: adopt the simulated degradation method based on the image space and the simulated degradation method based on the frequency domain to simulate the degradation process of the target modal high-resolution MRI image; among them, in the simulated degradation based on the image space, Gaussian blur and bicubic downsampling (Cubic Downsampling) are used to obtain the low-resolution MRI image; in the simulated degradation method based on the frequency domain, the target modal high-resolution MRI image is transformed to the frequency domain through Fourier transform, the edge part of the frequency domain data is truncated according to the super-resolution reconstruction coefficient, and the truncated part is filled with zeros. Subsequently, the filled frequency domain data is subjected to inverse Fourier transform to transform it to the image space, and spatial downsampling is performed on the MRI image to generate the final low-resolution MRI image.
[0032] Preferably, in step S4, the reference modal high-resolution MRI image is input into the reference image gradient map extraction module of the cross-modal reference branch network to obtain the gradient map of the reference modal MRI image. The specific method is as follows: A convolutional layer is used to implement the operation of gradient map extraction. Among them, the expression formula for the gradient extraction operation is as follows:
[0033] G H (I Ref ) = I Ref (H + 1, W, L) - I Ref (H - 1, W, L),
[0034] G ww (I Ref ) = I Ref (H, W + 1, L) - I Ref (H, W - 1, L),
[0035] G L (I Ref ) = I Ref (H, W, L + 1) - I Ref (H, W, L - 1),
[0036]
[0037]
[0038] In the formula, I Ref represents the reference modal high-resolution MRI image, H represents the height, W represents the width, L represents the length, G H (·), G w (·), GL(·) represent the operations of extracting gradients in the corresponding directions; represents the operation of obtaining gradient information with gradient intensity and gradient direction, GI(·) represents the operation of extracting a gradient map containing only gradient intensity information, ||·|| 2 represents the operation of taking the square root of the sum of the squares of the gradient intensities.
[0039] Preferably, in step S5, the specific method of constructing the feature extraction module by combining the 3D convolutional layer and the activation function is as follows: The leaky rectified linear unit (LReLU) is selected as the activation function. Among them, different from ReLU which maps negative inputs to 0, LReLU multiplies negative inputs by a very small weight, and the value range of the weight is 0.001 - 0.01, so that it outputs a very small negative number, and prevents the problem of neuron inactivation caused by all negative outputs being 0.
[0040] Preferably, in step S6, the target-modal low-resolution MRI image data obtained in S3 is input into the shallow feature extraction module of the main branch network to extract the shallow features of the target-modal low-resolution MRI image. The specific method is as follows: The activation function is selected as LReLU, and the shallow feature extraction process is expressed as:
[0041] X 0 =F Conv (I LR ),
[0042] wherein, X 0 represents the shallow features of the target-modal low-resolution MRI image, I LR is the target-modal low-resolution image, and F Conv (·) represents the operation of extracting the shallow features of the target-modal low-resolution MRI image.
[0043] Preferably, in step S7, the shallow features extracted in S6 are input into the deep feature extraction module of the main branch network to obtain multi-level deep features. The specific method is as follows: Stack M (10 ≤ M ≤ 15) RCA blocks as the backbone, and the PMEF module is flexibly embedded between the RCA blocks to process and obtain the target-scale features and corresponding low-resolution scale features of the main branch network with high-resolution details. The features output by the M RCA modules and the low-resolution scale features output by the PMFE module are fused to obtain multi-level deep features;
[0044] The specific method of embedding the PMEF module between the mth (1 ≤ m ≤ M - 1) and the (m + 1)th RCA blocks to process the features is as follows: The features X m (the subscript m in X m corresponds to the mth RCA block) output by the mth RCA module are input into the PMEF module to explore the global self-similarity prior information of the different-scale features of the MRI image, and the features Y m (the subscript m in Y m corresponds to the mth RCA block) with rich high-resolution details are obtained; Y m and X m are fused to obtain the target-scale features Y' m with high-resolution details of the main branch network output. Y' m and the reference-modal MRI image features I Ref are fused to obtain the low-resolution scale features X' m . Among them, the feature information X m is input into the PMFE module, and downsampling is performed with a stride s (1 ≤ s ≤ 3) to obtain the downsampled features X m , Perform a convolution operation to obtain Q and K (Q stands for Query, representing the query; K stands for Key, representing the vector of the relevance between the information to be queried and other information). Cut Q and K into blocks with a step size of g (1 ≤ g ≤ 3) and a block size of p (1 ≤ p ≤ 3) respectively to obtain q (several blocks cut from Q, and the i-th block is called q i ), k (several blocks cut from K, and the j-th block is called k j ). Calculate the similarity weights between each q and k to explore the global cross-scale dependence relationship between Q and K. The calculation formula is as follows:
[0045]
[0046] In the formula, <·, ·> represents the inner product operation, q i represents the i-th q block, k j represents the j-th k block, w i,j represents the similarity weight between the i-th q block and the j-th k block, and exp(·) represents the exponential function with the natural constant e as the base;
[0047] Perform a convolution operation on X m to obtain V. Cut V into blocks with a step size of s×g and a block size of q (1 ≤ q ≤ 3) to obtain v (several blocks cut from V, and the j-th block is called v j ). Perform a convolution operation with the similarity weights. The calculation formula is as follows:
[0048]
[0049] In the formula, v j represents the j-th v block, w i,j represents the similarity weight between the i-th q block and the j-th k block, v′ i represents the high-resolution patch obtained after the i-th attention operation, represents the element-wise multiplication operation;
[0050] All the high-resolution patches obtained after the attention operation are fused to obtain the feature Y m with rich high-resolution details; Y m and X m are fused to obtain the output target scale feature Y′ m of the main branch network with high-resolution details. Y′ m and the reference modality MRI image feature I Ref are fused to obtain the low-resolution scale feature X′ m . The specific formula is as follows:
[0051] Y′ m = F up (X m ) + Y m,
[0052] X′ m = F down (Y′ m + X Ref ),
[0053] [X′ m , Y′ m = F PMFE ([X m , I Ref ),
[0054] In the formula, m represents the m-th RCA module, X m represents the output feature of the m-th RCA module, Y m represents the feature with rich high-resolution details, Y′ m and X′ m represent the output target-scale feature and the corresponding low-resolution scale feature of the main branch network with high-resolution details, F up (·) is a transposed convolution upsampling operation with a stride of s (1 ≤ s ≤ 3), F down (·) is a convolution downsampling operation with a stride of s (1 ≤ s ≤ 3), F PMFE (·) is the PMFE module function;
[0055] The specific method for connecting the output features of the M RCA modules and the low-resolution scale features output by the PMFE module to obtain multi-level depth features is as follows: Using the low-resolution scale feature X′ m as the input of the (m + 1)-th RCA block to obtain X m+1 , and then concatenating the output features of the M RCA modules and X′ m along the channel direction to obtain multi-level depth feature X c (the subscript c represents Concatenation).
[0056] Preferably, in step S8, the specific method for inputting the reference modality MRI image features output by S5, the target-scale features of the cross-scale feature transfer module, and the multi-level depth features obtained in S7 into the main branch network upsampling and feature fusion module to adaptively adjust and fuse the features from different branches to obtain the fused features is as follows: Fusing the reference modality MRI image feature information I Ref , the output Y′ of the PMFE module m and the multi-level depth feature X c to obtain the preliminary fused feature Y a (the subscript a represents addition); where, I Ref and Y′ mhas the same scale as the target scale, and the multi-level depth feature X with the same scale as the target modality low-resolution MRI image c is upsampled using a 3D sub-pixel convolution layer to obtain the upsampled feature at the target scale; feature Y a obtains feature Y through spatial attention (SA) SA (the subscript SA represents spatial attention); feature Y is obtained through channel attention (CA) CA (the subscript CA represents channel attention); along the channel direction, feature Y SA and feature Y CA are concatenated and then, after a convolution operation, the fused feature Y f (the subscript f represents fusion) is obtained. The calculation formula is as follows:
[0057] Y f = F NConv ([Y CA , Y SA ),
[0058] In the formula, Y f is the fused feature, [Y CA , Y SA is the process of concatenating feature Y SA and feature Y cA , and F Nconv (·) is a convolution operation without an activation function.
[0059] Preferably, in step S9, the fused feature obtained in S8 and the target modality low-resolution MRI image obtained in S3 are input into the image reconstruction module of the main branch network to obtain the reconstructed high-resolution image. The specific method is as follows: The target modality low-resolution image is upsampled to obtain the upsampled feature of the target modality low-resolution image (the subscript LR represents the target modality low-resolution image) and is fused with the fused feature Y f to obtain the reconstructed high-resolution image. The calculation formula is as follows:
[0060]
[0061] In the formula, SR represents super-resolution, LR represents low-resolution, I SR represents the reconstructed high-resolution image, Y f represents the fused feature, represents the upsampled feature of the target modality low-resolution image.
[0062] Preferably, in step S10, the loss function is set, and the specific method for iteratively training the cross-modal and cross-scale feature fusion super-resolution network model is as follows: The super-resolution network model is iteratively trained, and iterated repeatedly until the network model converges. The mean absolute error (MAE) with a regularization term is selected as the loss function, and its expression formula is as follows:
[0063]
[0064] In the formula, F(·) represents the mapping function between I LR and I HR , LR is the low resolution; HR represents the high resolution, and N represents the amount of training data.
[0065] The 3D magnetic resonance super-resolution method provided by the present invention has the following characteristics: The present invention designs a 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion. Specifically, a network with multiple residual channel attention blocks as the backbone is constructed; the gradient information of the high-resolution reference image is used as the reference branch input to provide high-frequency information for the backbone network; a cross-scale feature transfer module is designed and flexibly embedded in the backbone network to capture the global cross-scale self-similarity within the features; all feature maps are fused using channel and spatial attention to adaptively adjust the high-resolution features, and the cross-modal self-similarity prior information of the high-resolution MRI image in the reference modality and the global cross-scale self-similarity prior information of the low-resolution MRI image in the target modality can be fully utilized.
[0066] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion. By adaptively adjusting and fusing the cross-modal self-similarity prior information of the high-resolution MRI image in the reference modality and the global cross-scale self-similarity prior information of the low-resolution MRI image in the target modality, rich high-resolution details are provided for reconstructing high-resolution images. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a schematic flow chart of the method described in the present invention;
[0068] Figure 2 is a structural diagram of the super-resolution network model constructed by the method described in the present invention;
[0069] Figure 3 is a structural diagram of the cross-scale feature transfer module of the main branch;
[0070] Figure 4 is a structural diagram of the upsampling and feature fusion module of the main branch network;
[0071] Figure 5Qualitative analysis results of the method of the present invention and other solutions. Detailed implementation manners
[0072] The accompanying drawings are only for illustrative purposes and should not be construed as a limitation to this patent.
[0073] To better illustrate this embodiment, some components in the accompanying drawings are omitted, enlarged or reduced, which does not represent the actual situation.
[0074] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0075] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0076] Specifically, the flowchart of the 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion is as shown in the appendix Figure 1 shown, and the structural diagram of the cross-modal and cross-scale feature fusion network model is as shown in the appendix Figure 2 shown. The following embodiments are experimented with PyTorch on an NVIDIA 3090 GPU, and the reconstruction scale is set to 2.
[0077] Embodiment 1
[0078] S1: Construct a super-resolution network model based on cross-modal and cross-scale feature fusion;
[0079] Specifically, as shown in the appendix Figure 2 shown, the super-resolution network model based on cross-modal and cross-scale feature fusion includes a cross-modal reference branch network and a main branch network.
[0080] The cross-modal reference branch network includes a cross-modal reference branch network reference image gradient map extraction module and a cross-modal reference branch network reference image feature extraction module; the output end of the cross-modal reference branch network reference image gradient map extraction module is connected to the cross-modal reference branch network reference image feature extraction module, and the output end of the cross-modal reference branch network reference image feature extraction module is respectively connected to the cross-scale feature migration module and the upsampling and feature fusion module of the main branch network.
[0081] The main branch network includes a main branch network shallow feature extraction module, a main branch network deep feature extraction module, a main branch network upsampling and feature fusion module, a main branch network image reconstruction module, and a main branch network high-resolution image output module. The output end of the main branch network shallow feature extraction module is connected to the main branch network deep feature extraction module. The output ends of the cross-modal reference branch network reference image feature extraction module, the main branch network deep feature extraction module, and the cross-scale feature migration module are connected to the main branch network upsampling and feature fusion module. The output end of the main branch network upsampling and feature fusion module is connected to the main branch network image reconstruction module. The output end of the main branch network image reconstruction module is connected to the main branch network high-resolution image output module;
[0082] Among them, multiple RCA blocks are stacked as the backbone, and the cross-scale feature migration module is flexibly embedded between the RCA blocks to form the main branch network deep feature extraction module;
[0083] S2: Obtain high-resolution MRI images of the target modality and the reference modality from the public dataset;
[0084] Specifically, the present invention selects the Kirby21 dataset (KKI06-KKI42) provided by the F.M. Kirby Functional Brain Imaging Center of the Kennedy Krieger Institute as the MRI image data. Among them, the T1W is used as the low-resolution image of the target modality, and the T2W and FLAIR are used as the high-resolution reference images of the reference modality;
[0085] S3: Perform simulated degradation preprocessing on the high-resolution MRI image data of the target modality to obtain the low-resolution MRI image data of the target modality;
[0086] Specifically, a simulated degradation method based on the image space is used to simulate the degradation process of the high-resolution MRI image of the target modality; among them, Gaussian blur and bicubic downsampling (CubicDownsampling) are used in the simulated degradation based on the image space to obtain the low-resolution MRI image;
[0087] More specifically, for the training dataset, low-resolution blocks of 26×26×26 are cropped from the low-resolution MRI image with a stride of 13. The T1W low-resolution image is interpolated to the same size as the high-resolution image using the Cubic method, and the 'imregister' function of MATLAB2017 is used to register the T1W low-resolution image with the reference modality MRI image.
[0088] S4: Input the high-resolution MRI image of the reference modality into the cross-modal reference branch network reference image gradient map extraction module to obtain the gradient map of the reference modality MRI image;
[0089] Construct a feature extraction module by combining a 3D convolutional layer and an activation function;
[0090] Add the feature extraction module to the cross-modal reference branch network as the feature extraction module of the cross-modal reference branch network;
[0091] Specifically, select the leaky rectified linear unit LReLU as the activation function. Among them, different from ReLU which maps negative inputs to 0, LReLU multiplies negative inputs by a very small weight (0.001 - 0.01) to make its output a very small negative number, thus preventing the problem of neuron inactivation caused by all negative outputs being 0.
[0092] More specifically, obtain the gradients in the three directions of height, width, and length, and obtain the gradient map of the reference modality MRI image through the operation of taking the square root of the sum of the squares of the gradients. Use a convolutional layer with 48 channels and a convolutional kernel size of 3×3×3 to obtain the features of the reference modality MRI image. The specific formula is as follows:
[0093] G H (I Ref )=I Ref (H + 1, W, L)-I Ref (H - 1, W, L),
[0094] G ww (I Ref )=I Ref (H, W + 1, L)-I Ref (H, W - 1, L),
[0095] G L (I Ref )=I Ref (H, W, L + 1)-I Ref (H, W, L - 1),
[0096]
[0097]
[0098] In the formula, I Ref represents the reference modality high-resolution MRI image, H represents height, W represents width, L represents length, G H (·), G w (·), GL(·) represent the operations of extracting gradients in the corresponding directions; represents the operation of obtaining gradient information with gradient intensity and gradient direction, GI(·) represents the operation of extracting a gradient map containing only gradient intensity information, ||·|| 2 represents the operation of taking the square root of the sum of the squares of the gradient intensity.
[0099] S5: Input the gradient map obtained in S4 into the reference branch network feature extraction module to capture the structural dependencies and spatial relationships of the reference high-resolution image, and output the reference modality MRI image features;
[0100] S6: Input the target modality low-resolution MRI image data obtained in S3 into the shallow feature extraction module of the main branch network to extract the shallow features of the target modality low-resolution MRI image;
[0101] Specifically, a feature extraction module is constructed by combining a 3D convolutional layer and an LReLU activation function, and the shallow feature extraction process is expressed as:
[0102] X 0 = F Conv (I LR ),
[0103] wherein, X 0 represents the shallow features of the target modality low-resolution MRI image, I LR is the target modality low-resolution image, and F Conv (·) represents a convolutional layer with 48 channels and a convolutional kernel size of 3×3×3 to implement the convolutional operation.
[0104] S7: Input the shallow features extracted in S6 into the deep feature extraction module of the main branch network to obtain multi-level deep features;
[0105] Specifically, stack 10 RCA blocks as the backbone, and flexibly embed the PMEF module between the 5th and 6th RCA blocks to process and obtain the target scale features and corresponding low-resolution scale features of the main branch network with high-resolution details. Fuse the output features of the 10 RCA modules and the low-resolution scale features output by the PMFE module to obtain multi-level deep features;
[0106] More specifically, input the output features X 5 of the 5th RCA module into the PMEF module, perform downsampling with a stride of 2 to obtain the downsampled features Input X 5 , into the convolutional operation to obtain Q and K (Q is short for Query, representing the query; K is short for Key, representing the vector of the correlation between the information to be queried and other information). Cut Q and K into chunks with a stride of 2 and a block size of 3 to obtain q (several chunks cut from Q, and the i-th chunk is called q i ), k (several chunks cut from K, and the j-th chunk is called k j ). Calculate the similarity weights between each q and k to explore the global cross-scale dependence relationship between Q and K. The calculation formula is as follows:
[0107]
[0108] wherein, <·,·> represents an inner product operation, q i represents the i-th q block, k j represents the j-th k block, w i,j represents the similarity weight between the i-th q block and the j-th k block, and exp(·) represents the exponential function with the natural constant e as the base;
[0109] Performing a convolution operation on X 5 to obtain V, which is expanded into patches with a stride of 4, and sliced with a stride of 4 and a block size of 3 to obtain v (several blocks cut from V, and the j-th block is called v j ), and performing a convolution operation with the similarity weight, and the calculation formula is as follows:
[0110]
[0111] wherein, v j represents the j-th V block, w i,j represents the similarity weight between the i-th q block and the j-th k block, v′ i represents the high-resolution patch obtained after the i-th through the attention operation, represents the element-wise multiplication operation;
[0112] All the high-resolution patches obtained after the attention operation are fused to obtain the feature Y with rich high-resolution details 5 ; Fusing Y 5 and X 5 to obtain the output target-scale feature Y′ of the main branch network with high-resolution details 5 , fusing Y′ 5 and the reference modality MRI image feature I Ref to obtain the low-resolution scale feature X′ 5 , and the specific formula is as follows:
[0113] Y′ 5 = F up (X 5 ) + Y 5 ,
[0114] X′ 5 = F down (Y′ 5 + X Ref ),
[0115] [X′ 5 , Yq′ 5 = F PMFE ([X 5 , I Ref ),
[0116] In the formula, 5 represents the 5th RCA module, X 5 represents the output feature of the 5th RCA module, Y 5 represents the feature with rich high-resolution details, Y' 5 and X' 5 represent the output target-scale features and corresponding low-resolution scale features of the main branch network with high-resolution details, F up (·) is a transposed convolution upsampling operation with a stride of 2, F down (·) is a convolution downsampling operation with a stride of 2, F PMFE (·) is the PMFE module function;
[0117] More specifically, the low-resolution scale feature X' 5 output by the PMFE module is used as the input of the 6th RCA block to obtain X 6 , and then the output features of 10 RCA modules and X' 5 are concatenated along the channel direction to obtain the multi-level depth feature X c .
[0118] S8: Input the reference modality MRI image features output by S5, the target-scale features output by the PMFE module, and the multi-level depth features obtained in S7 into the upsampling and feature fusion module of the main branch network, and adaptively adjust and fuse the features from different branches to obtain the fusion features;
[0119] Specifically, the reference modality MRI image feature information I Ref , the output Y' 5 of the PMFE module, and the multi-level depth feature X c are fused to obtain the preliminary fusion feature Y a (the subscript a represents addition); among them, the scales of I Ref and Y' m are the same as the target scale. The multi-level depth feature X c with the same scale as the target modality low-resolution MRI image is upsampled to the target scale using a 3D sub-pixel convolution layer with 48 channels and a convolution kernel size of 3×3×3; the feature Y a obtains the feature Y SA (the subscript SA represents spatial attention) through spatial attention (SA); the feature obtains the feature Y CA (the subscript CA represents channel attention) through channel attention (CA); the feature Y SA and the feature YCA After splicing and convolution operation, the fused feature Y is obtained. f (The subscript f represents fusion), and the calculation formula is as follows:
[0120] Y f = F Conv ([Y CA , Y SA ),
[0121] In the formula, Y f is the fused feature, [Y CA , Y SA is the process of splicing feature Y SA and feature Y CA , and F Conv (·) is a convolutional layer without an activation function, with 48 channels and a convolution kernel size of 3×3×3;
[0122] S9: Input the fused feature obtained in S8 and the target modality low-resolution MRI image obtained in S3 into the image reconstruction module of the main branch network to obtain the reconstructed high-resolution image;
[0123] Specifically, upsample the target modality low-resolution image to obtain the upsampled feature of the target modality low-resolution image and fuse it with the fused feature Y f to obtain the reconstructed high-resolution image. The calculation formula is as follows:
[0124]
[0125] In the formula, SR represents super-resolution, LR represents low-resolution, I SR represents the reconstructed high-resolution image, Y f represents the fused feature, represents the upsampled feature of the target modality low-resolution image.
[0126] S10: Set the loss function and perform iterative training on the super-resolution network model based on cross-modal and cross-scale feature fusion;
[0127] Specifically, select the mean absolute error (MAE) with a regularization term as the loss function, and its expression formula is as follows:
[0128]
[0129] In the formula, F(·) represents the mapping function between I LR and I HR , θ is a parameter, N is the amount of training data, which is 17982 in this example, and the hyperparameter λ of the regularization term is set to 1e-6.
[0130] More specifically, the batch size is set to 16, the learning rate is set to 10, and it is trained for 100 epochs. The Adam optimizer is used to train the reference super-resolution network.
[0131] S11: Repeat S4 - S10 until the model converges to obtain a trained super-resolution network model based on cross-modal and cross-scale feature fusion;
[0132] S12: Input the low-resolution MRI image into the trained network model based on cross-modal and cross-scale feature fusion to obtain the corresponding super-resolution MRI image.
[0133] In the case of spatial domain degradation, the present invention quantitatively analyzes and compares the performance of the 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion with other methods in terms of PSNR and SSIM. See Table 1 for details.
[0134] Figure 5 This is the qualitative analysis result of the 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion of the present invention and other solutions under spatial domain degradation. It can be seen from the comparison in the figure that the method proposed by the present invention can retain more details of anatomical content, can better restore the glioma part, eliminate the blurred edges (indicated by the red arrow), and produce the visual effect most similar to the ground truth image.
[0135] Table 1
[0136]
[0137] The higher the PSNR and SSIM indices, the better the image quality. It can be seen from the results in Table 1 that the method proposed by the present invention achieves the best results in the dataset.
[0138] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all should be covered by the protection scope of the present invention; without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. 3D Magnetic Resonance Super-Resolution Method Based on Cross-Modal and Cross-Scale Feature Fusion, Characterized in that, The operation is as follows: Input the low-resolution MRI image into the trained super-resolution network model based on cross-modal and cross-scale feature fusion to obtain the corresponding super-resolution MRI image; The construction of the super-resolution network model based on cross-modal and cross-scale feature fusion includes the following steps: S1: Construct a super-resolution network model based on cross-modal and cross-scale feature fusion; The super-resolution network model includes a cross-modal reference branch network and a main branch network; The cross-modal reference branch network includes a cross-modal reference branch network reference image gradient map extraction module and a cross-modal reference branch network reference image feature extraction module; The main branch network includes a main branch network shallow feature extraction module, a main branch network deep feature extraction module, a main branch network upsampling and feature fusion module, a main branch network image reconstruction module, and a main branch network high-resolution image output module; S2: Obtain high-resolution MRI images of the target modality and the reference modality from the public dataset; Among them, in the public dataset, T1W is used as the high-resolution MRI image of the target modality, and T2W and FLAIR are used as the high-resolution reference images of the reference modality; S3: Perform simulated degradation preprocessing on the high-resolution MRI image data of the target modality to obtain low-resolution MRI image data of the target modality; S4: Input the high-resolution MRI image of the reference modality into the cross-modal reference branch network reference image gradient map extraction module to obtain the gradient map of the reference modality MRI image; Construct a feature extraction module by combining a 3D convolutional layer and an activation function; Add the feature extraction module to the cross-modal reference branch network as the cross-modal reference branch network feature extraction module; S5: Input the gradient map obtained in S4 into the reference branch network feature extraction module to capture the structural dependencies and spatial relationships of the reference high-resolution image and output the reference modality MRI image features; Add the feature extraction module to the main branch network as the main branch network shallow feature extraction module; S6: Input the low-resolution MRI image data of the target modality obtained in S3 into the main branch network shallow feature extraction module to extract the shallow features of the low-resolution MRI image of the target modality; S7: Input the shallow features extracted in S6 into the main branch network deep feature extraction module to obtain multi-level deep features; Among them, stack multiple residual channel attention blocks as the backbone, and flexibly embed the cross-scale feature transfer module between the residual channel attention blocks to form the main branch network deep feature extraction module; S8: Input the reference modality MRI image features output in S5, the target scale features output by the cross-scale feature transfer module, and the multi-level deep features obtained in S7 into the main branch network upsampling and feature fusion module to adaptively adjust and fuse the features from different branches to obtain fusion features; S9: Input the fusion features obtained in S8 and the low-resolution MRI image of the target modality obtained in S3 into the main branch network image reconstruction module to obtain the reconstructed high-resolution image; S10: Set a loss function and perform iterative training on the super-resolution network model based on cross-modal and cross-scale feature fusion; S11: Repeat S4 - S10 until the model converges to obtain a trained super-resolution network model based on cross-modal and cross-scale feature fusion.
2. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 1, wherein, the output end of the reference image gradient map extraction module of the cross-modal reference branch network is connected to the reference image feature extraction module of the cross-modal reference branch network, and the output end of the reference image feature extraction module of the cross-modal reference branch network is respectively connected to the cross-scale feature transfer module and the upsampling and feature fusion module of the main branch network; the output end of the shallow feature extraction module of the main branch network is connected to the deep feature extraction module of the main branch network, the output end of the reference image feature extraction module of the cross-modal reference branch network, the output end of the deep feature extraction module of the main branch network, and the output end of the cross-scale feature transfer module are connected to the upsampling and feature fusion module of the main branch network, the output end of the upsampling and feature fusion module of the main branch network is connected to the image reconstruction module of the main branch network, and the output end of the image reconstruction module of the main branch network is connected to the high-resolution image output module of the main branch network.
3. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 1, wherein, Step S3 performs simulated degradation preprocessing on the high-resolution MRI image data of the target modality to obtain the low-resolution MRI image data of the target modality. The specific method is as follows: The degradation process of the high-resolution MRI image of the target modality is simulated by using a simulated degradation method based on the image space and a simulated degradation method based on the frequency domain; among them, in the simulated degradation based on the image space, Gaussian blur and bicubic downsampling are used to obtain the low-resolution MRI image data; in the simulated degradation method based on the frequency domain, the high-resolution MRI image of the target modality is transformed to the frequency domain through Fourier transform, the edge part of the frequency domain data is truncated according to the super-resolution reconstruction coefficient, and the truncated part is filled by zero-padding, and then the filled frequency domain data is subjected to inverse Fourier transform to transform it to the image space, and spatial downsampling is performed on the MRI image to generate the final low-resolution MRI image data.
4. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 1, wherein, Step S4 inputs the high-resolution MRI image of the reference modality into the reference image gradient map extraction module of the cross-modal reference branch network to obtain the gradient map of the reference modality MRI image. The specific method is as follows: A convolutional layer is used to implement the operation of gradient map extraction. Among them, the expression formula of the gradient extraction operation is as follows: G H (I Ref ) = I Ref (H + 1, W, L) - I Ref (H - 1, W, L), G w (I Ref ) = I Ref (H, W + 1, L) - I Ref (H, W - 1, L), G L (I Ref ) = I Ref (H, W, L + 1) - I Ref (H, W, L - 1), Wherein, I Ref represents the reference modal high-resolution MRI image, H represents the height, W represents the width, L represents the length, and G H (·), G W (·), G L (·) represents the operation of extracting the gradient in the corresponding direction; represents the operation of obtaining gradient information with gradient intensity and gradient direction, GI(·) represents the operation of extracting a gradient map containing only gradient intensity information, and ||·|| 2 represents the operation of taking the square root of the sum of the squares of the gradient intensities.
5. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 1, wherein, In step S5, the specific method of constructing the feature extraction module by combining the 3D convolutional layer and the activation function is as follows: The leaky rectified linear unit (LReLU) is selected as the activation function. Among them, different from the fact that ReLU maps negative inputs to 0, LReLU multiplies negative inputs by a weight, and the value range of the weight is 0.001 - 0.01, so that it outputs extremely small negative numbers, thereby preventing the problem of neuron inactivation caused by all negative numbers outputting 0.
6. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 1, wherein, In step S6, the target-modal low-resolution MRI image data obtained in S3 is input into the shallow feature extraction module of the main branch network to extract the shallow features of the target-modal low-resolution MRI image. The specific method is as follows: The activation function LReLU is selected, and the shallow feature extraction process is expressed as: X 0 = F Conv (I LR ), Wherein, X 0 represents the shallow features of the target modality low-resolution MRI image, I LR is the target modality low-resolution MRI image, and F Conv (·) represents the operation of extracting the shallow features of the target modality low-resolution MRI image.
7. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 1, wherein, In step S7, the shallow features extracted in S6 are input into the deep feature extraction module of the main branch network to obtain multi-level deep features. The specific method is as follows: Stack M residual channel attention blocks as the backbone, where 10 ≤ M ≤ 15. The cross-scale feature transfer module is flexibly embedded between the residual channel attention blocks for processing to obtain the output target-scale features and corresponding low-resolution scale features of the main branch network with high-resolution details. The features output by the M residual channel attention blocks and the low-resolution scale features output by the cross-scale feature transfer module are fused to obtain multi-level deep features; The specific method of embedding the cross-scale feature transfer module between the m-th and (m + 1)-th residual channel attention blocks to process the features is as follows: The output feature X of the m-th residual channel attention block m is input into the cross-scale feature transfer module to explore the global self-similarity prior information of the features at different scales of the MRI image, and the feature Y with rich high-resolution details is obtained m ; Y m and X m are fused to obtain the output target-scale feature Y' of the main branch network with high-resolution details m , and Y' m and the feature I of the reference modality MRI image Ref are fused to obtain the low-resolution scale feature X' m , where the feature information X m is input into the cross-scale feature transfer module and downsampled with a stride s to obtain the downsampled feature 1 ≤ s ≤ 3; X m , are subjected to convolution operations to obtain Q and K, and are sliced into q and k with strides g and block sizes p respectively, 1 ≤ g ≤ 3, 1 ≤ p ≤ 3, and the similarity weights between each q and k are calculated to explore the global cross-scale dependence relationship between Q and K. The calculation formula is as follows: wherein, <·, ·> represents an inner product operation, q i represents the i-th q block, k j represents the j-th k block, w i,j represents the similarity weight between the i-th q block and the j-th k block, and exp(·) represents the exponential function with the natural constant e as the base; Convolve X m to obtain V through convolution operation, segment it into v with a stride of s×g and a block size of q, and perform a convolution operation with the similarity weight. The calculation formula is as follows: where \(v\) j represents the \(j\)-th \(v\) block, \(w\) i,j represents the similarity weight between the \(i\)-th \(q\) block and the \(j\)-th \(k\) block, \(v'\) i represents the high-resolution patch obtained after the \(i\)-th attention operation, represents the element-wise multiplication operation; All the high-resolution patches obtained after the attention operation are fused to obtain the feature Y with rich high-resolution details. m , fuse Y m and X m to fuse the output target-scale feature Y' of the main branch network with high-resolution details. m , fuse Y' m and the reference modality MRI image feature I Ref to fuse the low-resolution scale feature X'. m The specific formula is as follows: Y′ m = F up (X m ) + Y m , X′ m = F down (Y′ m + X Ref ) [X′ m , Y′ m = F PMFE ([X m , I Ref ) where m represents the m-th residual channel attention block, and X m represents the output feature of the m-th residual channel attention block, Y m represents the feature with rich high-resolution details, Y' m and X' m represent the output target-scale feature and the corresponding low-resolution-scale feature of the main branch network with high-resolution details, F up (·) is the transposed convolution upsampling operation with a stride of s, F down (·) is the convolution downsampling operation with a stride of s, F PMFE (·) is the cross-scale feature transfer module function; The specific method for connecting the output features of M residual channel attention blocks and the low-resolution scale features output by the cross-scale feature transfer module to obtain multi-level depth features is as follows: The low-resolution scale features X' output by the cross-scale feature transfer module m are used as the input to the (m + 1)-th residual channel attention block to obtain X m+1 . Then, the output features of the M residual channel attention blocks and X' m are concatenated along the channel dimension to obtain the multi-level depth feature X c .
8. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 6, wherein, In step S8, the specific method of inputting the reference modal MRI image features output by S5, the target scale features of the cross-scale feature transfer module, and the multi-level depth features obtained in S7 into the upsampling and feature fusion module of the main branch network to adaptively adjust and fuse the features from different branches to obtain the fused features is as follows: the reference modal MRI image feature information I Ref , the target scale features Y' of the cross-scale feature transfer module m , and the multi-level depth features X c are fused to obtain the preliminary fused feature Y a ; among them, the scales of I Ref and Y' m are the same as the target scale, and the multi-level depth features X c with the same scale as the target modal low-resolution MRI image are upsampled by a 3D sub-pixel convolutional layer to obtain the upsampled features at the target scale; the feature Y a obtains the feature Y SA through spatial attention; the feature obtains the feature Y CA through channel attention; the feature Y SA and the feature Y CA are concatenated along the channel direction and then obtained the fused feature Y f after convolution operation, and the calculation formula is as follows: Y f = F NConv ([Y CA , Y SA ) Where Y f is the fused feature, [Y CA , Y SA is the process of splicing feature Y SA and feature Y CA , and F NConv (·) is a convolution operation without an activation function.
9. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 1, wherein, In step S9, the fused features obtained in S8 and the target-modal low-resolution MRI image obtained in S3 are input into the image reconstruction module of the main branch network to obtain the reconstructed high-resolution image. The specific method is as follows: the target-modal low-resolution image is upsampled to obtain the upsampled features of the target-modal low-resolution image and fused with the fused features Y f to obtain the reconstructed high-resolution image. The calculation formula is as follows: Wherein, SR represents super resolution, LR represents low resolution, and I SR represents the reconstructed high-resolution image, and Y f represents the fused feature, represents the upsampled feature of the target modal low-resolution image.
10. The 3D magnetic resonance super-resolution method based on cross-modal and cross-scale feature fusion according to claim 1, wherein, In step S10, a loss function is set, and the specific method for iteratively training the super-resolution network model based on cross-modal and cross-scale feature fusion is as follows: The super-resolution network model is iteratively trained, and iterated repeatedly until the network model converges. The mean absolute error (MAE) with a regularization term is selected as the loss function, and its expression formula is as follows: Wherein, F(·) represents the mapping function between I LR and I HR , LR is the low resolution, HR represents the high resolution, and N represents the amount of training data.