Transformer-Based Multimodal MRI Reconstruction Method
Through the multimodal MRI reconstruction method based on Transformer, the spatial registration and image reconstruction modules are used to solve the problem of long MRI scanning time and artifacts, and the reconstruction and artifact removal of high-resolution images are realized, improving the efficiency and quality of MRI reconstruction.
Patent Information
- Application Number
- CN202411764978.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-12-04
AI Technical Summary
The MRI scan time is too long and is sensitive to the physiological and physical movement of the person being scanned, resulting in different degrees of artifacts in the image, limiting the promotion and development of MRI.
A multimodal MRI reconstruction method based on Transformer is designed, and a network model is constructed by constructing an initial multimodal MRI image reconstruction, and a low-resolution target modal MRI image is reconstructed into a high-resolution image using the spatial registration module and the image reconstruction module.
Improve the quality of reconstructed image details, ensure high signal-to-noise ratio and fidelity, effectively remove tissue artifacts, and improve the efficiency and quality of MRI reconstruction.
Smart Images

Figure CN119251346B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image reconstruction, and particularly to a multi-modal MRI reconstruction method based on Transformer. Background Art
[0002] Image reconstruction belongs to an innovative direction in computer vision and artificial intelligence, and is mainly applied to multiple fields such as remote sensing imaging, security monitoring, and autonomous driving. Image reconstruction utilizes the prior knowledge of the degradation process and, through computer processing, removes or reduces degradation effects such as blurring and noise in the image, enabling the image to recover or approach the original true state. Image reconstruction technology relies on advanced deep learning models and algorithms, such as convolutional neural networks (CNNs) and Vision Transformers, to extract local and global features and generate high-quality images. Therefore, image reconstruction has become a research hotspot in the fields of computer vision and artificial intelligence, and is an important part of modern digital media and image content restoration systems.
[0003] Nuclear Magnetic Resonance Imaging (MRI) technology provides a more reliable basis for medical and biological research with its characteristics of high resolution and multi-parametric imaging. However, compared with other imaging methods, the MRI scanning time is too long and the imaging is slow, usually taking 15 minutes to 1 hour. In addition, during the MRI scanning process, it is very sensitive to the physiological and body movements of the scanned person. For example, heartbeat, breathing, coughing, and swallowing will all cause different degrees of artifacts in the image, which have restricted the popularization and development of MRI.
[0004] In early research, the reconstruction process of MRI mainly relied on parallel imaging and compressed sensing. Parallel imaging technology uses multiple coils to sample a certain part simultaneously and utilizes sensitivity information to assist in spatial positioning. Compressed sensing recovers the original data from the undersampled k-space by utilizing the sparsity of the image in the total variation (TV) or wavelet transform (WT). Recent research has shown that considering the complementary information between magnetic resonance images with different contrasts, fully sampled auxiliary mode MRI can help reconstruct the undersampled target mode MRI more effectively. In view of this, it is considered that the key to optimizing MRI reconstruction lies in two aspects: reconstructing anatomical details and removing tissue artifacts. Summary of the Invention
[0005] Objective of the present invention: to provide a multi-modal MRI reconstruction method based on Transformer, which can improve the detail quality of the reconstructed image and ensure high signal-to-noise ratio and fidelity;
[0006] To achieve the above functions, the present invention designs a multi-modal MRI reconstruction method based on Transformer, which performs the following steps S1 - S3 to reconstruct a low-resolution target-modal MRI image into a high-resolution image:
[0007] Step S1: Acquire a low-resolution target-modal MRI image and the corresponding high-resolution auxiliary-modal MRI image;
[0008] Step S2: Construct an initial multi-modal MRI image reconstruction network model, form an MRI image pair with the target-modal MRI image and the auxiliary-modal MRI image, and input the MRI image pair into the initial multi-modal MRI image reconstruction network model for training to output a reconstructed image, thereby obtaining a trained multi-modal MRI image reconstruction network model;
[0009] The multi-modal MRI image reconstruction network model includes a spatial registration module and an image reconstruction module. The spatial registration module aligns the spatial positions of the auxiliary-modal MRI image with the target-modal MRI image, and the image reconstruction module extracts auxiliary information from the MRI image pair;
[0010] Step S3: Input the paired undersampled and fully sampled MRI image pairs into the trained multi-modal MRI image reconstruction network model, and output the high-resolution image of the reconstructed target modality, thus completing the reconstruction of the target-modal MRI image.
[0011] Beneficial effects: Compared with the prior art, the advantages of the present invention include:
[0012] The present invention designs a multi-modal MRI image reconstruction network model, which is cascaded and predicts the flow field from coarse to fine, and can obtain accurate registration results;
[0013] The present invention improves the CSWin Transformer (alternating cross-window Transformer), uses alternating horizontal and vertical window attention instead of cross-window attention, and avoids window direction order bias;
[0014] The present invention adopts many different mask types, including Cartesian, Gaussian, and Poisson, as well as different undersampling rates, and conducts extensive experiments and result visualizations to verify the robustness of the model designed by the present invention under different conditions. Description of the Drawings
[0015] Figure 1It is a model framework diagram of a multi-modal MRI image reconstruction network model provided according to an embodiment of the present invention;
[0016] Figure 2 It is a schematic structural diagram of a spatial registration module provided according to an embodiment of the present invention;
[0017] Figure 3 It is a schematic diagram of a spatial transformation provided according to an embodiment of the present invention;
[0018] Figure 4 It is a schematic structural diagram of an image reconstruction module provided according to an embodiment of the present invention;
[0019] Figure 5 It is a schematic structural diagram of SCRWin Transformer Unet provided according to an embodiment of the present invention;
[0020] Figure 6 It is a schematic diagram of CSWin Transformer provided according to an embodiment of the present invention;
[0021] Figure 7 It is a graph of the results of a qualitative comparison experiment using a 50% Cartesian sampling pattern provided according to an embodiment of the present invention;
[0022] Figure 8 It is a graph of the results of a qualitative comparison experiment using a 50% Poisson sampling pattern provided according to an embodiment of the present invention;
[0023] Figure 9 It is a graph of the results of an ablation experiment of the spatial registration module provided according to an embodiment of the present invention;
[0024] Figure 10 It is a graph of the results of an ablation experiment of CSWin Transformer provided according to an embodiment of the present invention. Detailed implementation manners
[0025] The following further describes the present invention with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0026] The multi-modal MRI reconstruction method based on Transformer provided by the embodiment of the present invention performs the following steps S1 - S3 to reconstruct a low-resolution target-modal MRI image into a high-resolution image:
[0027] Step S1: Acquire a low-resolution target-modal MRI image and a corresponding high-resolution auxiliary-modal MRI image;
[0028] Step S2: Construct an initial multi-modal MRI image reconstruction network model. Form an MRI image pair with the target-modal MRI image and the auxiliary-modal MRI image, and input the MRI image pair into the initial multi-modal MRI image reconstruction network model for training to output a reconstructed image, thereby obtaining a trained multi-modal MRI image reconstruction network model;
[0029] The multi-modal MRI image reconstruction network model includes a spatial registration module and an image reconstruction module. The spatial registration module aligns the spatial positions of the auxiliary-modal MRI image to the target-modal MRI image, and the image reconstruction module extracts auxiliary information from the MRI image pair;
[0030] As Figure 1 shown, the training of the multi-modal MRI reconstruction network model is further introduced. Input the undersampled target-modal MRI image and the auxiliary-modal MRI image into the model. After registration, output the auxiliary-modal MRI image and the undersampled target-modal MRI image to the image reconstruction module. Through the learning and processing of the image reconstruction module, the final reconstructed image is obtained. Specifically, first input the auxiliary-modal MRI image with clearer relevant information and the low-resolution target-modal MRI image into the spatial registration module in pairs to achieve spatial feature alignment between the auxiliary-modal MRI image and the target-modal MRI image, so as to effectively extract the true pixel information of the target-modal MRI image from the auxiliary-modal MRI image.
[0031] After registration, the auxiliary-modal MRI image is fused with the target-modal MRI image to obtain initial auxiliary features. Due to the inaccuracy of this function, subsequent ASCRWin Transformer is required for further extraction. At the same time, to avoid the interference of the auxiliary modality, only the target-modal MRI image is input into the USCRWin Transformer alone to extract unique features. Add the auxiliary feature U and the unique feature A to obtain a residual image, and then the residual image is supplemented and added to the original target-modal MRI image to generate a reconstructed image. Regard the reconstructed image generated in the previous stage as the original target-modal MRI image, update the auxiliary and unique features in the same way as the initialization, generate the residual image of this stage, and add it to the original target-modal MRI image to obtain the reconstructed image of this stage.
[0032] The specific steps of Step S2 are as follows:
[0033] Step S2.1: The structural diagram of the spatial registration module is as Figure 2 shown. The architecture of this module aims to predict the flow field through the input image pair, perform spatial transformation on the auxiliary-modal MRI image, eliminate the spatial position differences between the auxiliary-modal MRI image and the target-modal MRI image, and achieve the alignment of the same tissue structures.
[0034] In the spatial registration module, after performing flow field prediction on the auxiliary modality MRI image, a spatial transformation (SpatialTransformer, STN) is implemented to obtain a registered image;
[0035] The specific steps of step S2.1 are as follows:
[0036] Step S2.1.1: The target modality MRI image and the auxiliary modality MRI image are concatenated in channels to obtain a combined image, and the combined image is input into the VoxelMorph network to obtain a predicted flow field;
[0037] Specifically, the undersampled target modality MRI image and the auxiliary modality MRI image are input into the model. After registration, the auxiliary modality MRI image and the undersampled target modality MRI image are input into the reconstruction module. After learning and processing by the reconstruction module, the final reconstructed image is obtained. Figure 2 In it, flow is the deformation field predicted and generated by the VoxelMorph network, and warp is the result of registering the moving image with the flow field using the spatial transformer. The VoxelMorph network consists of a U-Net network. The encoder in it uses max-pooling for downsampling. The combined image is input into the encoder of the U-Net network for three times of max-pooling downsampling. Each time the downsampling reduces the size of the feature map to half of the original, and the number of channels increases to 32, 64, and 128. The obtained high-level feature map is input into the decoder after passing through the bottleneck layer. The decoder uses transposed convolution to upsample the high-level features. Each time the upsampling doubles the size of the feature map and reduces the number of channels to half of the original. Skip connections are made between the encoder and the decoder to fuse high-level features and low-level features. Finally, a 2-channel predicted flow field is output.
[0038] Step S2.1.2: Use the predicted flow field to perform a spatial transformation on the auxiliary modality MRI image and correct it using the spatial transformation network to obtain the registered image for this time;
[0039] Refer to Figure 3 For, the specific steps of step S2.1.2 are as follows:
[0040] Step S2.1.2.1: The localization network predicts the affine transformation matrix through the image features extracted by the CNN network θ ;
[0041] Step S2.1.2.2: The grid generator generates the position transformation mapping relationship before and after the transformation according to the affine transformation matrix θ ; Figure 3 In represents the position transformation mapping relationship;
[0042] Step S2.1.2.3: Perform pixel value sampling according to the position transformation mapping relationship and perform bilinear interpolation to solve the problem of decimal positions in the grid sampler, as shown in the following formula:
[0043] ;
[0044] ;
[0045] In the formula, is the pixel value at the position of the output image , ; and are the height and width of the output image respectively, C represents the channel dimension; is the pixel value at the position of the combined image (n, m, c), , H and W are the height and width of the combined image respectively, θ 11 ~ θ 23 represent the parameters of the affine transformation, k represents bilinear interpolation, and represent the parameters of the interpolation function, represents the coordinates of the original image.
[0046] Step S2.1.3: Use the previous registered image as the auxiliary modal MRI image, repeat Steps S2.1.1 - S2.1.2 to generate a coarse-to-fine flow field, and obtain a registered image without position deviation.
[0047] Step S2.2: Input the registered image and the target modal MRI image into the initial multi-modal MRI image reconstruction network model. After passing through the image reconstruction module, output the reconstructed image to obtain a trained multi-modal MRI image reconstruction network model.
[0048] As Figure 4 shown, the reconstruction module receives the original low-resolution target modal MRI image, the high-resolution auxiliary modal MRI image, and the fused features as all inputs of the reconstruction module. The unique features and the auxiliary features are added to obtain the reconstructed image of this iteration. The reconstructed image is used as the new target modal MRI image to generate the fused features of the next iteration, and continue to repeat the iteration. Figure 4In the USCRWin Transformer Unet, it is used to output unique features, and in the ASCRWin Transformer Unet, it is used to output auxiliary features. The architecture diagram of the SCRWin Transformer (Separable Cross-shaped Window Transformer) Unet used to output unique features and auxiliary features in the reconstruction module is as Figure 5 shown. Residual connections are used to retain the context features of the encoder and decoder to accelerate the convergence speed during training. The encoder consists of a patch embedding layer, an SCRWin Transformer, and a patch merging layer. The patch embedding layer uses a 7x7 convolutional kernel, which neutralizes between the 16x16 convolutional kernel used in vision transformers and the 3x3 convolutional kernel used in EarlyConv to improve the model's non-linear expression ability while avoiding low receptive fields. Through patch embedding, the input image is divided into several non-overlapping blocks of size 4x4, and an additional linear projection layer can be used to project the feature dimensions of the feature blocks to any dimension. There is a patch merging layer between every two SCRWin Transformer layers, which is used to fuse different patches in the feature map and plays a role in downsampling and expanding the feature dimensions to construct feature maps of different scales. During the encoding stage, a total of three downsamplings are performed on the embedded sequence. For each scale of the feature map, it is input into the corresponding feature map generated by the decoder using residual connections. The decoder consists of interleaved SCRWin Transformers and patch expansion layers. The patch expansion layer upsamples the feature map while reducing the feature dimensions, gradually restoring the original resolution of the image.
[0049] The specific steps of step S2.2 are as follows:
[0050] Step S2.2.1: Input the registered image and the target modality MRI image into the CNN network to obtain fused features;
[0051] Step S2.2.2: Input the fused features and the target modality MRI image into a U-Net network with Transformers as basic components with non-shared parameters, namely the SCRWin Transformer Unet, for the prediction of auxiliary features and unique features;
[0052] The CSWin Transformer (Cross-shaped Window Transformer) proposes a cross-shaped window multi-head self-attention mechanism, as Figure 6As shown in the figure. The attention heads are divided into two parts: one part calculates the local attention of the vertical sliding window, and the other part calculates the local attention of the horizontal sliding window. The window width is adjusted according to the resolution of the feature map. For high-resolution feature maps, the window width is set smaller to consider image details, while for low-resolution feature maps, the window width is set larger to fuse more information. The present invention improves the calculation method of the CSWin Transformer to obtain the SCRWin Transformer (Separated Cross-shaped Window Transformer). In the original model, multiple cross-window and multi-head self-attention calculations are performed in each stage. The method of calculating attention by dividing the attention heads only considers the unidirectional information in two feature maps. To solve this problem, the present invention calculates the sliding window attention in an alternating manner of "first horizontal then vertical, then vertical then horizontal" in each CSWin Transformer layer. Specifically, if there are 24 attentions at this time, they are divided into a group according to the first 12 attentions, and the last 12 attentions are used as the second group. In the CSWin Transformer, each group of feature maps only performs attention on a single-direction sliding window in all stages, while the SCRWin Transformer performs direction-alternating sliding window attention calculations. The alternating order of the two groups of feature map windows is opposite to minimize the deviation of the model from the fixed-direction window and the fixed order. Each attention head fully considers the horizontal and vertical information, and compared with calculating attention only in a unidirectional window, it can generate a more accurate weight map. In addition, based on the input image size of 256x256, the window widths are set to 1, 2, and 8 respectively.
[0053] The formula of the SCRWin Transformer is as follows:
[0054] ;
[0055] ;
[0056] ;
[0057] ;
[0058] Wherein, d k is the dimension of the column vectors in the matrices Q and K , Q is the query matrix, V is the value matrix, K is the content of attention, head i represents the i th head attention, and is the weight acting on Q matrix and K matrix, and the subscript i represents that the calculation is the i head attention, W O is the output weight characterized by the convolution kernel, QK T is a dot product operation used to calculate Q the attention weight on V By scaling with softmax, it aims to avoid an overly large dot product. Because when the dot product is too large, the gradient through softmax will be very small. In addition, the advantage of softmax is that it facilitates the calculation of the gradient for backpropagation and smooths the result to the range of 0 - 1. 1…c / 2 is the first half of the channels, and c / 2 + 1…c is the second half of the channels. stage If it is odd, calculate the longitudinal local window attention; otherwise, calculate the lateral window attention.
[0059] Step S2.2.3: Input the auxiliary feature and the unique feature into the CNN network to obtain the residual feature;
[0060] Step S2.2.4: Add the residual feature and the target modality MRI image to obtain the current reconstructed image;
[0061] Step S2.2.5: Use the current reconstructed image as the target modality MRI image, input the target modality MRI image and the registered image into the CNN network to obtain the next fusion feature, and iterate steps S2.2.1 - S2.2.4; until the residual feature can recover the details lost in the undersampled image and remove the redundant artifacts, stop the iteration to obtain the trained multi-modal MRI image reconstruction network model.
[0062] The iteration process of step S2.2.5 is as follows:
[0063] ;
[0064] ;
[0065] ;
[0066] ;
[0067] ;
[0068] where U is the auxiliary feature and A is the unique feature, ED A , ED UThey are encoders and decoders for extracting auxiliary features and unique features respectively. The superscript t is the number of iterations, diff is the residual map, Rec x is the reconstructed image, Rec A is the fused feature, is the fully sampled auxiliary modality MRI image.
[0069] Step S3: Input the paired undersampled and fully sampled MRI image pairs into the trained multi-modal MRI image reconstruction network model, and output the high-resolution image after reconstruction of the target modality, thus completing the reconstruction of the target modality MRI image.
[0070] Figures 7 - 8 shows the qualitative experimental results of reconstruction using the IXI brain dataset on different Transformer-based MRI reconstruction algorithms.
[0071] The first group of experiments: Referring to Figure 7 , adopting a 50% Cartesian sampling pattern, the reconstructed brain nuclear magnetic resonance images of 5 baseline models including the method of the present invention are all relatively clear, and the obvious artifacts in the initial state are basically removed. Some residual vertical line ghosts due to undersampling can still be seen in MTrans. The details of the Task-Transformer method are jumbled and it looks relatively blurred. The edges of the SwinRN method are not sharp enough. The edges of the internal tissues of the DuDReTLU-net method are not clear enough. The images restored by the method of the present invention have clear edges and texture details, and detailed tissue structures can be seen, while aliasing and artifacts are basically completely eliminated.
[0072] The second group of experiments: Referring to Figure 8 , adopting a 50% Poisson sampling pattern, artifacts can still be seen in the empty areas of the images reconstructed by the MTrans and Task-Transformer methods, and both have jagged edges. The internal tissues of the SwinRN method are darker. The DSFormer performs poorly in areas where the tissue structure is relatively dense. The edges of some external contours and internal tissue structures of the DuDReTLU-net are blurred. The method of the present invention reconstructs images that basically do not have the above problems.
[0073] When the sampling rates are the same, the aliasing phenomenon of the images generated by Poisson sampling is not as serious as that of Cartesian sampling. The reason is that k important imaging information is contained in the spatial center. If k more information in the spatial center is lost, then the collected pictures will also have more ghosting.
[0074] To quantitatively compare the results, the present invention uses PSNR, SSIM, and RMSE as metrics to measure the quality of the reconstructed images. Different Transformer-based MRI reconstruction models were tested using the IXI brain dataset, and all models were tested under the same conditions. The sampling masks were Cartesian and Poisson patterns, and the sampling rate was set to 50%, with an acceleration factor AF = 2. The experimental results are shown in Table 1:
[0075]
[0076] The PSNR, SSIM, and RMSE obtained by the method of the present invention all reach the optimal values. PSNR is the peak signal-to-noise ratio. A high PSNR indicates that the difference between the reconstructed image and the original image is very small, the image details and color restoration are relatively good, the degree of distortion is low, and the visual effect is close to the original image. SSIM takes into account the changes in three aspects: brightness, contrast, and structure. A high SSIM indicates that the structural feature similarity between the reconstructed image and the original image is close. Even if there are slight color or texture differences, it is difficult for the human eye to perceive obvious distortion. RMSE is the root mean square difference between the actual value and the predicted value, and a lower value indicates a smaller gap between the predicted value and the true value.
[0077] To fully illustrate the effectiveness of the method of the present invention, the present invention conducted ablation experiments on the spatial registration module and CSWin Transformer (alternating cross-window Transformer). Specifically, in order to study the impact of image registration on the reconstruction algorithm, the present invention conducted experiments by deleting the spatial registration module while keeping the rest of the original model settings unchanged. Compared with the original model and other models with the spatial registration module retained, the performance of the multi-modal MRI image reconstruction model without the spatial registration module is the worst. Without the spatial registration module, the MRI images lack spatial alignment, which makes it difficult for the image reconstruction module to extract effective complementary features. This not only increases the learning burden of the image reconstruction module but also results in poor reconstruction quality. The experimental results are as Figure 9 shown.
[0078] Similarly, using the Cartesian and Poisson sampling patterns, the edges of the reconstructed images of the model without introducing the spatial registration module are relatively blurred compared to the model with the spatial registration module, the anatomical tissues are not fine enough, and the lines in the complex parts are slightly distorted. The reconstructed image of the model with the spatial registration module is almost identical to the real image. From the residual map between the two, it can be seen that the reconstructed image perfectly restores the high-resolution original image.
[0079] The present invention also conducted ablation experiments on CSWin Transformer, and the experimental results are as Figure 10As shown. To verify the effectiveness of the proposed SCRWin Transformer of the present invention, CSWin Transformer was used to replace SCRWin Transformer, while the remaining settings of the original model were retained for experiments. The results show that the proposed SCRWin Transformer of the present invention can indeed improve the performance of the model to a certain extent. Compared with the model based on CSWin Transformer, the PSNR and SSIM of the IXI brain dataset are increased by 1.31% and 1.02% respectively. SCRWin Transformer can achieve more stable loss reduction, which helps the model to converge quickly and optimize its performance, improve the generalization ability, and ensure reliable model evaluation, so that the model can learn data features more effectively and perform well on new data. Figure 10 The bar chart in [Figure] clearly shows the difference between the PSNR and SSIM reconstructed using SCRWin Transformer and CSWin Transformer.
[0080] In summary, the present invention proposes a two-stage end-to-end model. The first stage is the registration stage, in which the auxiliary-modal MRI image is aligned with the target-modal MRI image to eliminate the spatial position difference between the two. The second stage is the reconstruction stage, and a reconstruction model is designed based on the difference image. For the undersampled target-modal MRI images, the residual images between them and the potential high-resolution magnetic resonance images are modeled. The residual images are respectively modeled as the unique and auxiliary features extracted from the input target image and the fused feature map. This ensures the accuracy of feature extraction and the accuracy of the consistent information extracted from the auxiliary-modal MRI image.
[0081] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A Transformer-based multimodal MRI reconstruction method, characterized in that: Perform the following steps S1 to S3 to reconstruct the low-resolution target modality MRI image into a high-resolution image: Step S1: acquiring a low-resolution target modality MRI image and a corresponding high-resolution auxiliary modality MRI image; Step S2: constructing an initial multimodal MRI image reconstruction network model, forming an MRI image pair with the target modality MRI image and the auxiliary modality MRI image, and inputting the initial multimodal MRI image reconstruction network model for training, outputting the reconstructed image, and obtaining a trained multimodal MRI image reconstruction network model; The multimodal MRI image reconstruction network model includes a spatial registration module and an image reconstruction module, wherein the spatial registration module performs spatial position alignment of the auxiliary modality MRI image to the target modality MRI image, and the image reconstruction module extracts auxiliary information from the MRI image pair; The specific steps of step S2 are as follows: Step S2.1: In the spatial registration module, the auxiliary modality MRI image is subjected to flow field prediction and then spatial transformation to obtain a registered image; Step S2.2: inputting the registration image and the target modality MRI image into the initial multimodal MRI image reconstruction network model, passing through the image reconstruction module, outputting the reconstructed image, and obtaining a trained multimodal MRI image reconstruction network model; The specific steps of step S2.2 are as follows: Step S2.2.1: Input the registration image and the target modality MRI image into the CNN network to obtain the fusion features; Step S2.2.2: Input the fused features and the target modality MRI image into a U-Net network with Transformer as the basic component and no parameter sharing to predict auxiliary features and unique features; Step S2.2.3: Input the auxiliary features and unique features into the CNN network to obtain residual features; Step S2.2.4: Add the residual features and the target modality MRI image to obtain the reconstructed image; Step S2.2.5: Use the reconstructed image as the target modality MRI image, input the target modality MRI image and the registration image into the CNN network, obtain the next fusion feature, and iterate steps S2.2.1-S2.2.4; until the residual feature can restore the lost details of the under-sampled image and remove unnecessary artifacts, stop the iteration, and obtain a trained multi-modality MRI image reconstruction network model; The iterative process of step S2.2.5 is as follows: ; ; ; ; ; Among them, U is the auxiliary feature, A is the unique feature, ED A , ED U are the codecs for extracting auxiliary features and unique features, respectively. t is the number of iterations, diff is the residual graph, Rec x is the reconstructed image, Rec A is the fusion feature, It is a fully sampled auxiliary modality MRI image; Step S3: input the paired under-sampled and fully-sampled MRI image pairs into the trained multi-modal MRI image reconstruction network model, output the high-resolution image after target modality reconstruction, and complete the reconstruction of the target modality MRI image.
2. The Transformer-based multimodal MRI reconstruction method according to claim 1, characterized in that: The specific steps of step S2.1 are as follows: Step S2.1.1: perform channel stitching on the target modality MRI image and the auxiliary modality MRI image to obtain a combined image, and input the combined image into the VoxelMorph network to obtain a predicted flow field; Step S2.1.2: spatially transform the auxiliary modality MRI image using the predicted flow field to obtain the current registration image; Step S2.1.3: Use the last registered image as the auxiliary modality MRI image, repeat steps S2.1.1-S2.1.2 to obtain the registered image.
3. The Transformer-based multimodal MRI reconstruction method according to claim 2, characterized in that: The specific method of step S2.1.1 is as follows: The combined image is input into the encoder of the U-Net network for three maximum pooling downsamplings. Each time the feature map is downsampled, the size is reduced to half of the original size, and the number of channels is increased to twice the original size. The obtained high-level feature map passes through the bottleneck layer and is input into the decoder. The decoder uses the transposed convolution high-level feature pairs for upsampling. Each time the feature map is upsampled, the size is increased to twice the original size, and the number of channels is reduced to half the original size. Finally, a 2-channel predicted flow field is output.
4. The Transformer-based multimodal MRI reconstruction method according to claim 2, characterized in that: The specific steps of step S2.1.2 are as follows: Step S2.1.2.1: The localization network predicts the affine transformation matrix using the image features extracted by the CNN network θ ; Step S2.1.2.2: The mesh generator is based on the affine transformation matrix θ Generate a position transformation mapping relationship before and after the transformation; Step S2.1.2.3: Sample pixel values according to the position transformation mapping relationship and perform bilinear interpolation, as shown in the following formula: ; ; In the formula, is the output image The pixel value of the position, ; and are the height and width of the output image, respectively. C represents the channel dimension; is the pixel value at position (n,m,c) of the combined image, , H and W are the height and width of the combined image, respectively. θ 11 ~ θ 23 represents the parameters of the affine transformation, k represents bilinear interpolation, and represents the parameters of the interpolation function, Represents the coordinates of the original image.
5. The Transformer-based multimodal MRI reconstruction method according to claim 4, characterized in that: The specific method of step S2.2.2 is as follows: The U-Net network predicts auxiliary features and unique features based on cross-window multi-head attention, as shown in the following formula: ; ; ; ; in, d k is a matrix Q and K The dimensions of the column vectors in , Q is the query matrix, V is the value matrix, K It is the content of attention, head i Representative i Head attention, and It works on Q Matrix and K The weight on the matrix, the subscript i represents the calculation of the first i Head attention, W O is the output weight represented by the convolution kernel, QK T Used for calculation Q exist V The attention weights on , 1…c / 2 is the first half of the channel, c / 2+1…c is the second half of the channel, stage If it is an odd number, the vertical local window attention is calculated, otherwise the horizontal window attention is calculated.
Citation Information
Patent Citations
Multi-modal MR image super-resolution method based on gradient attention enhancement
CN116823613A