Construction method and application of MRI (Magnetic Resonance Imaging) image reconstruction model based on residual degasser diffusion
By combining a deep separable gating network with a dual-path attention layer, the deficiencies of feature extraction and fusion in MRI image reconstruction models in glioma diagnosis are addressed, high-precision image detail recovery and global consistency are achieved, and the accuracy of glioma MRI image reconstruction is improved.
Patent Information
- Application Number
- CN202511136296.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing MRI image reconstruction models have difficulty in effectively retaining high-frequency details when dealing with gliomas, resulting in blurred tumor boundaries, inaccurate reproduction of complex morphologies, and insufficient generalization capabilities, which affects the accuracy of lesion identification and positioning.
A deep separable gating network is used to replace traditional skip connections. Combining deep separable convolution and gating mechanism, a dual-path attention layer is set to guide the model to focus on the tumor area, optimize feature extraction and fusion, and improve image detail recovery and global context consistency.
It improves the accuracy of MRI image reconstruction of brain gliomas, enhances the feature extraction capability of complex lesions and atypical anatomical structures, and improves the accuracy and stability of image reconstruction.
Smart Images

Figure CN120635246A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical image processing, and in particular to a method for constructing and applying an MRI image reconstruction model based on residual de-diffusion. Background Art
[0002] Gliomas, the most common primary intracranial tumor, account for 1.8% of human malignant tumors and contribute to 2.3% of cancer mortality. Their incidence has been increasing in recent years, and their impact continues to expand. According to the World Health Organization (WHO) histopathological classification, gliomas can be divided into grades I and II: low-grade gliomas (LGGs) and grades III and IV: high-grade gliomas (HGGs). Magnetic resonance imaging (MRI), a core tool for glioma diagnosis and treatment planning, can provide critical information such as tumor grade, location, morphology, and size, directly influencing treatment options and patient survival expectations.
[0003] In clinical practice, physicians often rely on multi-sequence MRI for reliable diagnosis, as a single sequence often fails to fully characterize lesions. However, due to factors such as patient psychological state, individual differences, and scanning noise, MRI images often exhibit artifacts (such as Gibbs artifacts and motion artifacts), sequence omissions, and difficulty in interpretation, severely impacting diagnostic accuracy.
[0004] In recent years, advances in deep learning technology have revolutionized the field of medical image reconstruction. Generative adversarial networks (GANs) significantly enhance image structure preservation through adversarial training mechanisms. Advanced variants have achieved multi-to-multimodal synthesis, breaking through the limitations of traditional single-modality conversion. However, inherent limitations of GANs, such as mode collapse, training instability, and hyperparameter sensitivity, have significantly hindered their clinical translation.
[0005] The denoising diffusion probabilistic model (DDPM) is a robust alternative, demonstrating significant advantages in MRI image reconstruction thanks to its stable training dynamics and high-fidelity generation capabilities. Compared to traditional methods, this iterative optimization model excels in maintaining anatomical consistency and reducing artifacts. Its application has expanded from generative tasks such as image super-resolution and cross-modal synthesis to discriminative tasks such as lesion segmentation and tumor classification, significantly reducing the risk of mode collapse in adversarial frameworks. Building on this foundation, the residual denoising diffusion model (RDDM), by introducing a dual framework of residual diffusion and noise diffusion, achieves precise control over the denoising process, effectively maintaining anatomical consistency and ensuring accurate restoration of detail and structure in the generated images.
[0006] Despite continuous technological advancements, clinical translation still faces two key challenges: 1. Limitations of the existing U-net denoising module: The standard U-net architecture relies on skip connections for feature fusion. However, when processing complex lesions such as glioma MRI images, it struggles to effectively preserve high-frequency details, resulting in blurred tumor boundaries and inaccurate reproduction of complex morphologies. In scenarios with blurred tumor boundaries or complex lesion morphologies, fine-grained feature loss is particularly prominent, impacting lesion identification and localization accuracy.
[0007] 2. Insufficient generalization ability of RDDM: When faced with atypical anatomical manifestations (such as irregular tumor boundaries and low contrast with surrounding tissues), existing models have difficulty effectively extracting detailed features, resulting in reduced image reconstruction accuracy and affecting the effectiveness of clinical diagnosis.
[0008] Therefore, there is an urgent need for a technical solution that can optimize feature extraction and fusion strategies and improve the reconstruction accuracy of complex lesions and atypical anatomical structures, so as to address the shortcomings of existing technologies in glioma MRI image reconstruction and provide a more reliable auxiliary tool for clinical diagnosis. Summary of the Invention
[0009] The embodiments of the present application provide a method for constructing and applying an MRI image reconstruction model based on residual de-diffusion. By adopting a deep separable gating network to replace the traditional skip connection, combining the effective capture of spatially adjacent pixel information and the improvement of computational efficiency by the deep separable convolution, and the precise adjustment of the inter-layer information flow by the gating mechanism, fine-tuning of features is achieved, cross-layer information fusion is enhanced, image detail recovery is effectively promoted and global context consistency is maintained, thereby improving the accuracy of MRI image reconstruction of brain gliomas.
[0010] In a first aspect, an embodiment of the present application provides a method for constructing an MRI image reconstruction model based on residual de-diffusion, the method comprising: Acquire a first data set storing multimodal MRI images, gradually add noise to any modality of each multimodal MRI image in the first data set to obtain a second data set, and use the second data set as a training data set; The noisy mode in the multimodal MRI image in the training data set is used as the target mode, and the non-noisy mode is used as the prior mode. A prior tumor mask is obtained based on the prior mode. Each multimodal MRI image in the training data set and the corresponding prior tumor mask are used as training samples. The training samples are input into the constructed MRI image reconstruction architecture to obtain a reconstructed MRI image. The MRI image reconstruction architecture is based on a U-net network, including an encoding module, a decoding module, and a residual module, and a dual-path attention layer is set after each encoding level of the encoding module and each decoding level of the decoding module, wherein the encoding level / decoding level processes the multimodal MRI image in the training sample, and the dual-path attention layer performs attention calculation on the output feature map of the encoding level / decoding level based on the prior tumor mask in the training sample; The reconstructed MRI image is verified based on the multimodal MRI image in the first data set and the loss function is calculated. The parameters of the MRI image reconstruction architecture are adjusted based on the loss function. When the loss function meets the set conditions or reaches the preset number of iterations, the parameters of the MRI image reconstruction architecture are saved to obtain the MRI image reconstruction model.
[0011] In a second aspect, an embodiment of the present application provides an MRI image reconstruction method, comprising: An MRI image to be reconstructed is obtained, where the MRI image to be reconstructed is a multimodal MRI image in which a single modality is missing or unusable. The missing or unusable modality in the MRI image to be reconstructed is marked as a target modality and input into a constructed MRI image reconstruction model. The MRI image reconstruction model reconstructs the target modality in the MRI image to be reconstructed to obtain a full-modality MRI image.
[0012] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute a method for constructing an MRI image reconstruction model based on residual de-diffusion.
[0013] In a fourth aspect, an embodiment of the present application provides a readable storage medium, in which a computer program is stored. The computer program includes a program code for controlling a process to execute a process, and the process includes a method for constructing an MRI image reconstruction model based on residual de-diffusion.
[0014] The main contributions and innovations of the present invention are as follows: The embodiment of the present application adopts a deep separable gated network to replace the traditional skip connection, combining deep separable convolution with a gating mechanism to achieve fine control of features, enhance cross-layer information fusion, effectively promote image detail recovery and maintain global context consistency; the embodiment of the present application sets a dual-path attention layer after each encoding level of the encoding module and each decoding level of the decoding module, performs attention calculation based on the prior tumor mask, guides the model to focus on the tumor area, and improves the feature extraction capability of complex lesions and atypical anatomical structures; the embodiment of the present application is based on the U-net network, the first three encoding levels of the encoding module are residual downsampling layers, and the last two are encoding overlapping block embedding networks, the first two decoding levels of the decoding module are decoding overlapping block embedding networks, and the last three are residual upsampling layers. Through a reasonable hierarchical structure design, the feature extraction and processing flow is optimized, which helps to improve the accuracy of brain glioma MRI image reconstruction; the residual module of the embodiment of the present application adopts a dual-path deep separable convolution to process the output feature map and combines it with a gating mechanism to further enhance cross-layer information fusion, promote image detail recovery, maintain global context consistency, and improve reconstruction accuracy.
[0015] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 1 is an overall structural diagram of an MRI image reconstruction model based on residual de-diffusion according to an embodiment of the present application; Figure 2 is a structural diagram of a coded overlapping embedding network according to an embodiment of the present application; Figure 3 is a structural diagram of a bottleneck module according to an embodiment of the present application; Figure 4 is a structural diagram of a residual module according to an embodiment of the present application; Figure 5 is a structural diagram of a dual-path attention layer according to an embodiment of the present application; Figure 6 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0017] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.
[0018] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments, and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0019] Example 1 The embodiment of the present application provides a method for constructing an MRI image reconstruction model based on residual de-diffusion, which replaces the traditional skip connection with a deep separable gating network, combines the effective capture of spatial neighboring pixel information and the improvement of computational efficiency by deep separable convolution, and the precise adjustment of inter-layer information flow by the gating mechanism, thereby achieving fine control of features, enhancing cross-layer information fusion, effectively promoting image detail recovery and maintaining global context consistency, thereby improving the accuracy of MRI image reconstruction of brain gliomas. Specifically, reference Figure 1 , the method comprising: Acquire a first data set storing multimodal MRI images, gradually add noise to any modality of each multimodal MRI image in the first data set to obtain a second data set, and use the second data set as a training data set; The noisy mode in the multimodal MRI image in the training data set is used as the target mode, and the non-noisy mode is used as the prior mode. A prior tumor mask is obtained based on the prior mode. Each multimodal MRI image in the training data set and the corresponding prior tumor mask are used as training samples. The training samples are input into the constructed MRI image reconstruction architecture to obtain a reconstructed MRI image. The MRI image reconstruction architecture is based on a U-net network, including an encoding module, a decoding module, and a residual module, and a dual-path attention layer is set after each encoding level of the encoding module and each decoding level of the decoding module, wherein the encoding level / decoding level processes the multimodal MRI image in the training sample, and the dual-path attention layer performs attention calculation on the output feature map of the encoding level / decoding level based on the prior tumor mask in the training sample; The reconstructed MRI image is verified based on the multimodal MRI image in the first data set and the loss function is calculated. The parameters of the MRI image reconstruction architecture are adjusted based on the loss function. When the loss function meets the set conditions or reaches the preset number of iterations, the parameters of the MRI image reconstruction architecture are saved to obtain the MRI image reconstruction model.
[0020] In some specific embodiments, the first data set described in this solution can be multimodal MRI images directly acquired in various hospital systems, or an open source data set can be directly acquired, such as BraTS2021.
[0021] In some embodiments, the present scheme gradually adds noise to any modality of each multimodal MRI image in the first data set at a preset time step. In the present scheme, the multimodal MRI images include T1 modality, T1ce modality, T2 modality and FLAIR modality. In the process of gradually adding noise to the multimodal MRI images in the first data set, the number of images added with noise in T1 modality, T1ce modality, T2 modality and FLAIR modality is made basically the same, thereby ensuring that the trained MRI image reconstruction model can reconstruct MRI images of any modality.
[0022] In some specific embodiments, when obtaining a prior tumor mask, this solution uses a segmentation algorithm to process any one or more non-noise-added modes in a multimodal MRI image to obtain a prior tumor mask. For example, when the T1 mode is a noisy mode, the T1ce mode, T2 mode and FLAIR mode are non-noise-added modes, and a segmentation algorithm is used to segment the T1ce mode or T2 mode or FLAIR mode to obtain a prior tumor mask, or a weighted average of the segmentation results of the T1ce mode, T2 mode and FLAIR mode is performed to obtain a prior tumor mask.
[0023] Specifically, using the prior tumor mask as model input can guide the model to pay more attention to the tumor area, thereby better reconstructing the missing modality.
[0024] In some embodiments, the encoding module includes 5 encoding levels, the input of each encoding level is used as an input feature map, and the output is an output feature map. The first 3 encoding levels are residual downsampling layers, and the last two encoding levels are encoding overlapping block embedding networks. The decoding module includes 5 decoding levels, the first two decoding levels are decoding overlapping block embedding networks, and the last 3 decoding levels are residual upsampling layers.
[0025] Specifically, before the input feature map is processed by the encoding layer, the input feature map is mapped to the initial feature dimension. For example, the input feature map is represented as , where B is the batch size, C is the number of channels, H is the width of the input feature map, and W is the height of the input feature map. A 7×7 convolutional layer is used to process the input feature map to provide appropriate feature representation for subsequent processing. The formula is expressed as:
[0026] Furthermore, a time vector of a preset time step is obtained, and the time vector is processed continuously through two linear transformation units to obtain a time embedding vector, and each encoding level and decoding level is weighted by the time embedding vector.
[0027] Specifically, use SinusoidalPosEmb to obtain the sine and cosine vectors of the preset time step. The formula is expressed as:
[0028] Among them, time i is the preset time step and is the corresponding frequency coefficient.
[0029] Specifically, the formula for obtaining the time embedding vector by processing the time vector through two linear transformation units continuously is expressed as follows:
[0030]
[0031] in, is the weight of the first linear transformation, is the bias of the first linear transformation, is the weight of the second linear transformation, is the bias of the second linear transformation, is the activation function.
[0032] Specifically, by using time embedding to weight each encoding level and decoding level, features can be modulated and optimized in the time dimension.
[0033] Specifically, in the encoding layer, the downsampled feature map is combined with the temporal embedding and further processed through the residual connection. The downsampling formula is expressed as:
[0034] Among them, X is the input feature map, Downsample means downsampling, is the downsampling result.
[0035] In some embodiments, the structure of the encoding overlapping embedding network is as follows Figure 2As shown, the encoding overlapping block embedding network consists of an overlapping patch embedding layer, a transformation layer, a KAN network and a normalization layer. The overlapping patch embedding layer converts the input feature map into an embedded feature map. The transformation layer shifts the embedded feature map to obtain a transformed feature map. The KAN network processes the transformed feature map and performs a residual connection with the input feature map, and outputs it by a normalization layer. The structure of the decoding overlapping block embedding network is the same as that of the encoding overlapping block embedding network.
[0036] Specifically, the overlapping patch embedding layer implements the overlapping patch embedding function, converts the input image into embedded features through the convolution layer, and adjusts the shape of the features in the forward propagation for subsequent KAN network processing.
[0037] Specifically, the transformation layer shifts the embedded feature map left and right and up and down through the Shift operation, helping the model capture richer contextual information and spatial security relationships, better understand the local and global features of the input data, and thus improve the overall performance.
[0038] Specifically, the KAN network is constructed from a series of connected KAN layers, each containing a set of learnable one-dimensional activation functions. By stacking multiple KAN layers, the model is able to learn and extract the complex features of high-dimensional data layer by layer. Each KAN layer applies a nonlinear transformation to the input using a learnable activation function, thereby enhancing the model's expressive power. Ultimately, these layers perform in-depth processing on the input data, effectively capturing the complex relationships and underlying patterns in the data and improving model performance. The specific calculation formula for the KAN network is as follows:
[0039] in, Represents the i-th layer of the entire KAN network. Each KAN layer has n in dimensional input and n out dimensional output, including n in *n out A learnable activation function.
[0040] In some other embodiments, the decoding overlapping block embedding network uses D_Single convolution for decoding, enhancing feature representation through convolutional layers and temporal embedding. This module design allows for complex transformations in feature space while preserving temporal information, thereby improving the model's performance in image reconstruction. By integrating temporal embedding, the model can better understand the dynamic changes in input data and enhance generation.
[0041] In some embodiments, a bottleneck module is set between the encoding module and the decoding module, and the structure of the bottleneck module is as follows: Figure 3As shown, the bottleneck module performs self-attention calculation on the output feature map of the encoding module and inputs it into the decoding module.
[0042] Specifically, the self-attention mechanism added to the bottleneck part can help the model capture global information and overcome the limitations of local areas by calculating the relationship between features and weighting the input features. The self-attention mechanism is expressed as follows:
[0043] Among them, d k Represents feature dimensions, Q, K, and V represent query, key, and value respectively.
[0044] In some embodiments, the structure of the residual module is as follows Figure 4 As shown, the residual module uses a two-way depth-separable convolution to process the output feature map to obtain a first depth-separable feature and a second depth-separable feature respectively. After smoothing the first depth-separable feature, it is calculated with the second depth-separable feature by element-by-element multiplication to obtain a gated feature. The gated feature is residually connected with the output feature map and then output.
[0045] Specifically, in the U-net structure, the residual module is used to perform residual connections between the output feature maps of each encoding level in the encoding module and the input of the decoding level of the corresponding size, thereby further enhancing the fusion of cross-layer information, promoting the recovery of image details, and maintaining the consistency of the global context, thereby improving the accuracy of image reconstruction.
[0046] Specifically, for depth-wise separable convolution, a convolution kernel is applied independently to each channel, and the formula of depth-wise separable convolution is expressed as:
[0047] in, To try the result of separable convolution, X is the input feature map, K is the convolution kernel, i and j are spatial positions, k is the channel index, m is the number of input channels, and n is the number of output channels. Through depth-wise separable convolution, it can effectively capture information from spatially adjacent pixels while greatly improving computational efficiency. The gating mechanism regulates the flow of information between layers, allowing the network to highlight or suppress specific features as needed.
[0048] Specifically, the gating mechanism controls the feature flow through element-by-element multiplication, which is expressed as:
[0049] in, is the gated feature, 、 are the learnable weights, is the first depth-separable feature, is the second depth-wise separable feature, and SiLU is the activation function.
[0050] Specifically, in the gating mechanism, the SiLU activation function provides a smooth gradient flow in the gating mechanism, avoids the dead neuron problem, and contributes to better convergence of the model. Furthermore, after the depth-wise separable convolution and gating mechanism, 1×1 point convolution is used to adjust the number of channels.
[0051] In some embodiments, the structure of the dual-path attention layer is as follows Figure 5 As shown, the dual-path attention layer takes the output feature map of the encoding layer / decoding layer and the prior tumor mask of the corresponding size as input. The dual-path attention layer includes an implicit tensor attention path and a spatial attention path. The implicit tensor attention path performs tensor splitting on the output feature map and maps it to the non-negative space for attention calculation to obtain an implicit tensor attention result. The spatial attention path takes the output feature map and the prior tumor mask of the corresponding size as input, and allocates attention to the output feature map based on the prior tumor mask to obtain a spatial attention result. The spatial attention result, the implicit tensor attention result and the multimodal MRI image feature map are added as the output of the dual-path attention layer. It is worth mentioning that the output of the dual-path attention layer is used as the input feature map of the next encoding layer / decoding layer.
[0052] Furthermore, in the implicit tensor attention path, the input feature map is first layer-normalized, and then a 1×1 convolution operation is used to convert the layer-normalized result into query, key, and value. The result after the convolution operation is sliced into three tensors, representing Q, K, and V respectively. The formula is expressed as:
[0053] Among them, chunk represents the tensor slicing operation, chunk3, dim=1 means slicing into 3 fragments along the direction of dimension 1, and the shapes of Q, K, and V are (b, h, c, h*w).
[0054] Then perform Softmax processing on the query Q and key K respectively, and normalize the value V. The formula is expressed as:
[0055]
[0056]
[0057] After obtaining Q, K, and V, Q and K are mapped to the non-negative space for attention calculation. The formula is expressed as:
[0058] in, is the attention calculation result, and is the Softmax feature mapping function, which is used to map Q and K to the non-negative space.
[0059] Specifically, when performing attention calculation, by first calculating and , and then multiply by , thus avoiding explicit calculation , reducing the computational complexity.
[0060] Specifically, after obtaining the attention calculation result, the channel dimension of the attention calculation result is unified with the channel dimension of the input feature map, and then layer normalization is performed to obtain the implicit tensor attention result. That is to say, the attention result is first reconstructed, and the multi-head output is spliced back to the input feature map format, that is, the [B, heads, dim_head, N] dimension is converted to [B, C, H, W], and the multi-head pathological features are integrated. Finally, the channel dimension is adjusted to match the input feature map through 1*1 convolution, and finally the LayerNorm layer normalization is used to stabilize the feature distribution and then the implicit tensor attention result is output.
[0061] Furthermore, in the spatial attention path, the global context features of the input feature map are calculated by average pooling and maximum pooling, and the global context features are processed using a convolutional layer to obtain a spatial attention map. The spatial attention map is feature multiplied with the prior tumor mask to obtain a spatial feature map, and the spatial feature map is then element-wise multiplied with the input feature map to obtain the spatial attention result.
[0062] Specifically, when processing the input feature map, the global context representation is calculated through average pooling and maximum pooling to obtain the mean and maximum value of each channel in the spatial dimension respectively. Then, the two are spliced together along the channel dimension to obtain the global context feature. The global context feature is a more expressive feature expression. Subsequently, the spliced features are passed through the convolutional layer to generate a spatial attention map, focusing on capturing the key areas in the feature map.
[0063] Furthermore, the prior tumor mask is resized to the same size as the corresponding input feature map using bilinear interpolation. The resized prior tumor mask is further processed through convolution to generate a mask of size (3, H, W). When calculating spatial attention, the convolved spatial attention map is normalized using a sigmoid activation. The resulting spatial attention map is multiplied with the prior tumor mask, thereby introducing anatomical prior information to guide attention allocation. Finally, the resulting spatial feature map is element-wise multiplied with the input feature map to adjust the feature map's focus area, ensuring that the model focuses on the most discriminative anatomical regions.
[0064] Specifically, the dual-path attention layer can effectively combine anatomical prior information and spatial attention mechanism to improve the performance of the model in medical image segmentation tasks, especially showing significant advantages in unclear boundaries and complex tissue structures.
[0065] In some specific embodiments, this solution can use any loss function as a condition for terminating training.
[0066] Table 1 shows the evaluation results of the MRI modality conversion performance, including the mean and standard deviation of the peak signal-to-noise ratio (PSNR (dB)), structural similarity index (SSIM), and learning-perceptual image patch similarity (LPIPS (*10-2)).
[0067] Table 1. Evaluation of MRI modality conversion indicators for glioma
[0068] As shown in Table 1, when T1 is the target modality, one-to-one modality conversion using T2 as the source modality achieves optimal performance in terms of the three key evaluation metrics: PSNR, SSIM, and LPIPS. PSNR, SSIM, and LPIPS reach 26.954±2.010 (dB), 0.928±0.026, and 6.400±2.439 (*10-2), respectively. T2 as the source modality performs slightly better in generating T1 than T1ce, likely due to interference from the enhanced information. When T1ce is the target modality, T1 performs best as the source modality, achieving PSNR, SSIM, and LPIPS of 28.668±1.866 (dB), 0.940±0.018, and 4.844±1.708 (*10-2), respectively. The T1-to-T1ce modality conversion is the most stable and consistent with the T1ce modality's characteristics. Since they share the same physical imaging mechanism, the conversion process only needs to consider how to simulate the effects of the contrast agent, without considering differences in other imaging mechanisms. When T2 is the target modality, T1ce performs best as the source modality, achieving PSNR, SSIM, and LPIPS metrics of 28.486±1.803 (dB), 0.932±0.025, and 5.615±1.810 (*10-2), respectively. Active tumor areas in T1ce images and high signal areas (such as edema or tumor) in T2 images exhibit spatial overlap or similarity. This spatial correlation allows for more accurate mapping of regional features during modality conversion, thereby improving conversion quality. When FLAIR was the target modality, the T1 modality achieved the best performance as a source modality, with PSNR, SSIM, and LPIPS reaching 27.244±2.253 (dB), 0.903±0.033, and 7.481±1.750 (*10-2), respectively. When FLAIR was the target modality, the conversion performance of all modalities was generally poor, likely due to FLAIR's characteristic suppression of water signals, which makes it difficult for other modalities to accurately generate its features. In summary, the optimal source modality for generating the T1 modality was T2, the optimal source modality for generating the T2 modality was T1ce, and the optimal source modality for generating both the T1ce modality and the FLAIR modality was T1.
[0069] In the experiment comparing the modality completion effect with other models, AttKAN-RDDM demonstrated excellent performance on the BraTS2021 dataset. As shown in Table 2, facing many models such as Pix2Pix, MMGAN, ResViT, Restormer, FgC2F-UDiff, RDDM, AttKAN-RDDM stands out in the key indicators of PSNR, SSIM and LPIPS. Specifically, when T1 and FLAIR are used as target modalities, the method we proposed is slightly higher than the existing SOTA models in terms of PSNR, SSIM and LPIPS indicators. In terms of PSNR indicator, it is 26.954±2.010 (dB) higher than FgC2F-UDiff's 26.588±1.942 (dB) and 27.244±2.253 (dB) higher than Restormer's 26.847±2.272 (dB); in terms of SSIM indicator, it is 0.9 higher than FgC2F-UDiff's 26.588±1.942 (dB) and 0.9 higher than Restormer's 26.847±2.272 (dB). The proposed method significantly outperforms existing state-of-the-art models in PSNR, SSIM, and LPIPS (7.511±1.758*10-2). Furthermore, the proposed method achieves significantly lower standard deviations of PSNR, SSIM, and LPIPS across all modalities than the competing models, demonstrating its robustness to diverse samples. Overall, the experimental results of AttKAN-RDDM on the BraTS2021 dataset fully verified its effectiveness and advancement in the task of glioma MRI modality completion, especially in tumor lesion area mapping and model stability, which is superior to existing mainstream modality completion methods.
[0070] Table 2. Quantitative comparison of sequence completion with SOTA methods on the BraTS2021 dataset
[0071] Table 3. Quantitative comparison of Gibbs artifact removal with SOTA methods on the BraTS2021 dataset
[0072] As shown in Table 3, in experiments comparing Gibbs artifact removal with other models, AttKAN-RDDM demonstrated superior performance on the Gibbs artifact dataset. Our AttKAN-RDDM outperformed many models, including Pix2Pix, MMGAN, ResViT, Restormer, FgC2F-UDiff, and RDDM, in key metrics such as PSNR, SSIM, and LPIPS. Specifically, when performing Gibbs artifact removal on the T2 modality, our proposed method achieved the best results in terms of PSNR, SSIM, and LPIPS, significantly outperforming existing state-of-the-art models in these metrics. Furthermore, when performing Gibbs artifact removal on the T1, T1ce, and FLAIR modalities, AttKAN-RDDM also outperformed existing state-of-the-art models in terms of PSNR, SSIM, and LPIPS.
[0073] Example 2 Based on the same concept, this application also proposes an MRI image reconstruction method based on an MRI image reconstruction model, comprising: An MRI image to be reconstructed is obtained, where the MRI image to be reconstructed is a multimodal MRI image in which a single modality is missing or unusable. The missing or unusable modality in the MRI image to be reconstructed is marked as a target modality and input into a constructed MRI image reconstruction model. The MRI image reconstruction model reconstructs the target modality in the MRI image to be reconstructed to obtain a full-modality MRI image.
[0074] Example 3 This embodiment also provides an electronic device, referring to Figure 6 , includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.
[0075] Specifically, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0076] Memory 404 may include a large-capacity memory 404 for data or instructions. By way of example, and not limitation, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to the data processing device. In certain embodiments, memory 404 is non-volatile memory. In certain embodiments, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM) or a flash memory (FLASH), or a combination of two or more of these. In appropriate circumstances, the RAM may be a static random access memory (SRAM) or a dynamic random access memory (DRAM), wherein the DRAM may be a fast page mode dynamic random access memory 404 (FPMDRAM), an extended data output dynamic random access memory (EDODRAM), a synchronous dynamic random access memory (SDRAM), etc.
[0077] The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402 .
[0078] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any one of the methods for constructing an MRI image reconstruction model based on residual de-diffusion in the above embodiments.
[0079] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .
[0080] Transmission device 406 can be used to receive or transmit data via a network. Specific examples of such networks may include wired or wireless networks provided by the electronic device's communications provider. In one embodiment, the transmission device includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0081] The input and output device 408 is used to input or output information. In this embodiment, the input information may be multimodal MRI images, prior tumor masks, etc., and the output information may be parameter information of the MRI image reconstruction model, etc.
[0082] Optionally, in this embodiment, the processor 402 may be configured to execute the following steps through a computer program: Acquire a first data set storing multimodal MRI images, gradually add noise to any modality of each multimodal MRI image in the first data set to obtain a second data set, and use the second data set as a training data set; The noisy mode in the multimodal MRI image in the training data set is used as the target mode, and the non-noisy mode is used as the prior mode. A prior tumor mask is obtained based on the prior mode. Each multimodal MRI image in the training data set and the corresponding prior tumor mask are used as training samples. The training samples are input into the constructed MRI image reconstruction architecture to obtain a reconstructed MRI image. The MRI image reconstruction architecture is based on a U-net network, including an encoding module, a decoding module, and a residual module, and a dual-path attention layer is set after each encoding level of the encoding module and each decoding level of the decoding module, wherein the encoding level / decoding level processes the multimodal MRI image in the training sample, and the dual-path attention layer performs attention calculation on the output feature map of the encoding level / decoding level based on the prior tumor mask in the training sample; The reconstructed MRI image is verified based on the multimodal MRI image in the first data set and the loss function is calculated. The parameters of the MRI image reconstruction architecture are adjusted based on the loss function. When the loss function meets the set conditions or reaches the preset number of iterations, the parameters of the MRI image reconstruction architecture are saved to obtain the MRI image reconstruction model.
[0083] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be repeated here.
[0084] In general, various embodiments may be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0085] The embodiments of the present invention may be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros may be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. A computer program product may include one or more computer executable components that are configured to perform an embodiment when the program is run. One or more computer executable components may be at least one software code or a portion thereof. In addition, it should be noted at this point that, for example, Figure 6Any block of the logic flow in the program may represent program steps, or interconnected logic circuits, blocks and functions, or a combination of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs, etc. Physical media are non-transitory media.
[0086] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0087] The above embodiments merely illustrate several embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for constructing an MRI image reconstruction model based on residual de-diffusion, characterized in that: The following steps are involved: Acquire a first data set storing multimodal MRI images, gradually add noise to any modality of each multimodal MRI image in the first data set to obtain a second data set, and use the second data set as a training data set; The noisy mode in the multimodal MRI image in the training data set is used as the target mode, and the non-noisy mode is used as the prior mode. A prior tumor mask is obtained based on the prior mode. Each multimodal MRI image in the training data set and the corresponding prior tumor mask are used as training samples. The training samples are input into the constructed MRI image reconstruction architecture to obtain a reconstructed MRI image. The MRI image reconstruction architecture is based on a U-net network, including an encoding module, a decoding module, and a residual module, and a dual-path attention layer is set after each encoding level of the encoding module and each decoding level of the decoding module, wherein the encoding level / decoding level processes the multimodal MRI image in the training sample, and the dual-path attention layer performs attention calculation on the output feature map of the encoding level / decoding level based on the prior tumor mask in the training sample; The reconstructed MRI image is verified based on the multimodal MRI image in the first data set and the loss function is calculated. The parameters of the MRI image reconstruction architecture are adjusted based on the loss function. When the loss function meets the set conditions or reaches the preset number of iterations, the parameters of the MRI image reconstruction architecture are saved to obtain the MRI image reconstruction model.
2. The method for constructing an MRI image reconstruction model based on residual elimination diffusion according to claim 1, characterized in that: The encoding module includes 5 encoding levels, the first 3 encoding levels are residual downsampling layers, and the last two encoding levels are encoding overlapping block embedding networks. The decoding module includes 5 decoding levels, the first two decoding levels are decoding overlapping block embedding networks, and the last 3 decoding levels are residual upsampling layers.
3. The method for constructing an MRI image reconstruction model based on residual elimination diffusion according to claim 1, characterized in that: The encoding overlapping block embedding network consists of an overlapping patch embedding layer, a transformation layer, a KAN network and a normalization layer. The overlapping patch embedding layer converts the input feature map into an embedded feature map. The transformation layer shifts the embedded feature map to obtain a transformed feature map. The KAN network processes the transformed feature map and performs a residual connection with the input feature map, and outputs it by a normalization layer. The decoding overlapping block embedding network has the same structure as the encoding overlapping block embedding network.
4. The method for constructing an MRI image reconstruction model based on residual elimination diffusion according to claim 1, characterized in that: A time vector of a preset time step is obtained, and the time vector is processed continuously through two linear transformation units to obtain a time embedding vector, and each encoding level and decoding level is weighted by the time embedding vector.
5. The method for constructing an MRI image reconstruction model based on residual elimination diffusion according to claim 1, characterized in that: The residual module uses a two-way depth-wise separable convolution to process the output feature map to obtain a first depth-wise separable feature and a second depth-wise separable feature respectively. After smoothing the first depth-wise separable feature, the gating mechanism is calculated with the second depth-wise separable feature by element-by-element multiplication to obtain a gated feature. The gated feature is residually connected with the output feature map and then output.
6. The method for constructing an MRI image reconstruction model based on residual elimination diffusion according to claim 1, characterized in that: The dual-path attention layer takes the output feature map output by the encoding layer / decoding layer and the prior tumor mask of the corresponding size as input. The dual-path attention layer includes an implicit tensor attention path and a spatial attention path. The implicit tensor attention path performs tensor splitting on the output feature map and maps it to the non-negative space for attention calculation to obtain an implicit tensor attention result. The spatial attention path takes the output feature map and the prior tumor mask of the corresponding size as input, and allocates attention to the output feature map based on the prior tumor mask to obtain a spatial attention result. The spatial attention result, the implicit tensor attention result and the multimodal MRI image feature map are added as the output of the dual-path attention layer.
7. The method for constructing an MRI image reconstruction model based on residual elimination diffusion according to claim 6, characterized in that: The bilinear interpolation method is used to adjust the size of the prior tumor mask to the same as the corresponding input feature map.
8. A MRI image reconstruction method, characterized in that: include: Obtain an MRI image to be reconstructed, where the MRI image to be reconstructed is a multimodal MRI image in which a single modality is missing or unusable, mark the missing modality or unusable modality in the MRI image to be reconstructed as a target modality and input the mark into an MRI image reconstruction model constructed by any one of the methods of claims 1 to 7, and reconstruct the target modality in the MRI image to be reconstructed by the MRI image reconstruction model to obtain a full-modality MRI image.
9. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and the processor is configured to run the computer program to execute the method for constructing an MRI image reconstruction model based on residual de-diffusion according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, and the process includes a method for constructing an MRI image reconstruction model based on residual de-diffusion according to any one of claims 1-7.
Citation Information
Patent Citations
Transform-based MRI (Magnetic Resonance Imaging) brain tumor image reconstruction method and system
CN118298067A
RDDM-based high-quality speaking face video generation method and system
CN118488266A
Depth estimation method and system based on indirect diffusion model
CN120031936A
Deep learning super resolution of medical images
WO2023183504A1
Cited By
Extremely low light image restoration method and device
CN121599860A