A Fast Magnetic Resonance Image Reconstruction Method Based on Latent Variable Diffusion Model
The MRI image is reconstructed in low-dimensional space through the latent variable diffusion model, combining full-frequency and high-frequency detail models, and solving the long scanning time and image artifact problems of magnetic resonance imaging, achieving rapid and high-quality MRI reconstruction.
Patent Information
- Application Number
- CN202510637556.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing magnetic resonance imaging technology has bottlenecks in the long scanning time and image artifact problems, which affect imaging quality and clinical applications. It is difficult to effectively describe the complexity of MRI image distribution.
Using a method based on latent variable diffusion model, the original k-space data is encoded into a low-dimensional latent space, combined with full-frequency structure reconstruction and high-frequency detail refinement model, and fast and high-quality MRI reconstruction is achieved through encoder and decoder pre-training.
Significantly reduces computational complexity and reconstruction time, while retaining image structural details and clear edges, improving reconstruction image quality.
Smart Images

Figure CN120182422B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing research, and particularly relates to a fast magnetic resonance image reconstruction method based on a latent variable diffusion model. Background Art
[0002] Magnetic resonance imaging (MRI), as one of the most important contemporary medical auxiliary means, has become an indispensable imaging technology for non-invasively obtaining diagnostic information in clinical practice because it can provide fine soft tissue contrast and high-precision spatial resolution. However, as an imaging technology based on nuclear magnetic resonance, its imaging process is affected by various factors, including physical properties, signal acquisition time, scanning parameter settings, and physiological factors. This makes its scanning often require collecting a large amount of signal data to construct an image, and this process involves a series of complex physical processes, including the generation of magnetic field gradients, the emission and reception of pulse sequences, signal processing, etc., and each signal acquisition requires a certain amount of time. The long acquisition time will increase the discomfort of patients, and at the same time, due to factors such as the patient's breathing and heartbeat, image artifacts are caused, which in turn leads to misdiagnosis. These factors together have led to the major bottleneck problem of slow magnetic resonance imaging speed, which not only affects the imaging quality but also is not conducive to its clinical application. To sum up, magnetic resonance imaging can provide fine medical diagnostic information. If high-quality reconstruction of scanned images can be achieved while significantly shortening the scanning time, it will help to promote the clinical application of MRI technology. In the digital age, fast magnetic resonance image reconstruction, as a branch of medical image processing, has always been a challenging and attractive problem, with great research value and social significance.
[0003] Fast magnetic resonance image reconstruction is an intuitive image reconstruction problem. Briefly speaking, it aims to accelerate imaging by collecting a small amount of data. However, the lack of data will cause problems such as imaging artifacts. At this time, reconstruction algorithms need to be used to obtain the final target image. There are three major categories of early fast magnetic resonance imaging techniques: One is to improve the performance of magnetic resonance imaging hardware. For example, increasing the main magnetic field strength and gradient switching speed of the magnetic resonance scanner. However, this technique is too costly and prone to image blurring during imaging. In addition, the high-speed switching magnetic field will generate an eddy current field, which will lead to the generation of image artifacts and have a certain nerve electrical stimulation on human muscles. The second is to adopt parallel imaging algorithms, using the sensitivity of multiple receive coils to provide additional encoding information for reconstructing magnetic resonance images, allowing successful reconstruction of undamaged images from sparsely sampled image data. However, this technique is limited by various factors such as the accurate measurement of coil sensitivity distribution, sampling multiples, and the number of fitting blocks. Moreover, when the signal-to-noise ratio is very low or the acceleration acquisition factor is relatively large, the performance of these algorithms drops significantly. The third is the compressed sensing method. During the image reconstruction process, the prior information in the form of regularization provides constraints for optimization, thus better solving the image reconstruction problem. This technique utilizes the sparsity of the image under a certain mathematical transformation and the incoherence between data acquisition and sparse transformation to accurately reconstruct images from undersampled K-space data. However, the large reduction in data acquisition volume is prone to loss of details and image artifacts during reconstruction.
[0004] Traditional magnetic resonance imaging reconstruction techniques have made significant progress in the reconstruction field by introducing prior information of images. However, these methods often only explore the prior information of the reconstructed images and rarely involve reference images. However, with the further improvement of MRI acceleration, these traditional methods often suffer from loss of details and image artifacts. Driven by the success of deep learning in computer vision and natural image processing, fast magnetic resonance imaging reconstruction methods based on deep learning have received great attention. Different from traditional reconstruction methods, deep learning implicitly learns prior information from a large amount of magnetic resonance raw data scanned by different people through training a prior model. Due to the high non-linear characteristics of the data-driven prior extraction strategy, some deep learning methods are superior to traditional methods. Research data from many domestic and foreign teams show that deep learning methods not only have great potential in terms of performance and computational efficiency, but also improve the imaging quality while reducing the reconstruction time. Due to the pursuit of imaging quality and rate, the network structure of deep learning algorithms is becoming more and more complex. Although deep learning has shown promising results in the MRI field, they use quantization metrics as loss functions during the training process. However, in actual scenarios, the distribution of MRI images is often complex and difficult to describe. Summary of the Invention
[0005] The object of the present invention is to propose a fast magnetic resonance image reconstruction method based on a latent variable diffusion model for practical needs. The original k-space data is encoded into a low-dimensional latent space, and prior knowledge is generated through diffusion in it. Combining a full-frequency structure reconstruction model and a high-frequency detail refinement model, and through pre-training of the encoder and decoder, and latent k-space diffusion training, fast and high-quality MRI reconstruction is achieved. It can not only effectively reduce the computational amount and computational time required for diffusion model reconstruction, but also retain the structural details of the image, retain clear image edges and textures, and improve the quality of the reconstructed image.
[0006] To achieve the above object, the present invention adopts the following technical solutions.
[0007] A fast magnetic resonance image reconstruction method based on a latent variable diffusion model includes the following steps:
[0008] Step S1, obtaining an image data set;
[0009] The image data set is obtained based on the SIAT data set and the fastMRI data set. The SIAT data set and the fastMRI data set are subjected to equal-proportion scaling processing and data augmentation processing to obtain an image data set with consistent image size and sufficient quantity, and the image data set is further divided into a training data set and a validation data set;
[0010] Step S2, constructing a diffusion model network;
[0011] Construct a latent k-space diffusion model network that can capture and represent features and structures in the k-space. Use the training data set to train the constructed latent k-space diffusion model network, use the validation data set to verify the intermediate results of the training of the latent k-space diffusion model network, and obtain a trained latent k-space diffusion model network after verifying the fitting of the latent k-space diffusion model network;
[0012] Step S3, magnetic resonance image reconstruction;
[0013] Input the undersampled k-space data, and use the trained latent k-space diffusion model network to perform magnetic resonance image reconstruction, and output the reconstructed result image.
[0014] Specifically, in step S1, the SIAT data set and the fastMRI data set are subjected to equal-proportion scaling processing and data augmentation processing, and the processing process is as follows:
[0015] Merge 500 12-coil MRI images in the SIAT dataset into single-coil images, and then enhance them to 4000 single-coil images; randomly select 92 individuals from the fastMRI dataset, merge each slice of each individual into single-coil data, obtain a total of 1424 T1-weighted single-coil images, and then enhance them to 8000 single-coil images; uniformly crop the image sizes of the obtained 12000 single-coil images to 256×256 to obtain an image dataset with consistent image sizes and sufficient quantities.
[0016] Furthermore, in step S1, the image dataset is further divided into a training dataset and a validation dataset, and the ratio of the training dataset to the validation dataset is 9:1.
[0017] Specifically, in step S2, construct a latent k-space diffusion model network structure that can effectively capture and represent features and structures in the k-space. This latent k-space diffusion model network structure includes a frequency information generation model, a high-frequency detail refinement model, and an encoder-decoder network:
[0018] The frequency information generation model is used to reconstruct the entire image structure from undersampled latent k-space data, and the high-frequency detail refinement model is used to recover the high-frequency details lost during the process of reconstructing the image structure by the frequency information generation model;
[0019] In the encoder-decoder network, the encoder is stacked through residual blocks and linear layers and is used to compress k-space data into a low-dimensional latent k-space. The decoder combines dynamic Transformer blocks with a U-shaped structure. The dynamic Transformer blocks include multi-head transposed attention and a gated feed-forward network, and use the latent k-space as a dynamic modulation parameter to retain the details of the restored image.
[0020] Furthermore, the training process of the latent k-space diffusion model network in step S2 is as follows:
[0021] Step S21: Train the overall parameters of the latent k-space diffusion model network using the three channels of RGB images, and input the training dataset into the latent k-space diffusion model network;
[0022] Step S22: Train the encoder-decoder of the frequency information generation model and the high-frequency detail refinement model simultaneously;
[0023] For the encoder of the frequency information generation model , first multiply the high-quality k-space data by the low-quality k-space data through a weight matrix To modulate the value range of the entire k-space data, enabling the encoder to more accurately extract features from the k-space data, and then the weighted high-quality k-space data and the weighted low-quality k-space data are concatenated and downsampled using the PixelUnshuffle operation to obtain the input of the encoder of the frequency information generation model. The encoder encodes the input into a latent k-space representation. The expression for the entire training process is as follows:
[0024] ;
[0025] In the above formula, is the latent k-space representation; and are the weighted high-quality k-space data and the weighted low-quality k-space data respectively after multiplying the encoder input by weights;
[0026] For the encoder of the high-frequency detail refinement model, first use the same weight strategy to obtain and , and then multiply them by the high-frequency information extraction operator to obtain and , ensuring that the encoder pays more attention to the high-frequency details of the k-space data. Subsequently and are concatenated and downsampled to obtain the input of the encoder of the high-frequency detail refinement model. The encoder encodes the input into a high-frequency latent k-space representation. The expression for the entire training process is as follows:
[0027] ;
[0028] In the above formula, is the high-frequency latent k-space representation; and are respectively and multiplied by the high-frequency information extraction operator to obtain the weighted high-frequency k-space data;
[0029] The decoders and of the frequency information generation model and the high-frequency detail refinement model are trained to respectively restore and to k-space data and , and the formula is expressed as:
[0030] ;
[0031] ;
[0032] Step S23: Train the frequency information generation model and the high-frequency detail refinement model;
[0033] For the frequency information generation model, first, the weighted k-space data and are input into the pre-trained encoder , then the encoder encodes the input into the initial latent k-space representation , and subsequently, the forward DDPM converts the initial latent k-space representation into Gaussian noise through iteration . This process is described by the mathematical formula:
[0034] ;
[0035] In the above formula, represents converting the initial latent k-space representation into Gaussian noise through iteration ; is a Gaussian distribution; , , are hyperparameters for controlling the noise variance; I is the identity matrix;
[0036] For the high-frequency detail refinement model, use the encoder trained in step S22 to capture the high-frequency latent representation . Then, the forward DDPM diffuses the initial high-frequency latent k-space representation into Gaussian noise . This process is described by the mathematical formula:
[0037] ;
[0038] In the above formula, represents diffusing the initial high-frequency latent k-space representation into Gaussian noise through iteration T;
[0039] The objective functions for training the frequency information generation model and the high-frequency detail refinement model are respectively expressed as:
[0040] ;
[0041] ;
[0042] In the above formula, is the expectation over all variables, where is the noise sampled from a standard Gaussian distribution, is the high-frequency noise sampled from a standard Gaussian distribution, I is the identity matrix, t is the diffusion time step; and are the latent k-space representation at time step t and the high-frequency latent k-space representation at time step t respectively; and are the full-frequency data latent k-space representation and the high-frequency data latent k-space representation respectively; and are the parameterized diffusion models.
[0043] Furthermore, the training loss functions of the encoder-decoder corresponding to the frequency information generation model and the high-frequency detail refinement model are respectively expressed as:
[0044] ;
[0045] ;
[0046] In the above formula, and are the weighted k-space images of the frequency information generation model and the high-frequency detail refinement model respectively; and respectively represent the high-quality k-space data recovered by the encoders of the two models from the latent k-space; represents the norm .
[0047] Specifically, the reconstruction process of the magnetic resonance image in step S3 is as follows:
[0048] Step S31: Input the undersampled k-space data. The encoder encodes the k-space data therein into a latent k-space representation. The frequency information generation model reconstructs the entire image structure from the latent k-space. The reconstruction process is expressed as:
[0049] ;
[0050] In the above formula, is the latent k-space representation at time step t -1; is the latent k-space representation at time step t ; , , is a hyperparameter for controlling the noise variance;
[0051] Step S32: Recover the high-frequency details lost during the image structure reconstruction in step S31 through a high-frequency detail refinement model. The recovery process is expressed as:
[0052] ;
[0053] In the above formula, is the high-frequency latent k-space representation at time step t -1;
[0054] Step S33: Introduce a data consistency operator to align the data reconstructed by the frequency information generation model and the high-frequency detail refinement model with the original measurement results. The expression of the data consistency operator is:
[0055] ;
[0056] ;
[0057] In the above formula, ifj represents the position where the target sample data is collected; and respectively represent the index sets for collecting k-space samples and high-frequency collecting k-space samples; and represent the values of the k-space data generated by the frequency information generation model and the high-frequency detail refinement model at ; j represents a flag for identifying the position in the k-space; represents the value of the k-space data generated by the frequency information generation model at position j; represents the k-space data generated by the frequency information generation model; represents the value of the k-space data generated by the high-frequency detail refinement model at position j; represents the k-space data generated by the high-frequency detail refinement model; represents the value of the data generated by the frequency information generation model for the undersampled k-space at position j; represents the data generated by the frequency information generation model for the undersampled k-space ; represents the value of the data generated by the high-frequency detail refinement model for the undersampled k-space at position j; represents the data generated by the high-frequency detail refinement model for the undersampled k-space ; corresponds to the undersampled k-space; A weight coefficient representing the influence of balanced k-space measurement noise on reconstruction;
[0058] Step S34. After data consistency processing, the outputs of the frequency information generation model and the high-frequency detail refinement model are combined through a weighting strategy, and the mathematical expression of the weighting strategy is:
[0059] ;
[0060] By removing the weight the combined weighted data is converted into k-space data ; Then, the multi-coil image is obtained by applying the inverse Fourier transform, denoted as and the multi-coil images are combined by the sum of squares to obtain the final reconstruction result .
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] 1. The method of the present invention provides a potential k-space domain diffusion model and introduces a potential space diffusion strategy that combines frequency information generation and high-frequency refinement, enabling the model to operate in the potential k-space domain rather than the original k-space domain, significantly reducing the computational complexity and shortening the required reconstruction time.
[0063] 2. The method of the present invention constructs a frequency information generation model and a high-frequency detail refinement model by using a potential space diffusion strategy that combines frequency information generation and high-frequency refinement. The two models can effectively learn and utilize global and local information, which helps to improve the quality and accuracy of reconstruction.
[0064] 3. The method of the present invention strictly verifies and evaluates the reconstruction model on three datasets, and the experimental results prove its excellent reconstruction performance, demonstrating its potential as a robust and effective solution for MRI reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is a flowchart of a fast magnetic resonance image reconstruction method based on a latent variable diffusion model of the present invention;
[0066] Figure 2 is a network model diagram of the potential k-space diffusion model of the present invention;
[0067] Figure 3 is a diagram of the test data image reconstruction process;
[0068] Figure 4 is a diagram of the test image reconstruction result. DETAILED DESCRIPTION OF THE INVENTION
[0069] For the convenience of those of ordinary skill in the art to understand and implement the present invention, the following will detail each step of the method proposed by the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0070] Embodiment
[0071] As Figure 1 shown, the present invention discloses a fast magnetic resonance image reconstruction method based on a latent variable diffusion model, including the following steps:
[0072] Step S1, obtain an image data set;
[0073] The image data set is obtained based on the SIAT data set and the fastMRI data set. The SIAT data set and the fastMRI data set are subjected to equal-proportion scaling processing and data augmentation processing to obtain an image data set with consistent image sizes and sufficient quantities, and the image data set is further divided into a training data set and a validation data set;
[0074] Step S2, construct a diffusion model network;
[0075] Construct a latent k-space diffusion model network capable of capturing and representing features and structures in the k-space. Use the training data set to train the constructed latent k-space diffusion model network, use the validation data set to verify the intermediate results of the training of the latent k-space diffusion model network, and obtain a trained latent k-space diffusion model network after verifying the fitting of the latent k-space diffusion model network;
[0076] Step S3, magnetic resonance image reconstruction;
[0077] Input the undersampled k-space data, use the trained latent k-space diffusion model network for magnetic resonance image reconstruction, and output the reconstructed result image.
[0078] Specifically, in step S1, the equal-proportion scaling processing and data augmentation processing of the SIAT data set and the fastMRI data set are as follows:
[0079] Merge 500 12-coil MRI images in the SIAT dataset into single-coil images, and then enhance them to 4000 single-coil images; randomly select 92 individuals from the fastMRI dataset, merge each slice of each individual into single-coil data, obtain a total of 1424 T1-weighted single-coil images, and then enhance them to 8000 single-coil images; uniformly crop the image sizes of the obtained 12000 single-coil images to 256×256 to obtain an image dataset with consistent image sizes and sufficient quantities.
[0080] Further, in step S1, the image dataset is further divided into a training dataset and a validation dataset, and the ratio of the training dataset to the validation dataset is 9:1.
[0081] As Figure 2 shown, Figure 2 in (a) is the encoder-decoder network structure, Figure 2 in (b) is the latent k-space diffusion model network structure. In step S2, a latent k-space diffusion model network structure that can effectively capture and represent features and structures in the k-space is constructed. The latent k-space diffusion model network structure includes a frequency information generation model, a high-frequency detail refinement model, and an encoder-decoder network:
[0082] The frequency information generation model is used to reconstruct the entire image structure from the undersampled latent k-space data, and the high-frequency detail refinement model is used to recover the high-frequency details lost during the reconstruction of the image structure by the frequency information generation model;
[0083] In the encoder-decoder network, the encoder is stacked by residual blocks and linear layers to compress the k-space data into a low-dimensional latent k-space, and the decoder combines the dynamic Transformer block with the U-shaped structure. The dynamic Transformer block includes multi-head transposed attention and a gated feed-forward network, and uses the latent k-space as a dynamic modulation parameter to retain the details of the restored image.
[0084] Further, the training process of the latent k-space diffusion model network in step S2 is as follows:
[0085] Step S21: Train the overall parameters of the latent k-space diffusion model network using the three channels of the RGB image, and input the training dataset into the latent k-space diffusion model network;
[0086] Step S22: Train the encoder-decoder of the frequency information generation model and the high-frequency detail refinement model simultaneously;
[0087] For the encoder of the frequency information generation model , first, the high-quality k-space data Multiply with low-quality k-space data and modulate the value range of the entire k-space data through a weight matrix so that the encoder can more accurately extract features from the k-space data. Then, the weighted high-quality k-space data and the weighted low-quality k-space data are concatenated and downsampled using the PixelUnshuffle operation to obtain the input of the encoder for the encoder of the frequency information generation model which encodes the input into a latent k-space representation. The expression for the entire training process is as follows:
[0088] ;
[0089] In the above formula, is the latent k-space representation; and are the weighted high-quality k-space data and the weighted low-quality k-space data respectively after multiplying the encoder input by the weights;
[0090] For the encoder of the high-frequency detail refinement model , first obtain and using the same weight strategy, and then multiply them by the high-frequency information extraction operator to obtain and , ensuring that the encoder pays more attention to the high-frequency details of the k-space data. Subsequently and are concatenated and downsampled to obtain the input of the encoder for the high-frequency detail refinement model which encodes the input into a high-frequency latent k-space representation. The expression for the entire training process is as follows:
[0091] ;
[0092] In the above formula, is the high-frequency latent k-space representation; and are respectively and multiplied by the high-frequency information extraction operator to obtain the weighted high-frequency k-space data;
[0093] Train the decoders and of the frequency information generation model and the high-frequency detail refinement model, which respectively restore and to the k-space data and , which is expressed by the formula:
[0094] ;
[0095] ;
[0096] Step S23: Train the frequency information generation model and the high-frequency detail refinement model;
[0097] For the frequency information generation model, first, the weighted k-space data and are input into the pre-trained encoder , and then the encoder encodes the input into the initial latent k-space representation . Subsequently, the forward DDPM converts the initial latent k-space representation into Gaussian noise through iteration. This process is described by the mathematical formula:
[0098] ;
[0099] In the above formula, represents converting the initial latent k-space representation into Gaussian noise through iteration; is a Gaussian distribution; , , are hyperparameters for controlling the noise variance; I is the identity matrix;
[0100] For the high-frequency detail refinement model, the encoder trained in step S22 is used to capture the high-frequency latent representation . Then, the forward DDPM diffuses the initial high-frequency latent k-space representation into Gaussian noise . This process is described by the mathematical formula:
[0101] ;
[0102] In the above formula, represents diffusing the initial high-frequency latent k-space representation into Gaussian noise through iteration T;
[0103] The objective functions for training the frequency information generation model and the high-frequency detail refinement model are respectively expressed as:
[0104] ;
[0105] ;
[0106] In the above formula, is the expectation over all variables, where is the noise sampled from a standard Gaussian distribution, is the high-frequency noise sampled from a standard Gaussian distribution, I is the identity matrix, t is the diffusion time step; and are the latent k-space representations at time step t and the high-frequency latent k-space representations at time step t respectively; and are the full-frequency data latent k-space representation and the high-frequency data latent k-space representation respectively; and are the parameterized diffusion models.
[0107] Furthermore, the training loss functions of the encoder-decoder corresponding to the frequency information generation model and the high-frequency detail refinement model are respectively expressed as:
[0108] ;
[0109] ;
[0110] In the above formula, and are the weighted k-space images of the frequency information generation model and the high-frequency detail refinement model respectively; and respectively represent the high-quality k-space data recovered by the encoders of the two models from the latent k-space; represents the norm .
[0111] Specifically, as shown in Figure 3 , the reconstruction process of the magnetic resonance image in step S3 is as follows:
[0112] Step S31: Input the undersampled k-space data. The encoder encodes the k-space data therein into a latent k-space representation. The frequency information generation model reconstructs the entire image structure from the latent k-space. The reconstruction process is expressed as:
[0113] ;
[0114] In the above formula, is the latent k-space representation at time step t -1; is the potential k-space representation at time step t ; , , is the hyperparameter for controlling the noise variance;
[0115] Step S32, the high-frequency details lost during the image structure reconstruction in step S31 are restored by the high-frequency detail refinement model, and the restoration process is expressed as:
[0116] ;
[0117] In the above formula, is the high-frequency potential k-space representation at time step t -1;
[0118] Step S33, a data consistency operator is introduced to align the data reconstructed by the frequency information generation model and the high-frequency detail refinement model with the original measurement results. The expression of the data consistency operator is:
[0119] [[ID=Z8]] ;
[0120] ;
[0121] In the above formula, ifj represents the position where the target sample data is collected; and respectively represent the index sets for collecting k-space samples and high-frequency collecting k-space samples; and represent the values of the k-space data generated by the frequency information generation model and the high-frequency detail refinement model at ; j represents the identifier for the position in the k-space; represents the value of the k-space data generated by the frequency information generation model at position j; represents the k-space data generated by the frequency information generation model; represents the value of the k-space data generated by the high-frequency detail refinement model at position j; represents the k-space data generated by the high-frequency detail refinement model; represents the value of the data generated by the frequency information generation model for the undersampled k-space at position j; represents the data generated by the frequency information generation model for the undersampled k-space ; represents the value of the data generated by the high-frequency detail refinement model for the undersampled k-space at position j; represents the data generated by the high-frequency detail refinement model for the undersampled k-space Generated data; Corresponding to the undersampled k-space; Weighting coefficients representing the impact of balanced k-space measurement noise on reconstruction;
[0122] Step S34: After data consistency processing, combine the outputs of the frequency information generation model and the high-frequency detail refinement model through a weighting strategy, and the mathematical expression of the weighting strategy is:
[0123] ;
[0124] By removing the weights Convert the combined weighted data to k-space data ; Then obtain the multi-coil image by applying the inverse Fourier transform, denoted as , and combine the multi-coil images through the sum of squares to obtain the final reconstruction result .
[0125] The following gives a specific embodiment of using the method proposed by the present invention for image reconstruction to further illustrate the technical effects of the method of the present invention.
[0126] The experimental configuration is as follows:
[0127] In this embodiment, the method proposed by the present invention is implemented using a single NVIDIA RTX3090 GPU in PyTorch. For the decoder, a 5-level encoder-decoder structure is adopted in this embodiment. The number of Transformer blocks from level 1 to level 5 is set to [1, 1, 1, 1, 9], and note that the number of heads is set to [1, 2, 4, 8, 16]. For the latent k-space diffusion model network, the channel dimension is 64, and the predefined hyperparameter increases linearly from 0.1 to 0.99. During the training process, the batch size is set to 4 in this embodiment, and the Adam optimizer is used for network training, where , . During the reconstruction process, the total number of iteration steps is 4, and the batch size input to the model is dynamically determined by the coils of the multi-coil MRI data.
[0128] To demonstrate the clinical feasibility of the reconstruction effect, in this embodiment, the ability of the present invention to retain lesion areas in images is systematically evaluated by comparing with traditional reconstruction methods. In this embodiment, the method of the present invention is compared with two traditional deep learning algorithms, E2E-Varnet and MoDL. As Figure 4 shows the intuitive images of the reconstruction results of different methods, and the random undersampling factor R = 3, 4 and the Poisson undersampling factor R = 4, 6 are used in the image reconstruction process. AsFigure 4 It can be seen that the method of the present invention has a very great advantage in removing image artifacts. However, a relatively large amount of noise remains in the reconstruction result of the E2E-Varnet algorithm. Although the MoDL algorithm suppresses noise and artifacts to a certain extent, the reconstructed image is too smooth and loses too much texture detail. By comparing the three, it can be clearly concluded that the method of the present invention not only effectively suppresses artifacts but also preserves relatively obvious texture details, and is more capable of effectively retaining the lesion area of the reconstructed image than traditional reconstruction methods, achieving high-quality MRI reconstruction.
[0129] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. A fast magnetic resonance image reconstruction method based on a latent variable diffusion model, characterized in that It includes the following steps: Step S1, obtain an image dataset; The image dataset is obtained based on the SIAT dataset and the fastMRI dataset. The SIAT dataset and the fastMRI dataset are subjected to equal-proportion scaling processing and data augmentation processing to obtain an image dataset with consistent image sizes and sufficient quantities, and the image dataset is further divided into a training dataset and a validation dataset; Step S2, construct a diffusion model network; Construct a latent k-space diffusion model network capable of capturing and representing features and structures in the k-space. Use the training dataset to train the constructed latent k-space diffusion model network, use the validation dataset to verify the intermediate results of the training of the latent k-space diffusion model network, and obtain a trained latent k-space diffusion model network after verifying the fitting of the latent k-space diffusion model network; The structure of the latent k-space diffusion model network includes a frequency information generation model, a high-frequency detail refinement model, and an encoder-decoder network: The frequency information generation model is used to reconstruct the entire image structure from undersampled latent k-space data, and the high-frequency detail refinement model is used to recover the high-frequency details lost during the process of reconstructing the image structure by the frequency information generation model; In the encoder-decoder network, the encoder is stacked through residual blocks and linear layers and is used to compress k-space data into a low-dimensional latent k-space. The decoder combines a dynamic Transformer block and a U-shaped structure. The dynamic Transformer block includes multi-head transposed attention and a gated feed-forward network, and uses the latent k-space as a dynamic modulation parameter to retain the details of the restored image; Step S3, magnetic resonance image reconstruction; Input the undersampled k-space data, use the trained latent k-space diffusion model network for magnetic resonance image reconstruction, and output the reconstructed result image.
2. The rapid magnetic resonance image reconstruction method based on a latent variable diffusion model according to claim 1, wherein In step S1, the equal-proportion scaling processing and data augmentation processing of the SIAT dataset and the fastMRI dataset are as follows: Merge 500 12-coil MRI images in the SIAT dataset into single-coil images, and then enhance them to 4000 single-coil images; randomly select 92 individuals from the fastMRI dataset, merge each slice of each individual into single-coil data to obtain a total of 1424 T1-weighted single-coil images, and then enhance them to 8000 single-coil images; uniformly crop the image sizes of the obtained 12000 single-coil images to 256×256 to obtain an image dataset with consistent image sizes and sufficient quantities.
3. A fast magnetic resonance image reconstruction method based on a latent variable diffusion model according to claim 1, characterized in that, In step S1, the image dataset is further divided into a training dataset and a validation dataset, and the ratio of the training dataset to the validation dataset is 9:
1.
4. A fast magnetic resonance image reconstruction method based on a latent variable diffusion model according to claim 3, characterized in that, In step S2, the training process of the latent k-space diffusion model network is as follows: Step S21, train the overall parameters of the latent k-space diffusion model network using the three channels of an RGB image, and input the training dataset into the latent k-space diffusion model network; Step S22, train the encoder-decoder of the frequency information generation model and the high-frequency detail refinement model simultaneously; For the encoder of the frequency information generation model , first multiply the high-quality k-space data by the low-quality k-space data and modulate the value range of the entire k-space data through the weight matrix so that the encoder can more accurately extract features from the k-space data. Then, concatenate the weighted high-quality k-space data and the weighted low-quality k-space data and perform downsampling using the PixelUnshuffle operation to obtain the input of the encoder . The encoder of the frequency information generation model encodes the input into a latent k-space representation. The expression for the entire training process is as follows: ; In the above formula, is the potential k-space representation; and are the weighted high-quality k-space data and weighted low-quality k-space data after multiplying the encoder inputs by weights, respectively; For the encoder of the high-frequency detail refinement model , first obtain and using the same weight strategy, and then multiply them by the high-frequency information extraction operator to obtain and , ensuring that the encoder pays more attention to the high-frequency details of the k-space data. Subsequently, and are concatenated and downsampled to obtain the input of the encoder . The encoder of the high-frequency detail refinement model encodes the input into a high-frequency latent k-space representation. The expression for the entire training process is as follows: ; In the above formula, is the high-frequency latent k-space representation; and are respectively and multiplied by the high-frequency information extraction operator to obtain the weighted high-frequency k-space data; Decoder for training frequency information generation model and high-frequency detail refinement model and , respectively restore and to k-space data and , which is expressed by the formula as: ; ; Step S23: Train the frequency information generation model and the high-frequency detail refinement model; For the frequency information generation model, first, the weighted k-space data and are input into the pre-trained encoder . Then, the encoder encodes the input into an initial latent k-space representation . Subsequently, the forward DDPM iteratively converts the initial latent k-space representation into Gaussian noise . This process is described by the mathematical formula as follows: ; In the above formula, denotes that through iteration the initial latent k-space representation is transformed into Gaussian noise ; is a Gaussian distribution; , , are hyperparameters for controlling the noise variance; I is an identity matrix; For the high-frequency detail refinement model, use the encoder trained in step S22 to capture high-frequency latent representations , after that, the forward DDPM diffuses the initial high-frequency latent k-space representation into Gaussian noise , and this process is described by the mathematical formula as follows: ; In the above formula, represents the diffusion of the starting high-frequency latent k-space representation into Gaussian noise ; The objective functions for training the frequency information generation model and the high-frequency detail refinement model are respectively expressed as: ; ; In the above formula, is the expectation over all variables, where is the noise sampled from a standard Gaussian distribution, is the high-frequency noise sampled from a standard Gaussian distribution, I is the identity matrix, t is the diffusion time step; and are the latent k-space representations at time step t and the high-frequency latent k-space representations at time step t respectively; and are the full-frequency data latent k-space representation and the high-frequency data latent k-space representation respectively; and are the parameterized diffusion models.
5. A fast magnetic resonance image reconstruction method based on a latent variable diffusion model according to claim 4, wherein The training loss functions of the encoder-decoder corresponding to the frequency information generation model and the high-frequency detail refinement model are respectively expressed as: ; ; In the above formula, and are the weighted k-space images of the frequency information generation model and the high-frequency detail refinement model, respectively; and respectively represent the high-quality k-space data recovered by the encoders of the two models from the latent k-space; represents the norm .
6. A fast magnetic resonance image reconstruction method based on a latent variable diffusion model according to claim 4, characterized in that, The reconstruction process of the magnetic resonance image in Step S3 is as follows: Step S31: Input the undersampled k-space data. The encoder encodes the k-space data therein into a latent k-space representation. The frequency information generation model reconstructs the entire image structure from the latent k-space. The reconstruction process is expressed as: ; In the above formula, is the latent k-space representation at time step t -1; is the latent k-space representation at time step t ; , , are hyperparameters for controlling the noise variance. Step S32: The high-frequency detail refinement model restores the high-frequency details lost during the image structure reconstruction process in Step S31. The restoration process is expressed as: ; In the above formula, is the high-frequency latent k-space representation at time step t -1; Step S33: Introduce a data consistency operator to align the data reconstructed by the frequency information generation model and the high-frequency detail refinement model with the original measurement results. The expression of the data consistency operator is: ; ; In the above formula, ifj represents the position for collecting the target sample data; and respectively represent the index sets for collecting k-space samples and high-frequency collected k-space samples; and represent the values of the k-space data generated by the frequency information generation model and the high-frequency detail refinement model at ; j represents an identifier for the position in the k-space; represents the value of the k-space data generated by the frequency information generation model at position j; represents the k-space data generated by the frequency information generation model; represents the value of the k-space data generated by the high-frequency detail refinement model at position j; represents the k-space data generated by the high-frequency detail refinement model; represents the value of the data generated by the frequency information generation model for the undersampled k-space at position j; represents the data generated by the frequency information generation model for the undersampled k-space ; represents the value of the data generated by the high-frequency detail refinement model for the undersampled k-space at position j; represents the data generated by the high-frequency detail refinement model for the undersampled k-space ; corresponds to the undersampled k-space; represents the weight coefficient for balancing the influence of k-space measurement noise on reconstruction; Step S34: After data consistency processing, combine the outputs of the frequency information generation model and the high-frequency detail refinement model through a weighting strategy. The mathematical expression of the weighting strategy is: ; By removing the weights , the combined weighted data is converted into k-space data ; then a multi-coil image is obtained by applying the inverse Fourier transform, denoted as , and the multi-coil images are combined by the sum of squares to obtain the final reconstruction result .
Citation Information
Patent Citations
Nuclear magnetic resonance image super-resolution recovery method and model construction method
CN117611453A
Magnetic resonance image reconstruction with deep reinforcement learning
US20190172230A1