A medical magnetic resonance image reconstruction method and system

By using a potential diffusion prior framework guided by multimodal information, combined with inverse Fourier transform and modified flow reconstruction model, the MRI undersampling reconstruction process is optimized, solving the problems of MRI image quality and efficiency, and achieving high-quality and fast image reconstruction suitable for clinical diagnosis.

CN121353470BActive Publication Date: 2026-02-24SHENZHEN TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511905107.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-02-24
Estimated Expiration
2045-12-17

AI Technical Summary

Technical Problem

Existing MRI imaging techniques suffer from poor image quality, low data acquisition efficiency, and low reconstruction efficiency in undersampling reconstruction, making it particularly difficult to balance image quality and acquisition speed in clinical applications.

Method used

A multimodal information-guided latent diffusion prior framework is adopted. By combining inverse Fourier transform, latent space coding, text coding and modified flow reconstruction model with pixel loss, VGG loss, KL divergence loss and GAN loss, the image reconstruction process is optimized to generate high-quality MRI images.

Benefits of technology

It improves the quality and efficiency of MRI image reconstruction, reduces artifacts and blurring, enhances the visual realism and structural consistency of images, adapts to the diagnostic needs of different anatomical sites, and meets the needs of rapid clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353470B_ABST
    Figure CN121353470B_ABST
Patent Text Reader

Abstract

The application provides a medical magnetic resonance image reconstruction method and system, and belongs to the technical field of image reconstruction. The method comprises the following steps: acquiring sampling encoding and undersampled k-space data; processing the undersampled k-space data through inverse Fourier transform to convert the k-space data into an image domain and obtain undersampled image data with blur or artifacts; performing preliminary repair on the undersampled image data based on the sampling encoding and further performing data consistency operation on the image to obtain a rough image, wherein the rough image is used as an image condition vector after latent space encoding; acquiring text as a prompt word, inputting the text into a text encoder and encoding the text into a high-dimensional semantic embedding vector; and inputting the high-dimensional semantic embedding vector and the image condition vector into a reconstruction model based on a modified flow to predict a full-sampling MRI image and generate a final reconstructed MRI image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical magnetic resonance image reconstruction technology, and particularly relates to a medical magnetic resonance image reconstruction method and system. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Magnetic resonance imaging (MRI), a crucial non-invasive imaging technique in modern medicine, provides high-resolution soft tissue images, playing an irreplaceable role in disease diagnosis and treatment planning. However, unlike CT or X-rays that directly capture density images of the body's interior, MRI measures the response of atomic nuclei, primarily hydrogen protons, in a magnetic field and calculates images using complex mathematical transformations. This process is inherently sequential and requires a large number of data points, rather than being instantaneous. Therefore, it suffers from long data acquisition times, severely limiting its efficient application, especially in clinical settings dealing with large numbers of patients, where this drawback becomes increasingly apparent.

[0004] To address the time-consuming data acquisition problem in magnetic resonance imaging (MRI), numerous techniques have been proposed. Parallel imaging technology leverages the sensitivity distribution characteristics of multi-coil sensors to accelerate the acquisition process; the combination of k-space undersampling strategies and compressed sensing further improves acquisition efficiency. However, these acceleration techniques have significant drawbacks in practical applications. When the acceleration is too aggressive, aliasing inevitably occurs, leading to residual distortion, blurring, and reduced reconstruction fidelity in the reconstructed image, severely impacting image quality.

[0005] Furthermore, existing image reconstruction algorithms also have many shortcomings. Conventional reconstruction algorithms often require significant computational overhead, placing extremely high demands on hardware performance and greatly increasing equipment costs for hospitals. Simultaneously, error accumulation is inevitable during iterative reasoning, and the gap between the reconstructed image and the original image gradually widens with each iteration. Moreover, existing algorithms have limited control over reconstruction details, making it difficult for doctors to perform refined reconstructions of specific regions based on diagnostic needs.

[0006] Diffusion models have also been explored for application in MRI image reconstruction. However, when operating in the spatial domain, diffusion models incur significant computational costs, slowing down reconstruction and increasing hardware requirements.

[0007] While implicit diffusion models (LDMs) offer unique advantages, their implicit spatial compression leads to information loss, reducing the fidelity of reconstructed images. For example, in reconstructing fine brain structures, crucial details of neural connections may be lost. Furthermore, due to their iterative nature, LDMs accumulate additional errors along curved paths, further degrading the quality of the reconstructed images.

[0008] Therefore, in MRI imaging, traditional MRI requires sampling each location in k-space according to the Nyquist frequency to avoid aliasing artifacts. k-space is the Fourier transform domain of the image; however, this operation is time-consuming. Undersampling, on the other hand, intentionally excludes data from certain k-space locations to speed up acquisition. However, undersampling can lead to visual problems such as artifacts and noise in the sampled images. Given the artifact interference, poor image quality, and low data acquisition efficiency in MRI undersampling image reconstruction, current techniques for MRI undersampling reconstruction have failed to effectively balance reconstruction efficiency, image quality, and the convenience and accuracy of practical applications. Summary of the Invention

[0009] To overcome the shortcomings of the prior art, this invention provides a medical magnetic resonance image reconstruction method and system. Based on a multimodal information-guided latent diffusion prior framework, it solves the technical solution for MRI undersampling reconstruction, generating images with clinical realism and further improving the quality and efficiency of MRI image reconstruction.

[0010] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0011] In the first aspect, a method for medical magnetic resonance image reconstruction is disclosed, including:

[0012] Obtain undersampled k-space data;

[0013] The undersampled k-space data is processed by inverse Fourier transform, converting the k-space data into the image domain to obtain undersampled image data with blur or artifacts.

[0014] The undersampled image data is initially repaired based on latent space coding, and a data consistency operation is further applied to the image to obtain a coarse image. The coarse image is then used as an image condition vector after latent space coding.

[0015] The text used as prompt words is obtained, and the text is input into the text encoder and encoded into a high-dimensional semantic embedding vector.

[0016] High-dimensional semantic embedding vectors and image conditional vectors are input into a modified flow-based reconstruction model to predict fully sampled MRI images and generate the final reconstructed MRI images.

[0017] As a further technical solution, the text used as prompt words includes metadata information, including anatomical location, MRI sequence type, and scan parameters.

[0018] As a further technical solution, the image conditional vector is optimized using four loss functions, including: pixel loss, VGG loss, KL divergence loss, and GAN loss.

[0019] The pixel loss and VGG loss are used to reduce pixel-level differences;

[0020] KL divergence loss is used to optimize the uniformity and continuity of the latent space distribution;

[0021] The GAN loss is used to enhance the visual realism and texture of the reconstructed output.

[0022] As a further technical solution, the reconstruction model based on the modified flow includes:

[0023] MR-VAE encoders, which include encoders and decoders, are used to map a coarse image to a latent code.

[0024] The modified flow path predictor predicts the speed using image conditional latent coding variables and high-dimensional semantic embedding vectors; after training and optimization, the modified flow path predictor yields the reconstructed latent code.

[0025] The decoder receives the reconstruction latent code and decodes it back into the image domain to generate the final reconstructed MRI image.

[0026] As a further technical solution, the modified flow path predictor is trained with a target L, which is reconstructed to optimize the prediction of near-linear transmission trajectories.

[0027] Secondly, a medical magnetic resonance image reconstruction system is disclosed, comprising:

[0028] The data acquisition module is configured to acquire undersampled k-space data;

[0029] The data processing module is configured to process the undersampled k-space data through inverse Fourier transform, convert the k-space data into the image domain, and obtain undersampled image data with blur or artifacts.

[0030] The Sketcher module is configured to: perform preliminary repair on undersampled image data based on latent space coding and further apply data consistency operation to the image to obtain a rough image, which is then used as an image condition vector after latent space coding.

[0031] The text processing module is configured to: acquire the text used as prompt words, input the text into the text encoder and encode it into a high-dimensional semantic embedding vector;

[0032] The reconstruction module is configured to input high-dimensional semantic embedding vectors and image condition vectors into a modified flow-based reconstruction model to predict fully sampled MRI images and generate the final reconstructed MRI images.

[0033] The above one or more technical solutions have the following beneficial effects:

[0034] The technical solution of this invention is to initially solve the problem of aliasing artifacts and blurring in low-quality images obtained by converting k-space data into the image domain. The low-quality image is initially repaired to suppress aliasing artifacts and restore low-frequency structures such as organ contours, generating a relatively clean image condition that still has missing information. In order to ensure that the image condition is physically consistent with the original undersampled k-space acquisition data and avoid introducing false information, the system further applies a data consistency operation to the image to correct it, thereby ensuring the reliability of subsequent reconstruction.

[0035] The technical solution of this invention also introduces text prompts as additional conditional information to improve reconstruction performance. These texts are encoded into high-dimensional semantic embedding vectors by the CLIP text encoder and then injected into the diffusion U-Net through cross-attention. This provides semantic-level guidance for subsequent diffusion processes, enhancing the consistency between structure and semantics.

[0036] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0037] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0038] Figure 1 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation

[0039] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0040] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0041] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0042] Definitions:

[0043] Stable diffusion;

[0044] LDPM-V2: A diffusion-prior magnetic resonance reconstruction model for latent space;

[0045] LDMs: Implicit Diffusion Models;

[0046] MRI: Magnetic Resonance Imaging;

[0047] VGG: Visual Geometry Group Network;

[0048] VGG Loss: VGG loss;

[0049] KL loss: Kullback-Leibler Divergence Loss;

[0050] GAN: Generative Adversarial Network;

[0051] GAN Loss: GAN loss;

[0052] Adversarial Loss: Combating Loss;

[0053] MR-VAE: A variational autoencoder adapted for magnetic resonance, consisting of a symmetrical pair of encoders and decoders.

[0054] Example 1

[0055] See appendix Figure 1 As shown, this embodiment discloses a method for medical magnetic resonance image reconstruction, including:

[0056] Step 1: Simulate undersampling on real full-sample magnetic resonance data to obtain undersampled k-space data;

[0057] Step 2: Process the undersampled k-space data through inverse Fourier transform to convert the k-space data into the image domain, resulting in undersampled image data with blur or artifacts.

[0058] Step 3: Perform preliminary repair on the undersampled image data and apply data consistency operation to obtain a coarse image. Then, perform latent space encoding on the coarse image to obtain the image condition vector.

[0059] Step 4: Obtain the text used as prompt words. The text is input into the text encoder and encoded into a high-dimensional semantic embedding vector.

[0060] Step 5: Input the high-dimensional semantic embedding vector and image conditional vector into the modified flow-based reconstruction model to predict the full-sampled MRI image and generate the final reconstructed MRI image.

[0061] In the above steps, the latent space encoding is obtained through a variational autoencoder with stable diffusion, and the undersampled k-space data is obtained by simulating undersampling with real full-sampled magnetic resonance data.

[0062] In this implementation example, since the input to the latent space diffusion prior magnetic resonance reconstruction model LDPM-V2 is undersampled k-space data, an undersampling mask is used to record which data was actually acquired and which data was not sampled. Then, the k-space data is converted to the image domain through inverse Fourier transform (iFFT), resulting in a low-quality image with blur or artifacts, i.e., undersampled image data.

[0063] In one implementation example, the k-space data is converted to the image domain. Specifically, the image domain image is obtained by directly performing an inverse Fourier transform on the k-space data, i.e., the frequency domain data. The reason for the conversion is that the basic diffusion model also operates on the image domain data, and the data distribution characteristics of the image domain are more conducive to the learning of the diffusion model than the distribution of the k-space data.

[0064] In this implementation example, step three involves image guidance with the Sketcher module: to initially address undersampled image data. The problem of aliasing artifacts and blurring exists. This example innovatively proposes a Sketcher module based on the Transformer architecture to process undersampled image data. Preliminary repairs are performed to suppress aliasing artifacts and restore low-frequency structures such as organ contours, generating a relatively clean image that still contains some missing information. This is the rough image shown in the attached figure, denoted as [image description missing]. This serves as a "sketch" for subsequent AI reconstruction. To ensure image conditions... To maintain physical consistency with the original undersampled k-space data and avoid introducing false information, a data consistency operation (DC) is further applied to the image to correct it, thereby ensuring the reliability of subsequent reconstruction.

[0065] .

[0066] in, It is a binary matrix, used as a full sampling mask in the formula, with all positions set to 1, and is used to generate the complementary mask IM; Essentially a binary matrix, it serves as an undersampling mask in the formula: positions actually sampled in k-space are 1, and positions not sampled are 0. It is the original undersampled k-space data matrix, containing only the measurements actually acquired by the MRI scanner; This indicates that only the actual measurements in the k-space are retained; This indicates blurry, undersampled image data. Inputting the data into the LDPM-v2 network generates a complete, high-resolution predicted image, filling in areas where information is missing. Fourier transform represents the conversion of image data from the image domain to the frequency domain. The inverse Fourier transform represents the conversion of image data from the frequency domain to the image domain.

[0067] In this implementation example, step four involves text-guided prompts: To improve the controllability of the method built on the text-to-image model backbone, text prompts are introduced as additional conditional information to enhance reconstruction performance. These prompts typically contain metadata information such as anatomical location, MRI sequence type, and scan parameters. This text is encoded into high-dimensional semantic embedding vectors using the CLIP text encoder. Then, the image is injected into the diffusion U-Net via cross-attention, a layer within the stable diffusion model used to fuse text and images. Cross-attention provides semantic-level guidance for subsequent diffusion processes, enhancing structural and semantic consistency. The image conditions C processed by the Sketcher module and DC operations, along with the high-dimensional semantic embedding vector p obtained from the text prompt encoding, are input into the subsequent model as joint conditions.

[0068] .

[0069] Where p is a high-dimensional semantic embedding vector. τ It is a text encoder of the diffusion model, used to encode the input text prompt word P into the feature dimension, so that it can be fused with image features through cross attention.

[0070] Before entering the diffusion model, the image condition C is first processed by a variational autoencoder MR-VAE applicable to various MRI-related tasks and optimized using four loss functions. These four loss functions include: pixel loss and VGG loss, KL divergence loss, and GAN loss. Pixel loss and VGG loss are used to reduce pixel-level differences; KL divergence is also introduced to enhance the statistical consistency between the learned latent variables and predefined prior distributions (such as Gaussian distributions). Furthermore, GAN loss is added to enhance the visual realism and texture of the reconstructed output. The total loss can be expressed as:

[0071] .

[0072] in, This represents the reconstructed image output by a variational autoencoder (MR-VAE) adapted to magnetic resonance imaging. This represents a true, fully sampled image, used only during training. The VGG network acts as a feature extractor here. -VGG(x) means that two images are input into the VGG network respectively to extract their deep feature maps; KL represents Kullback-Leibler divergence, which is used to measure the difference between two probability distributions and force the latent space to be regularized; A Gaussian distribution representing the mean u and variance σ² is used to encode the latent features of the input image; Represents a standard normal distribution; It is a loss-countermeasure that uses a discriminator to play a game with the generator in order to improve the visual realism of the reconstructed image.

[0073] Here, parameters μ, ρ, ω, and λ are all hyperparameters, artificially set values ​​that need to be verified experimentally. Parameter μ represents the weighting coefficient used to control the pixel-level reconstruction loss, and is set to 1 in the experiment; ρ represents the weighting coefficient, adjusting the feature matching loss based on the VGG network, used to constrain the consistency of the generated image and the real image in deep semantic features, and is set to 0.2 in the experiment. ω is used to balance the KL divergence term, forcing the latent variable distribution N(u, To approximate the standard normal distribution N(0,1) to normalize the latent space, in the experiment, we take 1× λ is the resistance to loss The discriminator is used to improve the visual realism of the generated images; in the experiments, it was set to 0.65. Four weight parameters (μ, ρ, ω, λ) work together to coordinate the optimal balance between image reconstruction accuracy, feature fidelity, latent space regularization, and adversarial training. The total loss is the loss function used to train the variational autoencoder MR-VAE. The four weight parameters (μ, ρ, ω, λ) work together to coordinate the optimal balance between image reconstruction accuracy, feature fidelity, latent space regularization, and adversarial training.

[0074] Variational self-encoder adapted to MR-VAE magnetic resonance Finally, the image condition C is compressed and mapped to a low-dimensional latent space representation, denoted as c, thereby achieving efficient inference while preserving the key features of the image.

[0075] In this implementation example, in step five, the correction flow: using conditional guidance from the two modalities mentioned above—structural and semantic modalities—a correction flow-based reconstruction model is employed to accurately predict fully sampled MRI images. This is achieved by using an MR-VAE encoder. Map the conditional image C to the latent space representation c. High-quality reconstruction results can be obtained under robust multimodal control, thus enabling a forward Euler inference process that requires only one step:

[0076] .

[0077] in, Indicates the source latent encoding, from Mid-sampling; This represents the target latent code, i.e., the VAE code of the real MRI image. v represents the latent encoding of the MRI image predicted by the model, where v is the velocity predicted by the modified flow path predictor using the image conditional latent space representation c and the high-dimensional semantic embedding vector p.

[0078] .

[0079] in, Indicates the interpolation potential state: the midpoint on the ODE trajectory. Let c represent the time step, c be the image conditional latent space representation, and p be the high-dimensional semantic embedding vector. It is the network predicts the velocity field.

[0080] The path predictor takes as input the latent space representation c of the conditional image C encoded by MR-VAE and outputs the reconstructed magnetic resonance latent representation.

[0081] The training strategy is as follows: the modified flow path predictor is trained with a target L, which is reconstructed to optimize the prediction of near-linear transport trajectories. This optimization is performed under multimodal conditions p and c within the modified flow paradigm:

[0082] .

[0083] in, Indicates the time step; Indicates the source latent encoding, from Mid-sampling; This represents the target latent code, i.e., the VAE code of the real MRI image; Indicates the interpolation potential state: the midpoint on the ODE trajectory; This is the network predicting the velocity field; to efficiently optimize the model under flow targets, the upsampling layers of the Diffusion U-Net are unfrozen, and the linear transformation within the cross-attention module of the flow path predictor is fine-tuned using low-rank adaptive LoRA. After training and optimization, the following is obtained: This latent representation is fed into the MR-VAE decoder and decoded back into the image domain to generate the final reconstructed MRI image. Ultimately, this innovatively enabled the generation of high-quality, clear reconstructed images from undersampled k-space data.

[0084] See appendix again Figure 1 As shown, the working principle of the sub-technical solution in this embodiment is as follows:

[0085] (1) Image input: Perform raw data and basic transformation.

[0086] Undersampled k-space data is the original frequency domain data from magnetic resonance imaging. Because of the "undersampling," not all information was captured, and direct reconstruction would result in blurriness. M: Sampling matrix, defining the undersampled locations in k-space, used to constrain the reconstruction results to match the original undersampled data.

[0087] The basic transform is the Fourier transform, which converts the spatial domain image into frequency domain / k-space data. The inverse Fourier transform converts the k-space data back into a spatial domain image. The undersampled k-space data can be directly transformed into an initial blurred image through the inverse Fourier transform, serving as the starting point for subsequent processing.

[0088] (2) The Sketcher module (SkM) is used to generate preliminary image drafts.

[0089] The basic structure is recovered from the blurry initial image, generating a "draft" for subsequent fine-tuning.

[0090] Repair Module: Performs initial repair on the initial image using SwinIR, learning the image's contours and main structures.

[0091] Data Consistency: Ensures that the restored image matches the original undersampled data in k-space, mathematically satisfying... This is to avoid the reconstruction results deviating from the original data.

[0092] The output is a preliminary image draft, which retains the basic structure of the image, but the details are still blurry, serving as "structural anchor points" for subsequent fine reconstruction.

[0093] (3) MRControlNet module (MRCN) is used for fine reconstruction of high-quality images.

[0094] Guided by the initial draft, and combining two-stage sampling and encoding / decoding networks, high-quality images with rich details and conforming to the constraints of the original data are reconstructed.

[0095] Undersampled k-space data and sampling matrix M are used to reuse input data for final data consistency constraints.

[0096] The MR-VAE codec (Magnetic Resonance Variational Autoencoder) transforms an image draft matrix into low-dimensional latent features, extracts abstract structural information, and compresses data dimensionality.

[0097] For Decoder: It transforms the features of the latent space back into the spatial domain image, and finally generates the image from the optimized latent variables.

[0098] Dual-Stage Sampler (DSS): Optimizes latent variables in two steps to progressively approximate the true image distribution.

[0099] Initialization: Generate initial latent variables as the starting point for sampling.

[0100] Phase 1, the first p steps, is Inference Guidance.

[0101] Input: Current latent variables, draft latent features.

[0102] The sampling direction is guided by the Generator (a trainable control module) combined with draft features, and then noise is removed by the Frozen Denoiser (a pre-trained model) to optimize the latent variables.

[0103] Output: Optimized latent variables, ensuring consistency with the draft structure.

[0104] Phase 2, after Step 1: Data Consistency.

[0105] Input: Current latent variables, original undersampled k-space data matrix, binary matrix (for sampling).

[0106] Processing: Based on stage 1, add k-space data consistency constraints to ensure that the k-space of the decoded image matches the original undersampled k-space data matrix, and continue to optimize through the generator and denoiser.

[0107] Output: More accurate latent variables, while maintaining the fidelity of structure and original data.

[0108] Final Decoding:

[0109] Input: The final latent variables after T-step sampling optimization.

[0110] Output: The final image generated by the decoder.

[0111] (4) The final output is a high-quality reconstructed image. The high-quality reconstructed image is rich in detail and has a clear structure. It is consistent with the original undersampled data in k-space.

[0112] The principle of the diffusion model in this example is to denoise the latent code of arbitrary noise in order to predict the latent code of the desired output image. Therefore, the input of this method model should be two: one is the coarse image C as the condition for controlling the prediction; the other is the noisy latent code as the basis, on which the final target is obtained by denoising.

[0113] The image conditional vector c and the noise latent code z0 are concatenated and input together into the red trapezoidal network structure, the downsampling module of U-Net; the noise latent code z0 is then input separately into the blue trapezoidal network structure, the downsampling module of U-Net; the semantic vector p interacts with z0 through cross-attention layers (these layers are inserted inside the U-Net structure), as shown in the following formula:

[0114] .

[0115] Where Q is short for Query, which is the query vector used to initiate matching; K is short for Key, which is the key vector used to be matched; and V is short for Value, which is the value vector that is weighted by attention weights and outputs the actual content.

[0116] Cross-attention calculation formula:

[0117] .

[0118] in, These are the learnable projection weights in the cross-attention layer. yes The feature length of the vector is set to 64. T represents matrix transpose. Q is short for Query, which is the query vector used to initiate matching. K is short for Key, which is the key vector used to be matched. V is short for Value, which is the value vector that is weighted by attention weights and outputs the actual content.

[0119] Experimental evaluation and execution details:

[0120] Data Preparation: Experimental evaluation was conducted on the fastMRI brain dataset, which contains multi-coil FLAIR, T1, and T2-weighted MR acquisition data. To mitigate the impact of low-quality data, the last three slices for each subject were removed. The training set consisted of images from 1793 subjects, totaling 23275 slices, of which 2500 slices were retained for validation. For evaluation, this example used data from 72 subjects, yielding 934 slices. All multi-coil measurements were reconstructed into amplitude images using root mean square RSS combining. Textual cues associated with each subject were primarily derived from an h5 file with the key “ismrmrd_header”, containing various scan parameters and metadata, including field strength, fast spin echo with echo interval, TR, TE, TI, flip angle, etc.

[0121] Training details: The sketch module is trained on the corresponding dataset. The learning rate and batch size of 8 were used to train for 30 epochs. Both the MR-VAE and the prediction network were initialized from a pre-trained Stable Diffusion 2-1-base model. The MR-VAE was fine-tuned for 50 epochs with a learning rate of 8. Batch size is 8, loss weight , , ω = , The fine-tuned MR-VAE serves as a component of the corrected flow path predictor. During the 35 epochs of training the flow model, the VAE and the downsampling unit of the denoising U-Net remain frozen to ensure training stability, with a learning rate of [missing information]. The batch size was 4. All training used the AdamW optimizer, with input images of size 320×320.

[0122] The LDPM-V2 proposed in this invention exhibits several outstanding advantages in solving the problem of undersampling reconstruction in MRI:

[0123] Superior performance at high acceleration: Compared to the latent diffusion model (LDM) benchmark LDPM, LDPM-V2 shows a 0.38 dB improvement in peak signal-to-noise ratio (PSNR) at 8x acceleration and a 0.89 dB improvement at 10x acceleration. This demonstrates that LDPM-V2 effectively maintains the quality of reconstructed images under high acceleration conditions, reducing signal loss and noise interference, and providing doctors with clearer and more accurate images for diagnosis.

[0124] High detail fidelity: LDPM-V2 generates clear and sharp images while maintaining rich detail and fidelity. Based on LDPM-V2, missing frequency components can be inferred, resulting in sharper edges and higher structural fidelity. This effectively reduces over-smoothing issues, enhances visual realism, and provides doctors with more reliable image evidence for accurate diagnosis.

[0125] Short inference time: Compared with other state-of-the-art models, LDPM-V2 exhibits a significant advantage. The image domain diffusion method Score-MRI is computationally burdensome, requiring 471.831 seconds to generate a standard result for a single slice; the latent space method LDPM takes 7.314 seconds per slice. In contrast, LDPM-V2 achieves one-step inference, requiring only 0.446 seconds per slice, close to the runtime of SwinMR (0.375 seconds), while retaining the advantages of generative reconstruction, meeting the needs of rapid clinical diagnosis.

[0126] Significantly accelerated inference: LDPM-V2 achieved a PSNR of 29.9490, SSIM of 0.7928, and NMSE of 0.0246 at an 8x speedup. Even at a 10x speedup, its PSNR remained at 29.7298, SSIM at 0.7928, and NMSE at 0.0258, demonstrating robust performance and superiority over other models. It achieved better k-space recovery, significantly accelerated the inference process, and achieved iterative curved trajectory performance compared to traditional implicit diffusion methods, fully demonstrating its great potential for efficient MRI reconstruction.

[0127] Superior metrics across different anatomical structures: Compared with models such as UNet, E2E-VarNet, SwinMR, and LDPM, LDPM-V2 achieved the best PSNR and NMSE across different anatomical structures, and the second-best SSIM. This demonstrates that LDPM-V2 can achieve high-precision reconstruction of MRI data with different anatomical structures, accurately restoring the details and structural information of the images, providing high-quality image support for physicians to diagnose diseases in different locations.

[0128] Significant improvements under high acceleration conditions: The flow-based LDPM-V2 method consistently outperforms traditional LDM under both x8 and x10 acceleration conditions. At x8 acceleration, the PSNR reaches 30.5765, SSIM is 0.8135, and NMSE is 0.0230; at x10 acceleration, the PSNR is 30.1681, SSIM is 0.8093, and NMSE is 0.0252. This allows for more effective learning of the optimal denoising trajectory, and the one-step inference mechanism reduces error accumulation. This demonstrates that LDPM-V2 can still provide high-quality reconstructed images in clinical scenarios with high acceleration requirements, without affecting the physician's accurate judgment of the patient's condition.

[0129] In summary, LDPM-V2's superior performance in reconstruction quality, inference efficiency, and multimodal modulation surpasses traditional supervised methods and other diffusion-based methods, highlighting the enormous research and practical potential of flow cytometry-based generative models in the field of efficient and high-fidelity MRI reconstruction. This is of great significance for promoting the application of MRI imaging technology in clinical diagnosis.

[0130] Example 2

[0131] The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the above-described method.

[0132] Example 3

[0133] The purpose of this embodiment is to provide a computer-readable storage medium.

[0134] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the above method.

[0135] Example 4

[0136] The purpose of this embodiment is to provide a medical magnetic resonance image reconstruction system, including:

[0137] The data acquisition module is configured to acquire undersampled k-space data;

[0138] The data processing module is configured to process the undersampled k-space data through inverse Fourier transform, convert the k-space data into the image domain, and obtain undersampled image data with blur or artifacts.

[0139] The Sketcher module is configured to: perform preliminary repair on undersampled image data based on latent space coding and further apply data consistency operation to the image to obtain a rough image, which is then used as an image condition vector after latent space coding.

[0140] The text processing module is configured to: acquire the text used as prompt words, input the text into the text encoder and encode it into a high-dimensional semantic embedding vector;

[0141] The reconstruction module is configured to input high-dimensional semantic embedding vectors and image condition vectors into a modified flow-based reconstruction model to predict fully sampled MRI images and generate the final reconstructed MRI images.

[0142] Example 5

[0143] The purpose of this embodiment is to provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods and functions involved in any of the above embodiments.

[0144] The steps and methods involved in the apparatus of the above embodiments correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0145] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using program code executable by a measuring device, and thus can be stored in a storage device for execution by a computing device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0146] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for medical magnetic resonance image reconstruction, characterized in that, include: Obtain undersampled k-space data; The undersampled k-space data is processed by inverse Fourier transform, converting the k-space data into the image domain to obtain undersampled image data with blur or artifacts. Preliminary repair of undersampled image data is performed based on latent space coding, and data consistency operation is further applied to the image data to obtain a coarse image. The coarse image is then used as an image condition vector after latent space coding. The text used as prompt words is obtained, and the text is input into the text encoder and encoded into a high-dimensional semantic embedding vector. The image conditional vector and the noise latent code are concatenated and then input together into the downsampling module of the first U-Net; the noise latent code is then input separately into the downsampling module of the second U-Net; the high-dimensional semantic embedding vector interacts with the noise latent code through a cross-attention layer. High-dimensional semantic embedding vectors and image conditional vectors are input into a modified flow-based reconstruction model to predict fully sampled MRI images and generate the final reconstructed MRI images. The reconstruction model of the corrected flow includes: MR-VAE encoder, which includes an encoder and a decoder, is used to map a coarse image to a latent code; The modified flow path predictor predicts the speed using image conditional latent coding variables and high-dimensional semantic embedding vectors; after training and optimization, the modified flow path predictor yields the reconstructed latent code. The decoder receives the reconstruction latent code and decodes it back into the image domain to generate the final reconstructed MRI image.

2. The medical magnetic resonance image reconstruction method as described in claim 1, characterized in that, The text used as a prompt contains metadata information, including anatomical location, MRI sequence type, and scan parameters.

3. The medical magnetic resonance image reconstruction method as described in claim 1, characterized in that, The image conditional vector is optimized using four loss functions, including pixel loss, VGG loss, KL divergence loss, and GAN loss. The pixel loss and VGG loss are used to reduce pixel-level differences; KL divergence loss is used to optimize the uniformity and continuity of the latent space distribution; The GAN loss is used to enhance the visual realism and texture of the reconstructed output.

4. The medical magnetic resonance image reconstruction method as described in claim 1, characterized in that, The modified flow path predictor is trained with a target L, which is then reconstructed to optimize the prediction of near-linear transport trajectories.

5. A medical magnetic resonance image reconstruction system, characterized in that it comprises: The data acquisition module is configured to acquire undersampled k-space data; The data processing module is configured to process the undersampled k-space data through inverse Fourier transform, convert the k-space data into the image domain, and obtain undersampled image data with blur or artifacts. The Sketcher module is configured to: perform preliminary repair on undersampled image data based on latent space coding and further apply data consistency operation to the image data to obtain a coarse image, which is then used as an image condition vector after latent space coding. The text processing module is configured to: acquire the text used as prompt words, input the text into the text encoder and encode it into a high-dimensional semantic embedding vector; The image conditional vector and the noise latent code are concatenated and then input together into the downsampling module of the first U-Net; the noise latent code is then input separately into the downsampling module of the second U-Net; the high-dimensional semantic embedding vector interacts with the noise latent code through a cross-attention layer. The reconstruction module is configured to input high-dimensional semantic embedding vectors and image conditional vectors into a correction flow-based reconstruction model to predict fully sampled MRI images and generate the final reconstructed MRI images. The reconstruction model of the corrected flow includes: MR-VAE encoder, which includes an encoder and a decoder, is used to map a coarse image to a latent code; The modified flow path predictor predicts the speed using image conditional latent coding variables and high-dimensional semantic embedding vectors; after training and optimization, the modified flow path predictor yields the reconstructed latent code. The decoder receives the reconstruction latent code and decodes it back into the image domain to generate the final reconstructed MRI image.

6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.

7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1-4.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it performs the steps of the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Rapid nuclear magnetic resonance image reconstruction method and device

    CN114842103A

  • Prompt-based magnetic resonance image recovery method and system and application thereof

    CN117522721A