Self-supervised super-resolution image enhancement method and system based on conditional diffusion model

By employing a self-supervised super-resolution image enhancement method based on a conditional diffusion model, and utilizing T1-weighted images to provide prior information and a self-supervised data generator, the problem of improving the quality and resolution of ASL images is solved, and higher-quality ASL image reconstruction is achieved.

CN120598808BActive Publication Date: 2025-11-21ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511116013.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

In existing technologies, arterial spin-labeled magnetic resonance imaging (ASL) has limitations in improving image quality and resolution, especially low signal-to-noise ratio, low spatial resolution, and motion artifacts caused by long acquisition time, which limits the application of deep learning algorithms in ASL enhancement tasks.

Method used

A self-supervised super-resolution image enhancement method based on conditional diffusion model is adopted. By introducing a 3D U-Net model with sine and cosine coding and self-attention mechanism, combined with a two-stage accelerated sampling strategy, T1 weighted images are used to provide prior information and a self-supervised data generator to synthesize high- and low-resolution image sample pairs, and the conditional diffusion model is trained to improve the resolution of ASL images.

Benefits of technology

It achieves improved image super-resolution reconstruction speed and quality without sacrificing performance, enhances the resolution and detail preservation of ASL images, reduces image noise, and is suitable for clinical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598808B_ABST
    Figure CN120598808B_ABST
Patent Text Reader

Abstract

The application discloses a kind of self-supervision super-resolution image enhancement method and system based on conditional diffusion model.The method obtains T1 weighted image and first arterial spin labeling image collected by magnetic resonance equipment, and is up-sampled and registered, to obtain second arterial spin labeling image;Then based on the conditional diffusion model with 3D U-Net model as denoising network and pre-training, combined with double conditional sampling strategy and two-stage average acceleration sampling strategy, super-resolution arterial spin labeling image is generated.The application can improve the calculation speed of ASL image super-resolution reconstruction, and improve the generation quality and stability of diffusion model, and has obvious resolution and image quality improvement for low-resolution ASL.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of magnetic resonance technology, specifically relating to a self-supervised super-resolution image enhancement method and system based on a conditional diffusion model. Background Technology

[0002] Arterial spin labeling imaging (ASL) is a perfusion imaging technique widely used in vascular diseases due to its non-invasive measurement of cerebral blood flow. Cerebral blood flow (CBF) reflects the efficiency of nutrient and oxygen exchange between brain tissue and the circulatory system; therefore, abnormal perfusion is often considered a biomarker for various neurodegenerative diseases and tumors. In arterial spin labeling magnetic resonance imaging, blood in the internal carotid artery is labeled proximally in the imaging plane with radiofrequency pulses. These pulses reverse the magnetic moments of water molecules in the arterial blood, thus achieving labeling. The labeled image is acquired after a time delay, called the post-labeling delay (PLD). A control image is used to eliminate background MR signals; it is obtained by using the same sequence of pulses but without phase modulation of the labeled pulses, achieving an unlabeled effect. The difference between the labeled and control images is proportional to brain perfusion; this difference yields the perfusion-weighted image.

[0003] Although ASL non-invasive measurement of CBF has clinical advantages, its limitations are also very obvious. In arterial spin labeling magnetic resonance imaging, the labeled inverted protons decay at a rate of T1 (the T1 of blood is 1.6 seconds), and it takes a certain amount of time for the inverted protons to reach the image area after labeling. This results in a very limited time available for acquisition, and the labeled signal is also very weak (usually less than 1% of the original signal), with a relatively low signal-to-noise ratio (SNR). The spatial resolution of a single whole-brain ASL MRI is also very low [5], generally 4-8 mm. CBF quantification is compromised by low spatial resolution and motion artifacts because of significant partial volume effect (PVE), i.e., the voxel estimation of CBF parameters may include perfusion signals from other tissue types. In ASL acquisition, multiple control-label images are usually acquired for averaging to improve the signal-to-noise ratio, but long acquisition times can lead to motion artifacts and outliers. The limitations of low signal-to-noise ratio and spatial resolution, long acquisition time, and partial volume effect prevent ASL from being widely used in clinical practice.

[0004] Traditional enhancement methods offer limited improvement in ASL image quality; while they can reduce noise to some extent, they cannot improve image resolution. Deep learning-based algorithms face two main challenges: First, the model requires high-quality ASL images as target images, but due to limitations in ASL acquisition technology, obtaining high-quality, high-resolution ASL images is difficult. Second, there is a lack of sufficient low-resolution and high-resolution image pairs to train the super-resolution model; ASL imaging time is long, making it difficult to obtain a large number of low-resolution and high-resolution ASL scans corresponding to the same subject. These challenges hinder the improvement of deep learning algorithms in ASL enhancement tasks. Summary of the Invention

[0005] The purpose of this invention is to solve the above-mentioned problems in the prior art and to provide a self-supervised super-resolution image enhancement method and system based on a conditional diffusion model.

[0006] The specific technical solution adopted in this invention is as follows:

[0007] In a first aspect, the present invention provides a self-supervised super-resolution image enhancement method based on a conditional diffusion model, comprising:

[0008] S1. Acquire a T1-weighted image and a first arterial spin labeling image with a resolution lower than that of the T1-weighted image, and then upsample and register the first arterial spin labeling image onto the T1-weighted image to obtain a second arterial spin labeling image;

[0009] S2. Based on the pre-trained conditional diffusion model constructed from the sampling process and the 3D denoising network, the T1 weighted image is input into the conditional diffusion model, and a pseudo-random image is obtained by performing a forward diffusion process. Then, the pseudo-random image is used as the initial noise image and input into the conditional diffusion model to perform a sampling process according to the accelerated sampling strategy to obtain the generated image. In each step, the input of the denoising network is the current sampling step and the stitched image of the second arterial spin label image and the denoised image output from the previous step along the channel dimension. Then, the generated image is used as the coarse enhancement result and re-inputted into the conditional diffusion model. The partial sampling step at the end of the sampling process is performed with a strategy of not accelerating sampling, and the output is a super-resolution arterial spin label image with enhanced resolution and details.

[0010] As a preferred embodiment of the first aspect, the 3D denoising network in the conditional diffusion model adopts the 3D U-Net model, and in each convolutional layer of the encoder and decoder of the 3D U-Net model, the sampling step information is introduced into the feature map obtained by convolution through a sine and cosine coding mechanism. At the same time, a self-attention layer is added between the last convolutional layer of the encoder and the first upsampling layer of the decoder.

[0011] As a preferred embodiment of the first aspect above, the conditional diffusion model is pre-trained on data based on self-supervised generated image samples, and the method for synthesizing image sample pairs is as follows:

[0012] The original ASL image and T1-weighted image acquired through pairing are read. The original ASL image is upsampled to the same resolution as the T1-weighted image and then registered onto the T1-weighted image. Gray matter masks and white matter masks are segmented from the T1-weighted image and applied to the registered ASL image to weight the pixel values ​​of different regions. The weighted image is then smoothed using a low-pass filter to obtain a high-resolution ASL image. A Gaussian window is then used to truncate the k-space of the high-resolution ASL image to generate a low-resolution ASL image. Finally, noise is added to the low-resolution ASL image, and the noise-added low-resolution ASL image is paired with the high-resolution ASL image to form the image sample pairs required for training the conditional diffusion model.

[0013] More preferably, the method for adding noise to the low-resolution ASL image is as follows: the mean and variance of the noise are statistically calculated from the background region of the original ASL image, a Gaussian noise image that satisfies the mean and variance and has the same size as the original ASL image is generated, and the Gaussian noise image is upsampled and added to the low-resolution ASL image.

[0014] More preferably, when weighting the registered ASL image using two masks, the weighting coefficient within the gray matter mask range is set to 1, the weighting coefficient within the white matter mask range is set to 0.3, and the weighting coefficients within other ranges are set to 0; the low-pass filter uses a Gaussian kernel; the full width and half-height of the Gaussian window is one-quarter of the effective range of the k-space.

[0015] As a preferred embodiment of the first aspect mentioned above, when the stitched image is input into the conditional diffusion model, a coarse enhancement result is obtained based on random image block sampling and averaging. Specifically, the following steps are taken: according to multiple preset sliding window overlap rates, the input stitched image is first divided into blocks under each sliding window overlap rate, and then a reverse diffusion process is performed according to an accelerated sampling strategy. The generated images of each image block are then re-fused into a complete generated image. Finally, the complete generated images under different sliding window overlap rates are averaged, and the average image is used as the coarse enhancement result.

[0016] More preferably, the overlap rate of the plurality of sliding windows ranges from 0.1 to 0.5, and the proportion of the sampling steps at the end of the back diffusion process to the total sampling steps of the back diffusion process is 2% to 10%.

[0017] Secondly, the present invention provides a self-supervised super-resolution image enhancement system based on a conditional diffusion model, comprising:

[0018] The image data acquisition module is used to acquire a T1-weighted image and a first arterial spin labeling image with a resolution lower than that of the T1-weighted image acquired by the magnetic resonance imaging device, and to upsample and register the first arterial spin labeling image onto the T1-weighted image to obtain a second arterial spin labeling image;

[0019] The super-resolution image enhancement module is used to input the T1-weighted image into the conditional diffusion model based on a pre-trained sampling process and a 3D denoising network. A pseudo-random image is obtained by performing a forward diffusion process. The pseudo-random image is then used as the initial noise image and input into the conditional diffusion model to perform a sampling process according to an accelerated sampling strategy to obtain a generated image. In each step, the input of the denoising network is the current sampling step and the stitched image of the second arterial spin label image and the denoised image output from the previous step along the channel dimension. The generated image is then used as the coarse enhancement result and re-inputted into the conditional diffusion model. A partial sampling step at the end of the sampling process is performed with a strategy that does not accelerate sampling, and a super-resolution arterial spin label image with enhanced resolution and detail is output.

[0020] Thirdly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in any of the solutions of the first aspect above.

[0021] Fourthly, the present invention provides a computer electronic device, which includes a memory and a processor;

[0022] The memory is used to store computer programs;

[0023] The processor is configured to, when executing the computer program, implement the self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in any of the first aspects above.

[0024] Fifthly, the present invention provides a magnetic resonance imaging device, which includes a magnetic resonance scanner and a control unit, wherein the magnetic resonance scanner and the control unit are connected in communication.

[0025] The magnetic resonance scanner is used to acquire T1-weighted images and first arterial spin label images with a resolution lower than that of the T1-weighted images for magnetic resonance imaging targets, and transmit them to the control unit;

[0026] The control unit stores a computer program. When the computer program is executed, it can implement the self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in any of the first aspects above, and store the generated super-resolution arterial spin-labeled images according to a preset storage path for external devices to call or read.

[0027] Compared with the prior art, the present invention has the following advantages:

[0028] This invention designs a three-dimensional diffusion model that uses the 3DU-Net model, incorporating sine and cosine coding and a self-attention mechanism, as the denoising network. It also combines a two-stage accelerated sampling strategy. In the first stage, the accelerated sampling diffusion model and a multi-sample averaging strategy are used to quickly generate coarse results. In the second stage, a small number of steps are used to refine the resolution through sampling, achieving a 4x speedup and higher quality results. This invention can improve the computational speed of image super-resolution reconstruction without sacrificing performance or requiring retraining of the conditional diffusion model, while also enhancing the generation quality and stability of the diffusion model.

[0029] The three-dimensional diffusion model designed in this invention can employ a dual-conditional diffusion sampling strategy. On one hand, low-resolution ASL is used as a conditional variable to conditionally control the diffusion model and achieve super-resolution generation. On the other hand, pseudo-random images obtained by forward diffusion of T1-weighted images are used as generation conditions to provide high-resolution prior information for multiple modalities, thereby improving the super-resolution effect of ASL. This invention significantly improves the resolution of low-resolution ASL, reduces image SNR, and preserves detailed information such as cortical blood vessels as much as possible.

[0030] This invention designs a self-supervised data generator to synthesize ASL images. It combines intensity information from low-resolution ASL with structural information from T1w and introduces background noise from real ASL, enabling the synthesized ASL images to approximate real ASL scans. Using this self-supervised data generator, a large number of high-quality high- and low-resolution ASL image sample pairs can be synthesized, thereby effectively training the denoising network in the diffusion model. Attached Figure Description

[0031] Figure 1 This is a schematic diagram illustrating the steps of a self-supervised super-resolution image enhancement method based on a conditional diffusion model.

[0032] Figure 2 This is a detailed flowchart of the overall method in an embodiment of the present invention;

[0033] Figure 3 This is a framework diagram of a self-supervised super-resolution image enhancement method according to an embodiment of the present invention, wherein (A) is a self-supervised data generator; (B) is a three-dimensional conditional diffusion model; and (C) is a two-stage average accelerated sampling strategy.

[0034] Figure 4 This is a schematic diagram illustrating the principle of the self-supervised data generator according to an embodiment of the present invention;

[0035] Figure 5This is a schematic diagram of the framework of the two-stage average accelerated sampling strategy in an embodiment of the present invention;

[0036] Figure 6 This paper presents a visual comparison of the method of the present invention with other existing super-resolution and denoising methods on the ADNI dataset, where (a) shows the test results on a synthetic dataset; and (b) shows the results of different methods on the real ASL dataset of the ADNI dataset.

[0037] Figure 7 The present invention provides a visual comparison of the method of the present invention with other existing super-resolution and denoising methods on locally acquired low-resolution ASL images, wherein (a) shows the results of different methods on acquired low-resolution ASL images; (b) shows a visual comparison of different axes; and (c) shows the results of different methods on low-resolution ASL images obtained by downsampling acquired high-resolution ASL.

[0038] Figure 8 Visual comparison of the method of this invention with other methods on tumor ASL images of two pediatric subjects;

[0039] Figure 9 The experimental results compare the four strategies;

[0040] Figure 10 This is a schematic diagram of the module composition of a self-supervised super-resolution image enhancement system based on a conditional diffusion model.

[0041] Figure 11 This is a schematic diagram of the structure of a computer electronic device. Detailed Implementation

[0042] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in various embodiments of the present invention can be combined accordingly without mutual conflict.

[0043] In the description of this invention, it should be understood that the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated.

[0044] like Figure 1 As shown, in a preferred embodiment of the present invention, a self-supervised super-resolution image enhancement method based on a conditional diffusion model is provided, which includes the following steps S1 and S2. The specific implementation methods of the two steps are described below.

[0045] S1. Acquire a T1-weighted (T1w) image and a first arterial spin label (ASL) image with a resolution lower than that of the T1-weighted image from the magnetic resonance imaging device, and then upsample and register the first arterial spin label image onto the T1-weighted image to obtain a second arterial spin label image.

[0046] It should be noted that the acquisition of images by the MRI scanner can be achieved either offline or online. Offline acquisition involves reading image data from a storage device that already contains images acquired by the MRI scanner, while online acquisition involves driving the MRI scanner to acquire images in real time and retrieving the transmitted image data.

[0047] In this invention, due to time constraints in ASL acquisition, the directly acquired images are often low-resolution ASL images, with a resolution lower than that of T1-weighted images. These low-resolution ASL images require further super-resolution reconstruction via step S2 of this invention to improve their accuracy and image quality.

[0048] S2. Based on the pre-trained conditional diffusion model constructed from the sampling process and the 3D denoising network, the T1 weighted image is input into the conditional diffusion model, and a pseudo-random image is obtained by performing a forward diffusion process. Then, the pseudo-random image is used as the initial noise image and input into the conditional diffusion model to perform a sampling process according to the accelerated sampling strategy to obtain the generated image. In each step, the input of the denoising network is the current sampling step and the stitched image of the second arterial spin label image and the denoised image output from the previous step along the channel dimension. Then, the generated image is used as the coarse enhancement result and re-inputted into the conditional diffusion model. The partial sampling step at the end of the sampling process is performed with a strategy of not accelerating sampling, and the output is a super-resolution arterial spin label image with enhanced resolution and details.

[0049] It should be noted that the Conditional Diffusion Model is an existing technology; its core is to construct the diffusion process and a 3D denoising network. The diffusion process includes a forward diffusion process and a reverse sampling process (also known as a reverse diffusion process). The forward diffusion process is only used during training, while only the sampling process needs to be executed during the inference phase.

[0050] Traditional conditional diffusion models use a randomly sampled Gaussian noise image as initial input, which is then denoised step-by-step by a denoising network to obtain the target image. However, in this invention, a pseudo-random image obtained through forward diffusion of a T1-weighted image replaces the randomly sampled Gaussian noise image, thus providing an implicit prior condition for the denoising process. Furthermore, the conditional diffusion model of this invention is a 3D conditional diffusion model, and its 3D denoising network input also incorporates another implicit variable: a low-resolution second arterial spin label image. This second arterial spin label image is concatenated with the denoised image output from the previous sampling step before being input to the 3D denoising network. However, for the first sampling step, since there is no previous sampling step, the second arterial spin label image is concatenated with the aforementioned pseudo-random image before being input to the 3D denoising network. Additionally, the 3D denoising network has two inputs: besides the concatenated image, the other input is the currently executed sampling step t. Assuming the total number of forward diffusion steps is T, the sampling process starts from sampling step t=T and eventually reaches sampling step t=1 to obtain the target image.

[0051] In an embodiment of the present invention, a 3D denoising network A 3D U-Net model with sine and cosine coding and attention mechanisms is employed to process 3D input data and predict image noise at each sampling step. The 3D U-Net model is an existing technology, an extension of the traditional 2D U-Net. The 3D U-Net model consists of an encoder and a decoder; its specific structure can be found in the existing technical paper "3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation". In this 3D U-Net model, sine and cosine coding and attention mechanisms are introduced. Specifically, the encoder's initial input is the stitched image of the current sampling step. Simultaneously, in each convolutional layer of both the encoder and decoder, sine and cosine coding mechanisms are used to introduce sampling step information into the feature map obtained from the convolution. Furthermore, a self-attention layer is added between the last convolutional layer of the encoder and the first upsampling layer of the decoder. The sine-positional encoding mechanism introduces sampling step t into the output feature map of each convolutional layer. Sampling step t is converted into a high-dimensional embedding vector through sine-positional encoding, and then mapped by the MLP network to the same dimension as the convolutional layer's output feature map before being directly superimposed onto it. For each sampling step, the final output of the 3D U-Net model's decoder is the noise corresponding to that sampling step.

[0052] Additionally, it should be noted that the conditional diffusion model of this invention can adopt the diffusion process of the DDIM diffusion model. The DDIM diffusion model relaxes the Markov chain assumption in the traditional DDPM diffusion model, eliminating the need for step-by-step sampling. This allows for accelerated sampling strategies to improve image generation efficiency during the back-diffusion process. However, since accelerated sampling sacrifices some image quality, it is necessary to consider how to improve the quality of the generated images under the accelerated sampling strategy. Furthermore, 3D MRI data and the 3D UNet model impose significant computational costs on the training and inference of the diffusion model. When the hardware performance of the computer running the method of this invention is insufficient to directly input the complete stitched image into the conditional diffusion model for computation (especially due to GPU memory limitations), the original input stitched image can be divided into image patches, fed into the conditional diffusion model for computation, and then re-fused. Therefore, to reduce the impact of accelerated sampling strategies on image quality and lower the hardware performance requirements of the computer electronic equipment, this invention introduces random image block sampling and averaging to generate coarse enhancement results based on the stitched image. Specifically, the stitched image is no longer directly input into the conditional diffusion model for computation. Instead, based on multiple preset sliding window overlap rates (the preset overlap rates can be obtained by sampling within the range of 0.1 to 0.5, and the specific number of samples can be optimized through experiments), the input stitched image is first divided into blocks under each sliding window overlap rate. Each image block undergoes a back-diffusion process according to the accelerated sampling strategy, resulting in a generated image corresponding to each image block. The generated images corresponding to all image blocks under each sliding window overlap rate are then re-stitched and fused to form a complete generated image. An averaging operation is performed on the complete generated images under different sliding window overlap rates (i.e., averaging the values ​​of the same voxel in different images) to obtain an average image. This average image can then replace the generated image obtained by completely inputting the stitched image as the coarse enhancement result. Based on this approach, for the same image, the input image blocks for DDIM are not entirely identical each time, resulting in image block misalignment. This strategic image patch misalignment ensures that each image patch location has different contextual conditions during the diffusion process, thereby improving the quality of the generated results and reducing the performance requirements of the hardware devices.

[0053] Furthermore, the S2 step described above in this invention actually employs a two-stage averaging accelerated sampling strategy. In the first stage, multiple samples (in the form of image patches) are obtained from the same image. Each sample is accelerated using the DDIM model to obtain a generated result, and the results are then averaged to obtain a coarse enhancement result. In the second stage, the coarse result is fed into a conditional diffusion model based on the DDPM diffusion process, strictly following a stepwise sampling method, and executing the final step at the end of the back-diffusion process. The sampling process involves multiple sampling steps to obtain an enhanced and optimized super-resolution arterial spin-labeled image. The principle behind this two-stage averaging accelerated sampling strategy is that the coarse enhancement result may have some low resolution and blurred details due to the accelerated sampling strategy. Therefore, this coarse enhancement result is further input into the conditional diffusion model, performing a partial sampling step at the end of the backdiffusion process (i.e., a DDPM-based diffusion process) without accelerated sampling. Here, the partial sampling step at the end of the backdiffusion process refers to the last part of the sampling steps in the entire backdiffusion process, and this part accounts for 2% to 10% of the total number of sampling steps in the backdiffusion process. In a preferred embodiment, if the total number of sampling steps is 1000, the recommended number of this partial sampling step is 50 steps.

[0054] Furthermore, it should be noted that the aforementioned conditional diffusion model needs to be pre-trained on a dataset before being used for actual inference, so that it can meet the performance requirements of the high-resolution reconstruction task of ASL images. However, training the conditional diffusion model requires the collection of a large amount of sample data, while high-resolution ASL images are relatively scarce due to limitations in acquisition technology. Without sufficient high- and low-resolution ASL image sample data, the model cannot learn adequately. Therefore, this invention designs a self-supervised data generator (SSDG) to generate image sample pairs in batches. The conditional diffusion model is trained based on the image sample pairs generated by the self-supervised data generator. The method for synthesizing image sample pairs in the self-supervised data generator is as follows:

[0055] The process involves reading the original ASL image and T1-weighted image acquired by the MRI machine for the same object. The original ASL image has a lower resolution, while the T1-weighted image has a higher resolution. The original ASL image is upsampled to the same resolution as the T1-weighted image and then registered onto the T1-weighted image. Gray matter masks and white matter masks are segmented from the T1-weighted image and applied to the registered ASL image to weight the pixel values ​​of different regions. The weighted image is then smoothed to obtain the high-resolution ASL image. A Gaussian window is then used to truncate the k-space of the high-resolution ASL image to generate a low-resolution ASL image. Finally, the mean and variance of noise in the background region of the original ASL image are statistically analyzed to generate a Gaussian noise image that meets the mean and variance and has the same size as the original ASL image. This Gaussian noise image is upsampled and added to the low-resolution ASL image. The low-resolution ASL image with added noise is then paired with the high-resolution ASL image to form the image sample pairs required for training the conditional diffusion model.

[0056] It should be noted that applying gray matter and white matter masks to the registered ASL image to weight pixel values ​​in different regions involves assigning specified weighting coefficients to the white matter mask region, gray matter mask region, and non-gray matter white matter regions. Then, the voxel values ​​in the registered ASL image located within the white matter mask region, gray matter mask region, and other non-gray matter white matter regions are multiplied by their respective weighting coefficients. The weighted image serves as the initial high-resolution ASL image, which, after smoothing, becomes the high-resolution ASL image in the image sample pair. The weighting coefficients for the three regions can be optimized based on actual results. In this embodiment, when using two masks to weight the registered ASL image, the weighting coefficient within the gray matter mask range is set to 1, the weighting coefficient within the white matter mask range is set to 0.3, and the weighting coefficients in other ranges are set to 0. This weighting process can be implemented by directly multiplying the mask image values ​​onto the registered ASL image, or by mapping the mask image onto the registered ASL image and then performing weighted calculations on a voxel-by-voxel basis. The above smoothing operation can be implemented using a low-pass filter, and it is preferable to use a Gaussian kernel for smoothing.

[0057] In addition, the low-resolution ASL image of the present invention is not directly used as the original ASL image, but is obtained by truncating the k-space of the high-resolution ASL image with a Gaussian window. The full width and half height of the Gaussian window is one-quarter of the effective range of the k-space (i.e., FWHM = 1 / 4 k-space).

[0058] The following specific embodiment will demonstrate the implementation process and technical effects of the above-mentioned self-supervised super-resolution image enhancement method based on the conditional diffusion model.

[0059] Example

[0060] In this embodiment, the implementation flow of the self-supervised super-resolution image enhancement method based on the conditional diffusion model (referred to as the method of this invention) and the method for comparison are as follows: Figure 2 As shown. The implementation framework of the method of this invention is as follows. Figure 3 As shown below, each step of the process in this method will be described in detail.

[0061] 1. Construction of a self-supervised data generator and generation of image sample pairs:

[0062] This embodiment constructs a self-supervised data generator (SSDG) for obtaining synthetic ASL images. Traditional methods obtain simulated synthetic ASL images solely from structural MRI, resulting in the synthetic images' information being concentrated on the edge information of the structural MRI, leading to poor model generalization on real ASL. The self-supervised data generator designed in this invention, when generating image sample pairs, combines the intensity information of low-resolution ASL images and the structural information of T1w images, and adds background noise, thus more closely resembling actual ASL imaging.

[0063] like Figure 4 As shown, the self-supervised data generator constructs image sample pairs based on the ADNI dataset. The ADNI dataset contains original ASL images (denoted as Low-res ASL) and T1-weighted images (denoted as T1w) collected in pairs for different subjects. The original ASL images have lower resolution, while the T1w images have higher resolution. In the self-supervised data generator, the original ASL images from the ADNI dataset are first upsampled and registered with the corresponding T1w images of the same subject. Then, the SPM12 tool is used to segment the T1w images in the ADNI dataset into gray matter masks and white matter masks. Different weighting coefficients are set for the segmented mask regions: the weighting coefficient for non-gray matter white matter regions is 0, the weighting coefficient for gray matter mask regions is 1, and the weighting coefficient for white matter mask regions is 0.3, forming a mask image that covers the entire brain region but with different weighting coefficients for different regions. The coefficient setting for the white matter region is designed based on the fact that the signal intensity of the white matter part of the original ASL imaging is 1 / 3 of that of the gray matter. The masked image is then multiplied by the registered ASL image obtained on the same subject. This involves keeping the pixel values ​​in the gray matter regions unchanged, multiplying the pixel values ​​in the white matter regions by 0.3, and setting the pixel values ​​in the non-gray matter and white matter regions to 0, thus generating an initial high-resolution ASL image (denoted as ). Then use a Gaussian kernel (sigma=1) to... The image is slightly smoothed to obtain a synthesized high-resolution ASL image (denoted as ). Then, a k-space filter is introduced to degrade the k-space resolution, i.e., by using a Gaussian window to truncate the resolution. The initial low-resolution ASL image (denoted as ) is generated using the k-space of the image (FWHM = 1 / 4 k-space). ), which is equivalent to The simulated resolution of the image (1mm) was reduced to 4mm. This method of using K-space windowing can more closely approximate the actual acquisition process. Finally, the mean (Mean) and standard deviation (Std) of the noise were statistically obtained from the background region of the original ASL image. Based on this, a Gaussian noise image of the same size as the original ASL image was generated using Gaussian noise N(Mean,Std), and then upsampled to the same size. Images of the same size are superimposed on From the image, a composite low-resolution ASL image is formed (denoted as ). The final result will be a low-resolution ASL image. With high-resolution ASL images Pairing creates the image sample pairs needed for training the conditional diffusion model.

[0064] 2. Construction of the three-dimensional conditional diffusion model:

[0065] Through diffusion process and 3D denoising network To establish a three-dimensional conditional diffusion model, in which a denoising network is used. This is implemented using a 3D U-Net network with sine and cosine coding and attention mechanisms. During unconditional diffusion... Using a noise-damaged target y as input, the aim is to recover a clear, noise-free target. According to DDPM (J. Ho, J. Ho, AN Jain, A. Jain, P. Abbeel, and P. Abbeel, “Denoising diffusion probabilistic models.,” Neural Information Processing Systems, 2020), an unconditional forward diffusion process can be described by a Markov chain:

[0066]

[0067]

[0068] in: and The images are the diffused images at step t and step t-1, respectively. This is the initial image. The identity matrix has preset noise scheduling coefficients. , yes accumulation and , Under a constant variance, The noise obtained from the forward diffusion process is fitted using L2 loss. Since the objective of this invention is image enhancement, conditional variables need to be added to the diffusion process and the model. In the conditional diffusion model of this invention, the conditional sampling process (i.e., the reverse diffusion process) needs to start from step T and sample up to step 1, where step t (t=T,T-1,…,1) can be represented as:

[0069]

[0070] in .

[0071] During model training, yes , yes , yes The denoising result at the t-th diffusion step in the forward diffusion process. In the inference process, the traditional approach is to use the upsampled and registered ASL image and a random Gaussian noise image as input, and the enhanced high-resolution ASL image as output. However, this invention introduces a dual-conditional diffusion sampling strategy, replacing the random Gaussian noise with a pseudo-random image containing prior information about the T1w image, thereby reconstructing the 3D denoising network. encoder input That is, input of Defined as In this input In this process, the low-resolution ASL image that needs image enhancement (pre-upsampled and registered with the T1w image) is used as a conditional variable. Then, compare the denoising result with the result of the previous diffusion time step. The design principles and purpose of dual-condition diffusion sampling are described below, which involves splicing along the channel dimension.

[0072] In order to improve the super-resolution effect and make the generated results more stable, the sampling process of the conditional diffusion model is controlled by the following two conditions when introducing a dual conditional diffusion sampling strategy in this invention: The first part is to use the original acquired, upsampled and registered low-resolution ASL image as the conditional variable for each sampling step in the three-dimensional conditional diffusion model. The diffusion model is conditionally controlled to guide super-resolution generation. The second part involves prior implicit conditions, where a pseudo-random image is obtained by forward diffusion of the T1w image through all T steps (T is preferably 1000). This pseudo-random image replaces the original random Gaussian noise as the noise input for the 3D conditional diffusion model. However, it should be noted that this pseudo-random image is only used in the first sampling step. Input is performed at other time steps. It is the denoised image output from the previous sampling step.

[0073] Based on the principle of the diffusion process, the pseudo-random image obtained by forward diffusion of T1w is noise that closely approximates a Gaussian distribution. Compared to random Gaussian noise, it still contains some prior information from T1w. The high-resolution information of the ideal ASL image and the T1w image shares the same latent space. Therefore, using a pseudo-random image containing T1w prior information can add a high-resolution structural prior to the noise input while maintaining a randomness approximating a Gaussian distribution for sampling. Adjusting the sampling process through this implicit condition can yield better results. This invention tested the effects of using different noise images as latent variables: the first used a random Gaussian noise image; the second used a pseudo-random image obtained by forward diffusion of the ASL image; and the third used a pseudo-random image obtained by forward diffusion of the T1w image. Comparing the output results of the three approaches, it was demonstrated that the third approach, using a pseudo-random image obtained by forward diffusion of the T1w image as Gaussian noise, achieves the best ASL image enhancement effect.

[0074] 3. Two-stage averaging accelerated sampling strategy design:

[0075] Furthermore, 3D MRI data and the 3D UNet model impose significant computational costs on the training and inference of the diffusion model. Due to GPU memory limitations, the data needs to be divided into image patches before being fed into the model. Additionally, the generation results of the diffusion model exhibit some randomness. These two factors contribute to the long inference time and unstable results. Therefore, this invention constructs a two-stage averaging accelerated sampling strategy to obtain more stable results while accelerating the inference process. Its framework is as follows: Figure 5 As shown.

[0076] Phase 1: Based on the DDIM model (D. Ulyanov, A. Vedaldi, and V. Lempitsky, “Deepimage prior,” International Journal of Computer Vision, 2020, doi: 10.1007 / s11263-020-01303-4.), the sampling process is modified so that it can be directly used in pre-trained conditional diffusion models without retraining. The DDIM diffusion model relaxes the Markov chain assumption in DDPM, eliminating the need for step-by-step sampling. Its sampling process can be transformed into a shorter sequence {T, T − k, ..., 1}, where k is an acceleration factor. The sampling process is expressed as: where and The interval can be k iterations:

[0077]

[0078] Accelerated sampling strategies during the diffusion process sacrifice some diversity and image quality, and the faster the acceleration, the greater the image quality degradation. This invention uses a random image patch sampling and averaging method to partially address this problem. See also Figure 5 As shown, the pseudo-random image obtained by forward diffusion of the T1w image. and low-resolution ASL images The stitched image obtained along the channel is copied into N samples. Each sample is divided into image blocks using a specific sliding window overlap ratio. The divided image blocks are fed into the DDIM model to perform an accelerated sampling strategy to generate images. The generated images corresponding to all image blocks under each sliding window overlap ratio are then stitched and fused together to form a complete generated image. Finally, the complete generated images corresponding to the N samples are averaged to obtain the average image as the coarse enhancement result. For N samples, each corresponding to a different sliding window overlap rate can be uniformly sampled within the range of [0.1-0.5]. Based on this method, for the same image, the input image patches for DDIM are not exactly the same each time, but rather there will be misalignment of image patches. This strategic misalignment ensures that the position of each image patch is subject to different contextual conditions in the experiment, which can improve the quality of the generated results and is verified in experiments.

[0079] Second stage: The coarse enhancement results from the first stage are fed back into the DDPM-based conditional diffusion model (denoted as CMD), but only the final sampling process is performed. Step, its initial yes The final output is the optimized and enhanced super-resolution ASL image. This step aims to optimize resolution details and eliminate image blurring introduced by the first-stage averaging operation on multiple samples, thereby obtaining a higher resolution enhanced ASL result. The two-stage averaging accelerated sampling of this invention is performed only during the inference stage, without changing the training process of the 3D conditional diffusion model.

[0080] The hyperparameters designed in this embodiment for the three-dimensional conditional diffusion model are: N=5, k=20. =50, at which parameter, the speedup is 4 times compared to the original conditional diffusion model.

[0081] 4. Training of the three-dimensional conditional diffusion model:

[0082] In this embodiment, model training is implemented on a dataset of high- and low-resolution image samples synthesized by a self-supervised data generator (SSDG), while testing and validation are performed on a locally collected dataset. The training data is standardized using z-score before being input into the model and then linearly mapped to the [-1,1] interval.

[0083] In this embodiment, the 3D conditional diffusion model employs a patch-based training method. The original input image is pre-segmented into image patches of size [96, 96, 96], which are randomly selected from the training image (size [160, 200, 160]). During the training phase, random data augmentation operations such as rotation, shifting, and flipping are implemented. During the inference phase, sliding window prediction is used to obtain complete generated images. The sliding window is selected N times with different overlap rates, and complete generated images are generated separately and then averaged to improve the quality of the final generated image. The overlap rate between windows is 0.1~0.5. When the generated images of two adjacent image patches are fused and stitched, the overlapping area is weighted using a Gaussian score map method, which reduces the box artifacts caused by the sliding window by reducing the weight of the image patch edges.

[0084] In this embodiment, the denoising network 3D U-Net in the 3D conditional diffusion model adds a Time Embedding module and a self-attention mechanism module to the encoder-decoder structure. Time Embedding uses sine and cosine encoding, allowing 3D U-Net to obtain different outputs based on different input times, improving model performance. The self-attention mechanism module uses the multi-head attention mechanism from Transformer, splitting the feature map into tokens along the channel dimension, performing attention calculations, and then reassembling them back into a feature map. The 3D U-Net of this invention has a network depth of 4, corresponding to [32, 64, 128, 256] feature layers. Group Norm normalization is used throughout the 3D conditional diffusion model. During the diffusion process, the total number of diffusion steps is set to T=1000, and β0=1e-4.

[0085] In this embodiment, the Adam optimizer is used as the optimizer for training the 3D conditional diffusion model. The initial learning rate is 1e-4, and the L2 loss function is used for training. If the loss remains stable over 10 epochs, the learning rate is halved, with a lower bound of 10⁻⁷. The construction, training, and inference of the 3D conditional diffusion model are all implemented in PyTorch, and training is based on two RTX3090 GPUs.

[0086] To demonstrate the effectiveness of the method of the present invention, this embodiment introduces five algorithms—BM4D, NL-patch, DIP, EDSR3D, and UNet 3D—and compares them on the same dataset.

[0087] 5. Verification of experimental results:

[0088] 5.1 Results of synthesizing the dataset:

[0089] In this embodiment, a self-supervised data generator (SSDG) is used to generate a synthetic dataset consisting of 170 image sample pairs from the ADNI data, and the dataset is divided into a training set of 150 samples and a test set of 20 samples. Figure 6 The results of applying different methods to the synthetic dataset are shown. Although BM4D and DIP show some denoising effect, they reduce detail information compared to the original low-resolution synthetic ASL. NL-patch enhances structural edges using T1w images, but the overall image quality remains low. Deep learning-based UNet3D, EDSR3D, and the proposed diffusion model demonstrate excellent super-resolution performance, with the method of this invention producing the sharpest results. For example, in the synthetic image, the tortuosity of blood flow is noticeably sharp, showing a high similarity to the target image. In Table 1, the method of this invention consistently outperforms the other five methods in SSIM, PSNR, and MI, improving them by 12%, 33%, and 30% respectively compared to the NL-patch method.

[0090] Table 1. Comparative test results of different methods on synthetic datasets (20 cases)

[0091]

[0092] 5.2 Results of real ASL images in the ADNI dataset:

[0093] Since the ADNI dataset only includes low-resolution ASL images and lacks high-resolution target ASLs, the inception score (IS) and FID score are used to evaluate image quality. The FID target image is an ASL image with background noise removed. In Table 2, the diffusion model of this invention outperforms other methods in both maximum and average inception scores, indicating that the super-resolution ASL generated by the diffusion model is the clearest. Regarding the FID score, the method proposed in this invention lags behind EDSR3D and Unet3D, but outperforms other methods. This difference may be due to the target image still being a blurred low-resolution ASL, resulting in a generally low FID score. The sensitivity of FID to spatial blur and its two-dimensional design bias limit its effectiveness in evaluating three-dimensional ASLs. Nevertheless, the method of this invention... Figure 6The most significant enhancements are observed in the data. The performance of this invention on the ASL of the ADNI dataset is consistent with its performance on synthetic datasets, demonstrating that the designed self-supervised data generator can provide information close to the true ASL distribution for training deep learning models.

[0094] Table 2. IS and FID scores of different methods on low-resolution ASL images in the ADNI dataset.

[0095]

[0096] 5.3 Results of locally collected ASL dataset:

[0097] Figure 7 The results show the enhancements achieved by the method of this invention on locally acquired ASL data, compared to other methods. LR ASL and HR ASL represent locally acquired low-resolution and high-resolution ASL images, respectively. The trends of the different methods are consistent with those shown in the synthetic dataset. BM4D and NL-patch methods show effective denoising, but weak enhancement of high-resolution details. While the DIP method achieves slightly sharper edges, it also loses crucial contrast information in ASL. Deep learning-based ESR3D, UNet3D, and the diffusion model all show significant super-resolution effects, revealing clearer blood flow details in the ASL images. Among these, the diffusion model of this invention shows the best enhancement. These results demonstrate the superiority of the SSDG design, enabling the model to generalize across different datasets without relying on real high-resolution ASL. As shown in Table 3, the diffusion model of this invention outperforms the other five methods in both sets of metrics.

[0098] Table 3 Results of low-resolution ASL images acquired locally using different methods

[0099]

[0100] Within the locally acquired dataset, subtle differences in blood flow imaging were observed between the low-resolution and high-resolution ASL images, primarily due to subtle variations in brain perfusion at different acquisition periods. To mitigate the impact of these differences on the results, the high-resolution ASL was downsampled by truncating 1 / 4 of the k-space, resulting in low-resolution ASL images with blood flow imaging consistent with the original high-resolution ASL. Figure 7 Table (c) shows the results of the method of the present invention on low-resolution ASL. The super-resolution ASL is very similar in detail to the original high-resolution ASL, effectively recovering the original resolution information without introducing irrelevant structures. In Table 4, the method of the present invention achieves the highest values ​​on three metrics (SSIM, PSNR, and MI), indicating that the results are very close to the target high-resolution ASL.

[0101] Table 4. Results of low-resolution ASL obtained by different methods in downsampling high-resolution ASL.

[0102]

[0103] Furthermore, applying this invention to ASL images of pediatric tumors yielded the following results: Figure 8 As shown, compared to the structural distortion observed in UNet3D and EDSR 3D, the diffusion model-based method effectively preserves blood flow structures and contrast details near the tumor region. This result demonstrates the stability of the method, as it accurately preserves blood flow details from the original ASL image without distorting pathological details, thus avoiding interference with the diagnostic process.

[0104] This embodiment tests the effectiveness of the two-stage accelerated sampling strategy in this invention. To demonstrate this, four strategies were compared, and the experiment was conducted based on synthesized ASL high- and low-resolution image pairs. The four strategies are as follows:

[0105] Strategy (a) DDPM Only - Phase 2: Test the generation effect when the inference sampling is based solely on the DDPM method, with the values ​​T=[100,200,500,1000].

[0106] Strategy (b) DDIM-only acceleration: Acceleration sampling based only on DDIM, with acceleration factors set to [20, 10, 5, 2].

[0107] Strategy (c) DDIM + Random Sampling Averaging: Based on accelerated sampling using DDIM and random sample averaging, a coarse generation result is obtained using a random image patch sample averaging strategy. The parameter N is set to 50, and the acceleration factor of DDIM is [10, 20, 50].

[0108] Strategy (d) Complete Strategy: This means adopting the complete two-stage accelerated sampling strategy of this invention, including DDIM acceleration, random image patch sample averaging, and DDPM method refinement.

[0109] This embodiment plots all results in a line graph, with the horizontal axis representing the total number of sampling steps in each inference iteration and the vertical axis representing the graphical metrics of the generated results. Clearly, the ideal model should achieve optimal results with as few steps as possible; that is, the curve and points should be biased towards the upper left. Figure 9Four different strategies are presented. Strategy c performs the worst in SSIM, indicating that the refinement step is important for resolving the resolution. Strategy b uses only the DDIM method to accelerate sampling, which reduces computation time significantly, but at the cost of considerable performance loss. Strategy a uses only the refinement step and performs the worst in both PSNR and MI metrics, demonstrating the effectiveness of random sample averaging and DDIM acceleration. The proposed method, a two-stage accelerated sampling strategy, achieves a balance between computational speed and performance, proving its superiority.

[0110] For the dual-condition sampling strategy, this embodiment conducted a comparative study of different noise conditions, which was tested on synthesized ASL image pairs. The results are shown in Table 5. Group A is Gaussian-Noise: representing a random Gaussian noise image as the input noise image to the diffusion model. Group B is Asl-Latent variable: representing a pseudo-random image generated from an ASL image through forward diffusion (T=1000) as the input noise image to the diffusion model. Group C is T1w-Latent variable: representing a pseudo-random image generated from a T1w image through forward diffusion as the input noise image to the diffusion model. The results show that T1w-Latent variable performs best because the T1w diffused image, as the noise input, provides prior information about the MRI structure, providing constraints to resolve structural ambiguity in ASL super-resolution. This indicates that the latent variable of T1w forward diffusion provides potential high-resolution information, which can guide the diffusion model to generate images with better super-resolution performance.

[0111] Table 5 Comparison results under different noise conditions

[0112]

[0113] This embodiment separately tested the effectiveness of random image patch sampling and averaging using real ASL images. For comparison, this embodiment included a control group using direct averaging, the difference being that all overlap rates were set to 0.25. Both groups used the same number of samples N for averaging, N=5. The results are shown in Table 6. It can be seen that the random image patch sampling and averaging method outperforms the direct averaging method in all three metrics, fully demonstrating the effectiveness of this strategy.

[0114] Table 6. Comparison of results using different sample averaging strategies in the two-stage accelerated sampling strategy.

[0115]

[0116] In addition, it should be noted that the method steps shown in S1~S2 above can essentially be implemented in the form of computer programs or software functional modules.

[0117] Therefore, based on the same inventive concept, such as Figure 10 As shown, the present invention also provides a self-supervised super-resolution image enhancement system based on a conditional diffusion model, corresponding to the self-supervised super-resolution image enhancement method based on a conditional diffusion model provided in the above embodiments, which includes:

[0118] The image data acquisition module is used to acquire a T1-weighted image and a first arterial spin labeling image with a resolution lower than that of the T1-weighted image acquired by the magnetic resonance imaging device, and to upsample and register the first arterial spin labeling image onto the T1-weighted image to obtain a second arterial spin labeling image;

[0119] The super-resolution image enhancement module is used to input the T1-weighted image into the conditional diffusion model based on a pre-trained sampling process and a 3D denoising network. A pseudo-random image is obtained by performing a forward diffusion process. The pseudo-random image is then used as the initial noise image and input into the conditional diffusion model to perform a sampling process according to an accelerated sampling strategy to obtain a generated image. In each step, the input of the denoising network is the current sampling step and the stitched image of the second arterial spin label image and the denoised image output from the previous step along the channel dimension. The generated image is then used as the coarse enhancement result and re-inputted into the conditional diffusion model. A partial sampling step at the end of the sampling process is performed with a strategy that does not accelerate sampling, and a super-resolution arterial spin label image with enhanced resolution and detail is output.

[0120] Furthermore, based on the same inventive concept, such as Figure 11 As shown, this invention also provides a magnetic resonance imaging device corresponding to the self-supervised super-resolution image enhancement method based on a conditional diffusion model provided in the above embodiments. The device includes a magnetic resonance scanner and a control unit, which are connected in a communication manner to transmit data and control commands to each other. The communication connection can be established via a wired network or a wireless network; considering the reliability of data transmission, a wired network is preferred.

[0121] The aforementioned magnetic resonance scanner is used to acquire T1-weighted images and first arterial spin label images with a resolution lower than that of the T1-weighted images for magnetic resonance imaging targets, and transmit them to the control unit;

[0122] The aforementioned control unit stores a computer program. When the computer program is executed, it can implement the self-supervised super-resolution image enhancement method based on the conditional diffusion model as described above, and store the generated super-resolution arterial spin-labeled images according to a preset storage path for external devices to call or read.

[0123] It should be noted that the magnetic resonance imaging (MRI) device can be any MRI scanner capable of implementing parallel imaging methods. Its structure is existing technology, and mature commercial products can be used; the specific model is not limited. Furthermore, in addition to storing the aforementioned computer programs, the control unit of the MRI device should also contain the imaging sequences and other software programs necessary for implementing MRI.

[0124] Of course, the aforementioned control unit can be a standalone control unit or a control unit integrated into the magnetic resonance scanner. That is, the self-supervised super-resolution image enhancement method based on the conditional diffusion model can be integrated into the control unit of the magnetic resonance imaging device as a data processing program, allowing the magnetic resonance scanner to directly output the reconstruction results without the need for an additional control unit. The control unit of the magnetic resonance imaging device can be in the form of a host computer, which needs to include memory and a processor.

[0125] It should be noted that after the generated super-resolution arterial spin-labeled images are stored according to a preset storage path, they can be accessed or read by external devices according to the corresponding permissions. Such external devices can be local display devices, or servers, cloud platforms, mobile devices, etc., that remotely access and read image data.

[0126] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0127] Therefore, based on the same inventive concept, this invention provides a computer-readable storage medium corresponding to the self-supervised super-resolution image enhancement method based on the conditional diffusion model. The storage medium stores a computer program, which, when executed by a processor, can realize the self-supervised super-resolution image enhancement method based on the conditional diffusion model as described above.

[0128] Based on the same inventive concept, this invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the self-supervised super-resolution image enhancement method based on the conditional diffusion model as described above.

[0129] Specifically, in the computer-readable storage medium of the above embodiments, the stored computer program is executed by a processor, which can perform the aforementioned steps S1 to S2.

[0130] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.

[0131] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0132] It should also be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. In the embodiments provided in this application, the division of steps or modules in the system and method is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules or steps may be combined or integrated together, and a module or step may also be split.

[0133] The embodiments described above are merely some preferred implementations of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.

Claims

1. A self-supervised super-resolution image enhancement method based on a conditional diffusion model, characterized in that, include: S1. Acquire a T1-weighted image and a first arterial spin labeling image with a resolution lower than that of the T1-weighted image, and then upsample and register the first arterial spin labeling image onto the T1-weighted image to obtain a second arterial spin labeling image; S2. Based on the pre-trained conditional diffusion model constructed from the sampling process and the 3D denoising network, the T1 weighted image is input into the conditional diffusion model, and a pseudo-random image is obtained by performing a forward diffusion process. Then, a pseudo-random image is used as the initial noisy image and input into the conditional diffusion model to perform a sampling process according to an accelerated sampling strategy to obtain the generated image. In each step, the input of the denoising network is a stitched image of the second arterial spin label image and the denoised image output from the previous step along the channel dimension. The generated image is then used as a coarse enhancement result and re-inputted into the conditional diffusion model. A partial sampling step at the end of the sampling process is performed with a strategy of not accelerating sampling, and the output resolution and detail are enhanced and optimized into a super-resolution arterial spin label image. The 3D denoising network in the conditional diffusion model adopts the 3D U-Net model. In each convolutional layer of the encoder and decoder of the 3D U-Net model, the sampling step information is introduced into the feature map obtained by convolution through the sine and cosine coding mechanism. At the same time, a self-attention layer is added between the last convolutional layer of the encoder and the first upsampling layer of the decoder.

2. The self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in claim 1, characterized in that, The conditional diffusion model is pre-trained on data based on self-supervised generated image samples. The method for synthesizing image sample pairs is as follows: The original ASL image and T1-weighted image acquired through pairing are read. The original ASL image is upsampled to the same resolution as the T1-weighted image and then registered onto the T1-weighted image. Gray matter masks and white matter masks are segmented from the T1-weighted image and applied to the registered ASL image to weight the pixel values ​​of different regions. The weighted image is then smoothed using a low-pass filter to obtain a high-resolution ASL image. A Gaussian window is then used to truncate the k-space of the high-resolution ASL image to generate a low-resolution ASL image. Finally, noise is added to the low-resolution ASL image, and the noise-added low-resolution ASL image is paired with the high-resolution ASL image to form the image sample pairs required for training the conditional diffusion model.

3. The self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in claim 2, characterized in that, The method for adding noise to a low-resolution ASL image is as follows: the mean and variance of the noise are statistically calculated from the background region of the original ASL image, a Gaussian noise image that satisfies the mean and variance and has the same size as the original ASL image is generated, and the Gaussian noise image is upsampled and added to the low-resolution ASL image.

4. The self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in claim 2, characterized in that, When weighting the registered ASL image using two masks, the weighting coefficients within the gray matter mask range are set to 1, the weighting coefficients within the white matter mask range are set to 0.3, and the weighting coefficients within other ranges are set to 0; the low-pass filter uses a Gaussian kernel; the full width and half-height of the Gaussian window is one-quarter of the effective range of the k-space.

5. The self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in claim 1, characterized in that, When the stitched image is input into the conditional diffusion model, a coarse enhancement result is obtained based on random image block sampling and averaging. Specifically, the following steps are taken: according to multiple preset sliding window overlap rates, the input stitched image is first divided into blocks under each sliding window overlap rate, and then the sampling process is performed according to the accelerated sampling strategy. The generated images of each image block are then re-fused into a complete generated image. Finally, the complete generated images under different sliding window overlap rates are averaged, and the average image is used as the coarse enhancement result.

6. The self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in claim 5, characterized in that, The overlap rate of the multiple sliding windows ranges from 0.1 to 0.5, and the proportion of the sampling steps at the end of the sampling process to all sampling steps in the sampling process is 2% to 10%.

7. A self-supervised super-resolution image enhancement system based on a conditional diffusion model, characterized in that, include: The image data acquisition module is used to acquire a T1-weighted image and a first arterial spin labeling image with a resolution lower than that of the T1-weighted image acquired by the magnetic resonance imaging device, and to upsample and register the first arterial spin labeling image onto the T1-weighted image to obtain a second arterial spin labeling image; The super-resolution image enhancement module is used to input the T1-weighted image into the conditional diffusion model based on a pre-trained conditional diffusion model constructed from a sampling process and a 3D denoising network, and obtain a pseudo-random image by performing a forward diffusion process. Then, a pseudo-random image is used as the initial noisy image and input into the conditional diffusion model to perform a sampling process according to an accelerated sampling strategy to obtain the generated image. In each step, the input of the denoising network is the current sampling step and the stitched image of the second arterial spin label image and the denoised image output from the previous step along the channel dimension. The generated image is then used as the coarse enhancement result and re-inputted into the conditional diffusion model. A partial sampling step at the end of the sampling process is performed with a strategy of not accelerating sampling, and the output is a super-resolution arterial spin label image with enhanced resolution and detail. The 3D denoising network in the conditional diffusion model adopts the 3D U-Net model. In each convolutional layer of the encoder and decoder of the 3D U-Net model, the sampling step information is introduced into the feature map obtained by convolution through the sine and cosine coding mechanism. At the same time, a self-attention layer is added between the last convolutional layer of the encoder and the first upsampling layer of the decoder.

8. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in any one of claims 1 to 6.

9. A magnetic resonance imaging device, characterized in that, It includes a magnetic resonance scanner and a control unit, which are connected in communication. The magnetic resonance scanner is used to acquire T1-weighted images and first arterial spin label images with a resolution lower than that of the T1-weighted images for magnetic resonance imaging targets, and transmit them to the control unit; The control unit stores a computer program. When the computer program is executed, it can implement the self-supervised super-resolution image enhancement method based on the conditional diffusion model as described in any one of claims 1 to 6, and store the generated super-resolution arterial spin-labeled images according to a preset storage path for external devices to call or read.

Citation Information

Patent Citations

  • Polarization image super-resolution method based on conditional diffusion model

    CN119648533A

  • Zero-shot low-dose CT image denoising method and apparatus based on strip diffusion model

    US12333686B1