Arbitrary-scale heart cine MIR super-resolution reconstruction method and system

Through the implicit neural representation method based on diffusion priors and the implicit attention mechanism in space-time, implicit attention mechanism of space-time, the problems of fixed scale, insufficient frequency domain information, motion artifacts and high computing overhead in cardiac cine MIR super-resolution reconstruction are solved, and the super-resolution reconstruction of cardiac cine MRI at any scale is realized, improving image quality and reconstruction efficiency.

CN120070184AActive Publication Date: 2025-05-30YANTAI UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510219392.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing cardiac cine MIR super-resolution reconstruction methods have limitations in fixed scales, limitations in spatial resolution caused by insufficient frequency domain information, the impact of motion artifacts on image quality, and high computational overhead in the generation process of diffusion model.

Method used

Using the implicit neural representation method based on diffusion prior (DP-INR), the implicit neural network and the implicit attention mechanism of space-time is used to realize super-resolution reconstruction of cardiac cine MRI at any scale. The method includes an encoder extracting features of the boundary reference frame, a denoising diffusion model processes the intermediate frame to obtain diffusion prior features, and reconstructing in combination with spatial and temporal perception implicit attention.

Benefits of technology

Super-resolution reconstruction at any scale is realized, which improves the spatial resolution and clarity of the image, reduces the impact of motion artifacts, and significantly shortens the reconstruction time, improving the universality and operation efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070184A_ABST
    Figure CN120070184A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical image processing. The invention provides an arbitrary-scale heart cine MIR super-resolution reconstruction method and system, and the method comprises the steps: taking a diffusion prior feature and a key target frame feature as the input of spatial perception implicit attention; taking the diffusion prior feature, the key target frame feature and the output of the spatial perception implicit attention as the input of the time perception implicit attention, and processing the output of the time perception implicit attention by adopting a decoder to obtain a super-resolution reconstruction result corresponding to 2N + 1 continuous video frame sequences; according to the method, the problem of fixed-scale reconstruction in an existing super-resolution method, the problem of spatial resolution limitation caused by insufficient frequency domain information, the problem of influence of motion artifacts and the problem of high calculation overhead caused by multiple iteration steps in the diffusion model generation process are solved, and more accurate super-resolution reconstruction of heart cine MRI is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a method for super-resolution reconstruction of cardiac cine MIR at any scale, a system for super-resolution reconstruction of cardiac cine MIR at any scale, a computer device, a computer-readable storage medium, and a computer program product. Background Technique

[0002] The statements in this section merely provide background techniques related to the present invention and do not necessarily constitute prior art.

[0003] Cardiac magnetic resonance imaging (CMR) has advantages in quantitatively evaluating the morphology, motion, and blood flow conditions of the heart. High-resolution (HR) cardiac cine magnetic resonance images can dynamically present the motion process of the heart. However, due to the limited acquisition window for k-space (frequency domain) filling, insufficient frequency domain information can be obtained, which limits the spatial resolution of the images. In addition, long scans may cause involuntary movement of the patient, easily generating motion artifacts, thereby further reducing the clarity and diagnostic value of the images. Super-resolution (SR) technology can achieve the reconstruction of low-resolution (LR) MR images obtained by scanning and improve the resolution of the images.

[0004] The existing super-resolution reconstruction of cardiac cine MIR has the following problems:

[0005] (1) Limitations of super-resolution at a fixed scale. Existing MRI super-resolution methods usually can only perform reconstruction for a fixed upsampling scale (such as 2×, 3×, 4×); therefore, these methods have to train and store corresponding neural network models separately for each upsampling multiple, making it difficult for them to flexibly and efficiently meet different imaging conditions and specific clinical requirements.

[0006] (2) Problem of limited spatial resolution caused by insufficient frequency domain information. During the cardiac cine MRI imaging process, the acquisition window for k-space filling is limited, and this acquisition limitation directly leads to insufficient frequency domain information, thereby limiting the spatial resolution of the images. Insufficient frequency domain information will cause the details of the images to be blurred, making it difficult to clearly display key cardiac structural features, which poses a technical challenge to the quantitative analysis of cardiac function and the early diagnosis of lesions.

[0007] (3) Impact of motion artifacts on image quality. Long scans may cause patients to have motion artifacts due to involuntary movement. The presence of motion artifacts not only affects the clarity of the image but may also obscure key lesion information, thereby reducing the diagnostic value of the image. Especially in cardiac cine MRI, capturing the precise dynamic motion process is crucial, and the interference of motion artifacts poses additional technical challenges for reconstructing high-quality dynamic images.

[0008] (4) Limitations of existing diffusion models. Diffusion models have demonstrated excellent capabilities in the field of image enhancement, especially in high-fidelity image generation. However, in MRI super-resolution tasks, diffusion models still have some important limitations. On the one hand, diffusion models usually require a multi-step iterative generation process, which makes the reconstruction time long and difficult to meet clinical application scenarios with real-time requirements. On the other hand, due to limited control ability over details, the super-resolution images generated by diffusion models may have deficiencies in texture performance. Especially when dealing with complex data such as cardiac cine MRI, it is easy to result in the problem that the generated images are not realistic enough. Summary of the Invention

[0009] To solve the problems of fixed-scale reconstruction in existing super-resolution methods, spatial resolution limitations caused by insufficient frequency domain information, the impact of motion artifacts, and the high computational overhead brought by multiple iterative steps in the diffusion model generation process, the present invention provides a method and system for super-resolution reconstruction of cardiac cine MIR at any scale, achieving more accurate super-resolution reconstruction of cardiac cine MRI.

[0010] To achieve the above object, the present invention adopts the following technical solutions:

[0011] In the first aspect, the present invention provides a method for super-resolution reconstruction of cardiac cine MIR at any scale.

[0012] A method for super-resolution reconstruction of cardiac cine MIR at any scale includes the following processes:

[0013] Obtain a sequence of 2N + 1 consecutive video frames I τ-N ,…I τ …,I τ+N , where N is a positive integer greater than or equal to 1;

[0014] Use an encoder to extract the boundary features of the boundary reference frames I τ-N and I τ+N ;

[0015] Use a denoising diffusion model for the intermediate frame sequence I τ-N+1 ,…I τ …,I τ+N-1After processing, diffusion prior features corresponding to each intermediate frame are obtained, and key target frame features are obtained according to the diffusion prior and the boundary features;

[0016] Using the diffusion prior features and the key target frame features as the input of spatial-aware implicit attention, and using the diffusion prior features, the key target frame features, and the output of the spatial-aware implicit attention as the input of temporal-aware implicit attention, a decoder is used to process the output of the temporal-aware implicit attention to obtain the super-resolution reconstruction results corresponding to 2N+1 consecutive video frame sequences.

[0017] In the present invention, the training process includes a pre-training stage and a training stage. In the pre-training stage, an encoder is used to extract the boundary features of the boundary reference frames I τ-N and I τ+N , and a denoising diffusion model is used to process the intermediate frame I τ . In the training stage, an encoder is used to extract the boundary features of the boundary reference frames I τ-N and I τ+N , and a denoising diffusion model is used to process the intermediate frame sequence I τ-N+1 ,…I τ …,I τ+N-1 ...

[0018] In a second aspect, the present invention provides a cardiac cine MIR super-resolution reconstruction system at any scale.

[0019] A cardiac cine MIR super-resolution reconstruction system at any scale, comprising:

[0020] A data acquisition unit, configured to: acquire 2N+1 consecutive video frame sequences I τ-N ,…I τ …,I τ+N , where N is a positive integer greater than or equal to 1;

[0021] A first processing unit, configured to: use an encoder to extract the boundary features of the boundary reference frames I τ-N and I τ+N ;

[0022] A second processing unit, configured to: use a denoising diffusion model to process the intermediate frame sequence I τ-N+1 ,…I τ …,I τ+N-1 After processing, diffusion prior features corresponding to each intermediate frame are obtained, and key target frame features are obtained according to the diffusion prior and the boundary features;

[0023] A reconstruction result generation unit, configured to: use the diffusion prior feature and the key target frame feature as inputs of spatial perception implicit attention, use the diffusion prior feature, the key target frame feature, and the output of the spatial perception implicit attention as inputs of temporal perception implicit attention, and process the output of the temporal perception implicit attention by a decoder to obtain a super-resolution reconstruction result corresponding to a sequence of 2N + 1 consecutive video frames.

[0024] In a third aspect, the present invention provides a computer device, including: a processor and a computer-readable storage medium;

[0025] The processor is adapted to execute a computer program;

[0026] The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the arbitrary-scale cardiac cine MIR super-resolution reconstruction method as described in the first aspect of the present invention.

[0027] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the arbitrary-scale cardiac cine MIR super-resolution reconstruction method as described in the first aspect of the present invention.

[0028] Compared with the prior art, the beneficial effects of the present invention are:

[0029] 1. The implicit neural representation (DP-INR) method based on diffusion prior proposed by the present invention successfully realizes super-resolution reconstruction at any scale by introducing an implicit neural network for modeling. The present invention does not need to train a network model separately for each fixed multiple, and adopts a unified model framework, which can flexibly perform super-resolution reconstruction at different scales according to the requirements of the input image. This means that whether the target is a 2-fold, 3-fold or higher multiple resolution improvement, it can be processed by the same network model, greatly improving the generality and operation efficiency of the model, and being able to better adapt to different imaging requirements and clinical scenarios.

[0030] 2. During the cardiac cine MRI imaging process, due to the limitations of scanning time and k-space filling, it is usually impossible to collect enough frequency domain information. This lack of information directly affects the spatial resolution of the image, resulting in blurred image details, especially in the display of cardiac structure and motion characteristics. Therefore, the present invention designs a method for image super-resolution reconstruction to secondarily improve the spatial resolution of the image in the image domain, clearly presenting the detailed features of the heart, and further improving the accuracy of cardiac function quantitative analysis and the effect of early lesion diagnosis.

[0031] 3. Regarding the impact of motion artifacts on the quality of cardiac cine MRI images, the present invention combines the spatio-temporal relationship between dynamic frames and effectively alleviates the artifact problem caused by involuntary patient movement through spatio-temporal modeling technology. During the image reconstruction process, the present invention introduces a spatio-temporal implicit attention mechanism, which can not only accurately capture the dynamic changes between frames but also effectively separate and eliminate the interference of motion artifacts on the cardiac structure and motion. Through this method, the motion coherence of the reconstructed image is significantly improved, and at the same time, the key cardiac details are clearly presented, greatly enhancing the diagnostic value of the image.

[0032] 4. The generation process of traditional diffusion models requires multiple iterations, which not only incurs a large computational overhead but also significantly increases the reconstruction time, limiting its feasibility in clinical real-time applications. The present invention introduces a skip sampling strategy, taking large strides at key steps during the diffusion process, reducing the number of iterations. At the same time, key prior information is extracted from intermediate steps, avoiding the complete generation process, thus effectively shortening the generation time and retaining the high-quality image features of the diffusion model. Aiming at the deficiency of the diffusion model in detail restoration, the present invention combines implicit neural representation and spatio-temporal modeling technology to accurately model the spatio-temporal structure of cardiac cine MRI, significantly enhancing the detail restoration ability and the realism of the image.

[0033] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become apparent from the following description, or will be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0035] Figure 1 The overall structural diagram of the arbitrary-scale cardiac cine MIR super-resolution reconstruction method provided for Embodiment 1 of the present invention;

[0036] Figure 2 The working schematic diagram of the spatial perception implicit attention provided for Embodiment 1 of the present invention;

[0037] Figure 3 The working schematic diagram of the time perception implicit attention provided for Embodiment 1 of the present invention;

[0038] Figure 4 The comparison diagram of visualization results of multiple methods provided for Embodiment 1 of the present invention;

[0039] Figure 5 The comparison diagram of visualization results of other methods in the time dimension provided for Embodiment 1 of the present invention;

[0040] Figure 6 It is a comparison chart of visualization results outside the scale provided by Embodiment 1 of the present invention with other methods;

[0041] Figure 7 It is a schematic diagram of an arbitrary-scale cardiac cine MIR super-resolution reconstruction system provided by Embodiment 2 of the present invention;

[0042] Figure 8 It is a schematic diagram of a computer device provided by Embodiment 3 of the present invention. Detailed implementation manners

[0043] The present invention will be further described below in conjunction with the drawings and embodiments.

[0044] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0045] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0046] Embodiment 1:

[0047] This implementation proposes an arbitrary-scale cardiac cine MIR super-resolution reconstruction method, and designs an implicit neural representation method based on diffusion prior (DP-INR) for arbitrary-scale cardiac cine MRI super-resolution reconstruction. DP-INR includes two key components: diffusion module (DM) and spatio-temporal implicit attention (ST-IA). The purpose of designing DP-INR in the present invention is to integrate DM and ST-IA together, so as to make more full use of the advantages of DM and achieve MRI reconstruction at arbitrary scales.

[0048] In order to construct a continuous time series with limited spatial resolution as the input, the present invention uses a sliding window video frame reading mechanism, and the size of the sliding window is 2N + 1, and the step size is 1. Therefore, the input of the network consists of 2N + 1 consecutive time series frames: [τ - N, τ + N]. The overall structure of the DP-INR proposed by the present invention is as Figure 1 shown. It can be seen from the figure that the method proposed by the present invention adopts a phased training strategy: pre-training stage and training stage.

[0049] In the pre-training stage, DP-INR uses the intermediate sequence frame I τ to obtain the diffusion prior Z 0 to guide the ST-IA to decode the spatial coordinates and learn the motion trajectory. At the same time, the boundary reference frames I τ-N , I τ+NContext information is extracted by the encoder and provided for the target frame. During the training phase, the complete LR video sequence {I τ-N ,…,I τ+N} is used to optimize the model's parameters so that the model can more accurately reconstruct the high-resolution continuous time series. Specifically, the first and last frames are used as boundary references, but the focus is on the processing of the middle frame sequence [τ-N-1,τ+N-1]. The Time-Embedded Position Encoding (TPE) corresponding to each frame helps the model better understand and process the temporal and spatial position relationships between the middle sequence frames and the boundary reference frames, thereby learning the motion trajectory more effectively.

[0050] In addition, traditional DMs require training over a large number of time steps. The present invention uses an appropriate update rule to "skip" some steps and thus obtain accurate results with fewer iterations.

[0051] In the present invention, preferably, two cardiac cine data sets are used: the CMR×Recon public data set and the In-house clinical data set. For the CMR×Recon data set, the original k-space data of 44 subjects are selected and divided into a training set (30 subjects), a validation set (5 subjects), and a test set (9 subjects). The In-house data set contains 34 subjects and is divided into a training set / validation set / test set as 23 / 4 / 7. Both data sets contain cine data of 6 slices per subject, but each slice of the CMR×Recon data set contains 12 time frames, while the In-house data set contains 15 time frames. The In-house data set is obtained by full-sampling short-axis cardiac MR scanning using the TRUFI sequence, with parameters of repetition time (TR) = 47.88 ms, echo time (TE) = 1.5 ms, slice thickness = 8.0 mm, and field of view (FOV) = 344 mm.

[0052] In the present invention, a diffusion model is introduced to shape specific prior features. As Figure 1 shown in the upper left region, the denoising diffusion model (DDM) used in the present invention defines a diffusion process on the premise of Y∈{I τ-N-1 ,…,I τ+N-1}: Z T →Z T-1 →…→Z 1 →Z 0This inverse process (inference process) requires learning a denoising process, commonly known as the "Joint Distribution", which can be explained as follows:

[0053]

[0054] where Z t represents the prior at time step t, and Z 0 is the desired result, obtained from Z T to Z 0 is a forward process (diffusion process) of gradually adding noise, which can be described as:

[0055]

[0056] To better handle noise and missing values in the data, the present invention applies singular value decomposition (SVD) to project the original data Z t and Y into a new and more compact pixel space. This projection reduces redundancy, highlights important features, and improves data representation. The data representation in the pixel space is given by:

[0057]

[0058] where is used to represent the projection process. Further, for equations (1) and (2), the following definitions are given:

[0059]

[0060] where represents Gaussian noise, σ T is the noise at time step T, and s represents the singular value.

[0061] Classical spatio-temporal SR methods often have difficulty effectively processing information in both spatial and temporal dimensions simultaneously, which may lead to information confusion or loss. When studying and processing complex spatio-temporal data, spatio-temporal implicit neural representation has become an innovative and effective method. It aims to capture and understand the internal patterns and features where space and time are intertwined. Different from previous implicit representations, the present invention proposes to use an attention mechanism to learn parameters that can make the integration effect of local regions better.

[0062] Spatial Aware Implicit Attention (SIA), specifically, includes: for any target LR temporal frame, SIA focuses attention on both visual features and coordinate information simultaneously, thus more precisely capturing information crucial for upsampling at any scale. Specifically, when processing visual features, SIA can highlight those features that are significantly semantically and structurally important at different scales. For coordinate information, it can perceive changes in spatial positions and accordingly adjust the upsampling strategy. The specific structure of SIA is as Figure 2 shown.

[0063] Different from previous self-attention, SIA does not use visual features as queries, but instead uses coordinate information as queries. Specifically, the query spatial coordinates are upsampled, and then the relative offsets of the prior coordinates closest to the query coordinates are calculated. This process can be expressed as:

[0064]

[0065] where Δ is used to represent the solution of the coordinate distance, Z is from the prior feature represents the potential motion closest to the query coordinates , C z is the spatial coordinate of the potential motion Z. The sampling processes of Q, K, and V in SIA are defined as follows:

[0066]

[0067] where S represents the sampling operation for the target, ↑ represents upsampling, and LR and HR respectively represent that the processed features are in the low-dimensional space and the high-dimensional space. Due to the introduction of precise sampling and upsampling strategies for features of different dimensions, the information difference between low-dimensional and high-dimensional features is effectively balanced, thereby optimizing the representation and transmission of features, and thus the resolution gap between them can be alleviated. Since the queries, keys, and values are obtained, furthermore, SIA can be expressed as:

[0068]

[0069] where ψ is a multi-layer perceptron (MLP) used to predict pixel values, [·] is used to represent the concatenation operation, represents matrix multiplication, ω is the temporal position encoding (TPE), which aims to introduce time-related position weights into the model, thereby affecting the subsequent attention calculation and pixel value prediction processes, and its value is related to the position of the target frame in the sequence.

[0070] Time-Aware Implicit Attention (TIA), specifically, includes: when dealing with complex spatio-temporal data, the information in the time dimension is equally crucial. To better learn the dynamic changes and features in the time dimension, the present invention proposes Time-Aware Implicit Attention. As Figure 3 shown, the core idea of TIA lies in focusing on the key target frame features in the time series and the change trends of the front and back reference frames, and further learning the motion trajectories along the time axis, thereby continuously enhancing the model's understanding and representation ability of time information.

[0071] There is a great difficulty in directly relying on the network to construct the decoding characteristics of the target frame, because the network not only needs to learn the movement rules between consecutive time frames, but also needs to grasp the context information. Therefore, the present invention suggests allocating adaptive weights TPE for different time frames to introduce the information in the time dimension, enabling the network to more accurately focus on the target frame and its context, and thus more effectively capture the dynamic patterns in the time series. Formally, TIA can be defined as:

[0072]

[0073] where and The sampling process is similar to formula (7), ω is the weight TPE, and I br is derived from the boundary reference frames I τ-N and I τ+B . Specifically, after cascading and and TPE, rich information related to motion after fusion is obtained by means of MLP while also supplementing the low-frequency components in the target frame . Further, there is:

[0074]

[0075] where and serve as forward motion estimation and backward motion estimation respectively. Next, the Motion flow maps and the boundary reference frame I br perform Motion FlowGuided Grid Warping. The estimation of motion optical flow plays a key role in supplementing the low-frequency components of the target frame. However, considering the possible imperfections in flow estimation, this may lead to inaccuracies in some dynamic information. To make up for this deficiency, the present invention extracts prior features in the last step Local feature enhancement is performed, and at the same time, the high-frequency components in the target frame are supplemented. The specific steps are as follows:

[0076]

[0077] Among them, are the backward motion feature and the forward motion feature respectively. S represents the sampling operation, φ is used to represent Grid Warping, and δ is the convolutional layer.

[0078] In this implementation, the multi-layer perceptron (MLP) is directly incorporated into ST-IA, which is different from the traditional attention mode. The previous attention mechanism used the Softmax function in the calculation, which has certain non-linear characteristics. However, ST-IA in the present invention deliberately ignores this point when calculating attention. This is mainly because non-linear calculations may ignore the information characteristics of other regions or sacrifice the ability to model long-distance dependencies, which may cause the model to be more prone to overfitting in some scenarios. However, the introduction of MLP cleverly makes up for the insufficient model expression ability that may be caused by the abandonment of non-linearity. The powerful non-linear mapping ability of MLP can effectively capture the complex patterns and characteristics in the data and enhance the learning and extraction ability of the model.

[0079] The present invention uses two loss functions to optimize the performance of the model: L1-pixel loss and DC loss.

[0080] L1-pixel loss is a commonly used loss function for measuring the pixel-level difference between the predicted image and the real image. Its definition is as follows:

[0081]

[0082] Among them, represents the i-th frame of the real image, represents the i-th frame of the predicted image, 2N - 1 represents the total number of image frames participating in the calculation for each iteration, and ‖·‖ 1 is the L1 norm.

[0083] DC loss is used to ensure the consistency of the model with the original data when reconstructing the image. Its definition is as follows:

[0084]

[0085] Among them, represents the Fourier transform operation. DC loss emphasizes maintaining the consistency of image data in the frequency domain, thereby improving the quality of the reconstructed image.

[0086] Combining the above two loss functions, the total loss function of the present invention is defined as follows:

[0087]

[0088] where λ 1 and λ 2 are weight parameters used to balance the contributions of the two losses to the total loss.

[0089] The proposed method of the present invention is implemented on an NVIDIA A100 GPU workstation. Continuous frames are read in using a sliding window, and the window size is set to 5, i.e., 2N + 1 = 5. The Adam optimizer is used for network training for 60,000 iterations. In the two-stage training strategy of the present invention, the scaling factor in the first stage (the first 45,000 iterations) is set to 4, and the scaling factor in the second stage (the last 15,000 iterations) is flexibly selected within the range of [2, 4]. The initial learning rate is 10 -4 , the momentum decay coefficient is set to [0.9, 0.999], and the learning rate decays to 10 -7 after every 15,000 iterations. The batch size is set to 16. The time step T in the diffusion model is 5. In addition, we set the two hyperparameters in the training loss: λ 1 to 1 and λ 2 to 0.001. For the decoder, the hidden layer dimensions set by the present invention are [64, 64, 256, 256].

[0090] The present invention uses two metrics, peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), to evaluate the performance of the model. In addition, the present invention also uses normalized mean square error (NMSE) for comparison to measure the performance of the model in the MR reconstruction task. In Figure 4 , the present invention provides a qualitative visual comparison with other SR methods on two datasets, where the scale is 4×. The top and bottom rows are the SR results under the CMR×Recon and In-house datasets, respectively. The second and third rows are the local enlarged areas within the boxes in the first row and the corresponding error maps, respectively. The more texture there is in the error map, the worse the quality of the reconstructed SR result. It can be seen that the method proposed by the present invention shows superior performance in cardiac cine MRI tissue reconstruction. Compared with other methods, the cardiac structure of the present invention is clearer. At the same time, Figure 5 the visualization results also show that the DP-INR of the present invention can capture spatial and temporal detail information more accurately. In addition, Figure 6 shows the SR reconstruction results of the method of the present invention and other arbitrary-scale SR methods at a scale of 5.3× outside the scale. By selecting a non-integer scale factor, the present invention can better evaluate the ability of the model to accurately reconstruct details and structures at uncommon scales, demonstrating its feasibility in dealing with arbitrary-scale scenarios.

[0091] Quantitative comparisons of DP-INR with other reconstruction methods are given in Tables 1 and 2. On the in-house and CMR×Recon datasets, DP-INR has competitive performance compared with other SOTA models. Tables 1 based on in-scale data and Table 2 based on out-of-scale data show that the present invention has achieved the best results in all metrics, indicating its progress in modeling cardiac cycle motion and maximizing the recovery of motion information using spatio-temporal perception.

[0092] Table 1: Quantitative comparison with SOTA on CMR×Recon and in-house datasets (in-scale), with the best performance highlighted in bold. The evaluation metrics include PSNR (dB), SSIM, and NMSE (×10 -2 )

[0093]

[0094]

[0095] Table 2: Quantitative comparison with SOTA on CMR×Recon and in-house datasets (out-of-scale), with the best performance highlighted in bold. The evaluation metrics include PSNR (dB), SSIM, and NMSE (×10 -2 )

[0096]

[0097] The present invention can effectively improve the spatial resolution of images, making the image details clearer, especially in the display of cardiac structures and dynamic motion processes. Compared with traditional super-resolution methods, the technology of the present invention can achieve higher image clarity, improving the visualization effect of cardiac structures and the accuracy of motion processes, which is of great significance for cardiac function analysis, lesion detection, and clinical diagnosis. On multiple cardiac cine MRI datasets, the reconstructed images of the present invention have increased the peak signal-to-noise ratio (PSNR) by approximately 1 - 5.1 dB, the structural similarity index (SSIM) by approximately 0.001 - 0.11, and the maximum reduction in the normalized mean square error (NMSE) has reached 7.5 (×10 -2 ), showing higher image quality.

[0098] Traditional diffusion models require multiple iterations to complete image generation, resulting in large computational costs and long time consumption. By introducing skip sampling and leveraging the prior information of the intermediate steps of the diffusion model, the present invention can significantly shorten the reconstruction time. This method avoids the lengthy iterations in the complete diffusion process, reduces the computational overhead, thereby improving the reconstruction speed and meeting the clinical requirements for real-time and rapid diagnosis; compared with traditional diffusion models, the present invention reduces the reconstruction time by 40%-60%, greatly enhancing the efficiency in clinical use.

[0099] The present invention effectively solves the impact of motion artifacts on image quality by combining spatio-temporal modeling techniques. By adopting a spatio-temporal implicit attention mechanism, it can more precisely model the dynamic process of cardiac motion, suppress motion artifacts, and enhance the stability and detail performance of dynamic images; the present invention adopts a scale-free super-resolution method, which not only improves the efficiency of training and storage, but also can flexibly adapt to different imaging conditions and clinical needs. The present invention does not require training separate models for different upsampling multiples, greatly saving storage space and computing resources.

[0100] Optionally, in some other implementation manners, Transformer can be used for the recovery of frequency-domain information. In the reconstruction of cardiac cine MRI images, Transformer can infer the high-frequency information of the spatial image from the missing frequency-domain information through the self-attention mechanism. Transformer can automatically capture the detailed recovery pattern by learning the global features in the image, thereby making up for the deficiency of frequency-domain information and enhancing the image resolution.

[0101] Optionally, in some other implementation manners, the combination of temporal modeling and image reconstruction improves the image quality. In cardiac cine MRI images, the temporal information between frames is very important for the recovery of details. By introducing a temporal information model (such as long short-term memory network (LSTM) or convolutional-based temporal modeling method), the missing image details can be recovered from the time dimension, improving the spatial resolution. On the basis of temporal modeling, combined with motion compensation technology, by capturing motion information and correcting the image, the impact of motion artifacts can be reduced, further improving the effect of super-resolution reconstruction.

[0102] Optionally, in some other implementation manners, by combining traditional physical modeling methods (such as compressive sensing, extended conjugate gradient method, etc.), a physically constrained deep learning framework can be designed. By modeling the physical limitations (such as sampling pattern, noise model, etc.) in the image recovery process, it helps the network to perform super-resolution reconstruction in the case of insufficient frequency-domain information, but an additional scheme needs to be designed for the problem of arbitrary scale.

[0103] Example 2:

[0104] Such asFigure 7 As shown, this implementation provides a cardiac cine MIR super-resolution reconstruction system at any scale, including:

[0105] A data acquisition unit, configured to: acquire 2N + 1 consecutive video frame sequences I τ-N ,…I τ …,I τ+N , where N is a positive integer greater than or equal to 1;

[0106] A first processing unit, configured to: extract the boundary features of the boundary reference frames I τ-N and I τ+N using an encoder;

[0107] A second processing unit, configured to: process the intermediate frame sequences I τ-N+1 ,…I τ …,I τ+N-1 using a denoising diffusion model, obtain the diffusion prior features corresponding to each intermediate frame, and obtain the key target frame features based on the diffusion prior and the boundary features;

[0108] A reconstruction result generation unit, configured to: use the diffusion prior features and the key target frame features as the input of spatial-aware implicit attention, use the diffusion prior features, the key target frame features, and the output of the spatial-aware implicit attention as the input of temporal-aware implicit attention, and use a decoder to process the output of the temporal-aware implicit attention to obtain the super-resolution reconstruction results corresponding to 2N + 1 consecutive video frame sequences.

[0109] The specific working processes of the above units are described in Embodiment 1 and will not be elaborated here.

[0110] It can be understood that the above units can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the system can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0111] According to another embodiment of the present application, the system described in this embodiment can be constructed, and the method of Embodiment 1 of the present application can be implemented by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), and a Read-Only Memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0112] Embodiment 3:

[0113] As Figure 8 shown, this implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer-readable storage medium 1003. Among them, the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected through a bus or other means.

[0114] Among them, the communication interface 1002 is used to receive and send data. The computer-readable storage medium 1003 can be stored in the memory of the electronic device. The computer-readable storage medium 1003 is used to store a computer program, and the computer program includes program instructions. The processor 1001 is used to execute the program instructions stored in the computer-readable storage medium 1003.

[0115] The processor 1001 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the electronic device, and is adapted to implement one or more instructions. Specifically, it is adapted to load and execute one or more instructions to implement the corresponding method flow or corresponding function.

[0116] The processor 1001 is configured to execute the following process:

[0117] Obtain 2N + 1 consecutive video frame sequences I τ-N ,…I τ …,I τ+N , where N is a positive integer greater than or equal to 1;

[0118] Use an encoder to extract the boundary features of the boundary reference frames I τ-N and I τ+N ;

[0119] Use a denoising diffusion model for the intermediate frame sequence I τ-N+1 ,…Iτ …, I τ+N-1 After processing, diffusion prior features corresponding to each intermediate frame are obtained, and according to the diffusion prior and the boundary features, key target frame features are obtained;

[0120] Using the diffusion prior features and the key target frame features as the input of spatial perception implicit attention, and using the diffusion prior features, the key target frame features and the output of the spatial perception implicit attention as the input of temporal perception implicit attention, a decoder is used to process the output of the temporal perception implicit attention to obtain the super-resolution reconstruction results corresponding to 2N + 1 consecutive video frame sequences.

[0121] For the specific working process, see the introduction in Embodiment 1 and will not be elaborated here.

[0122] Embodiment 4:

[0123] This implementation provides a computer-readable storage medium (Memory). A computer-readable storage medium is a memory device in an electronic device for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space that stores the processing system of the electronic device.

[0124] Moreover, one or more instructions suitable for being loaded and executed by a processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0125] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes one or more instructions stored in the computer-readable storage medium to implement the following process:

[0126] Obtain 2N + 1 consecutive video frame sequences I τ-N , … I τ …, I τ+N , where N is a positive integer greater than or equal to 1;

[0127] Use an encoder to extract the boundary reference frames I τ-N and I τ+N of the boundary features;

[0128] Use a denoising diffusion model for the intermediate frame sequence I τ-N+1 , … Iτ …,I τ+N-1 After processing, diffusion prior features corresponding to each intermediate frame are obtained, and key target frame features are obtained according to the diffusion prior and the boundary features;

[0129] Taking the diffusion prior features and the key target frame features as the input of spatial perception implicit attention, and taking the diffusion prior features, the key target frame features and the output of the spatial perception implicit attention as the input of temporal perception implicit attention, a decoder is used to process the output of the temporal perception implicit attention to obtain the super-resolution reconstruction results corresponding to 2N + 1 consecutive video frame sequences.

[0130] The specific working process is described in Embodiment 1 and will not be elaborated here.

[0131] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0132] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data processing device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0133] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for super-resolution reconstruction of cardiac cine MIR at any scale, characterized in that: The process includes: Get 2N+1 consecutive video frame sequences I τ-N ,···I τ ···,I τ+N , where N is a positive integer greater than or equal to 1; The encoder is used to extract the boundary reference frame I τ-N and I τ+N The boundary characteristics of The denoising diffusion model is used to denoise the intermediate frame sequence I τ-N+1 ,···I τ ···,I τ+N-1 After processing, a diffusion prior feature corresponding to each intermediate frame is obtained, and a key target frame feature is obtained according to the diffusion prior and the boundary feature; The diffusion prior features and the key target frame features are used as the input of spatial-perceptual implicit attention, and the output of the diffusion prior features, the key target frame features and the spatial-perceptual implicit attention are used as the input of temporal-perceptual implicit attention. The output of the temporal-perceptual implicit attention is processed by a decoder to obtain a super-resolution reconstruction result corresponding to 2N+1 continuous video frame sequences.

2. The arbitrary scale heart cine MIR super-resolution reconstruction method according to claim 1, characterized in that: The training process includes a pre-training stage and a training stage. The pre-training stage uses an encoder to extract a boundary reference frame I τ-N and I τ+N The boundary features of the intermediate frame I are also analyzed by using the denoising diffusion model. τ Processing; the training stage uses an encoder to extract the boundary reference frame I τ-N and I τ+N The boundary features of the intermediate frame sequence I are also analyzed by using the denoising diffusion model. τ-N+1 ,···I τ ···,I τ+N-1 to be processed.

3. The arbitrary scale heart cine MIR super-resolution reconstruction method according to claim 1, characterized in that: The denoising diffusion model is used to denoise the intermediate frame sequence I τ-N+1 ,···I τ ···,I τ+N-1 Processing includes: The denoising diffusion model is used to define a τ-N-1 ,···I τ ···,I τ+N-1 Diffusion rules based on the premise of Z T →Z T-1 →…→Z1→Z0, by Z T To Z0 is the forward process of gradual noise addition, and the original data Z is decomposed by singular value decomposition. t and Y are projected into a new, more compact pixel space.

4. The arbitrary scale heart cine MIR super-resolution reconstruction method according to claim 1, characterized in that: The output of the spatially aware implicit attention includes: Among them, [·] is used to represent the connection operation, represents matrix multiplication, ω is the temporal position code, ψ qkv , k and ψ v is the corresponding multi-layer perceptron, d coord is the prior coordinate C closest to the query coordinate z The relative offset of is the diffusion prior feature, is the key target frame feature, S represents the sampling operation for the target, ↑ represents upsampling, LR and HR represent that the processed features are in low-dimensional space and high-dimensional space respectively.

5. The arbitrary scale heart cine MIR super-resolution reconstruction method according to claim 1, characterized in that: The output of the time-aware implicit attention includes: in, are the backward motion features and the forward motion features respectively, δ is the convolution layer, is the diffusion prior feature.

6. The arbitrary scale heart cine MIR super-resolution reconstruction method according to claim 5, characterized in that: Where S represents the sampling operation, φ is used to represent the mesh warping, ω is the weight TPE, is a multi-layer perceptron, are the query vector, key vector, and value vector of time-aware implicit attention, respectively.

7. The arbitrary scale heart cine MIR super-resolution reconstruction method according to any one of claims 1 to 6, characterized in that: The loss function is the sum of L1-pixel loss and DC loss.

8. An arbitrary scale cardiac cine MIR super-resolution reconstruction system, characterized in that: include: The data acquisition unit is configured to: acquire 2N+1 consecutive video frame sequences I τ-N ,···I τ ···,I τ+N , where N is a positive integer greater than or equal to 1; The first processing unit is configured to: extract a boundary reference frame I using an encoder τ-N and I τ+N The boundary characteristics of The second processing unit is configured to: use a denoising diffusion model to process the intermediate frame sequence I τ-N+1 ,···I τ ···,I τ+N-1 After processing, a diffusion prior feature corresponding to each intermediate frame is obtained, and a key target frame feature is obtained according to the diffusion prior and the boundary feature; The reconstruction result generating unit is configured to: use the diffusion prior feature and the key target frame feature as the input of the spatial perception implicit attention, use the diffusion prior feature, the key target frame feature and the output of the spatial perception implicit attention as the input of the temporal perception implicit attention, and use a decoder to process the output of the temporal perception implicit attention to obtain a super-resolution reconstruction result corresponding to 2N+1 continuous video frame sequences.

9. A computer device, characterized in that: include: a processor and a computer readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the arbitrary-scale cardiac cine MIR super-resolution reconstruction method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the arbitrary-scale heart cine MIR super-resolution reconstruction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Tooth segmentation method and system based on iterative boundary optimization and deep learning

    CN115953583A

  • Brain MRI super-resolution reconstruction method and system based on implicit nerve representation

    CN117237196A

  • Method of reconstruction of super-resolution of video frame

    US20220261959A1

  • Pet reconstruction method based on denoising score matching network

    WO2023279316A1