Heart cine MRI reconstruction method and system, medium and computer equipment

By introducing knowledge-guided multi-geometric window Transformer network and adaptive spatiotemporal attention mechanism in cardiac cine MRI reconstruction, the problem of insufficient spatial and temporal information integration in the existing technology is solved, and the reconstruction quality and diagnostic value are significantly improved.

CN120070640AActive Publication Date: 2025-05-30YANTAI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510218445.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The existing cardiac cine MRI reconstruction methods have shortcomings in spatiotemporal information integration, artifact removal, feature alignment and global dependency modeling, resulting in the lack of coherence and realism of details in the reconstruction image, which affects the diagnostic value.

Method used

A multi-geometric window Transformer (KGMgT) network based on knowledge guidance is proposed. By introducing a knowledge guidance mechanism and an adaptive spatiotemporal attention mechanism, it captures spatiotemporal and spatial correlation information of adjacent frames, enhances feature alignment capabilities, realizes global information integration, and improves reconstruction quality.

Benefits of technology

It significantly improves the quality of cardiac cine MRI reconstruction, restores dynamic information of the image under undersampling conditions, improves the diagnostic value of the reconstruction image, reduces the impact of motion artifacts, and ensures the spatial and temporal consistency and detail restoration of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070640A_ABST
    Figure CN120070640A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical image processing. The invention provides a heart Cine MRI reconstruction method and system, a medium and computer equipment, and the method comprises the steps: carrying out the preprocessing of obtained heart Cine MRI data, and obtaining a heart Cine MRI sequence; taking the current frame image and the adjacent previous frame image as the input of the Mentor network, and obtaining a guide feature generated after the Mentor network guides; and taking the guide feature and the adjacent latter frame of image as the input of a Learner network, and obtaining a reconstruction result after the Learner network guidance according to the prior feature in the Mentor network, the guide feature and the adjacent latter frame of image. According to the method, the defects of an existing heart cine MRI reconstruction method in the aspects of spatio-temporal information integration, artifact removal, feature alignment and global dependence modeling are overcome, spatio-temporal correlation information of adjacent frames can be effectively captured, the feature alignment capability is enhanced, global information integration is achieved, and therefore the reconstruction quality of the heart cine MRI is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and in particular to a cardiac cine MRI reconstruction method, a cardiac cine MRI reconstruction system, a computer device, and a computer-readable storage medium. Background Art

[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.

[0003] Cardiac short-axis cine and long-axis cine play important roles in cardiac MRI reconstruction. The short-axis cine provides dynamic images of the transverse section of the heart, allowing for accurate observation of the contraction and relaxation of the myocardium during the heartbeat, providing an important perspective for clinicians to understand cardiac function and diseases. The long-axis cine, on the other hand, shows the view of the heart from the base to the apex, enabling doctors to visually see the morphology and functional status of the main chambers of the heart. Due to the continuous beating of the heart, CMR imaging faces special challenges. Therefore, professional reconstruction techniques play a crucial role in eliminating the blurring and distortion caused by motion.

[0004] The current cardiac cine MRI reconstruction strategies still have the following problems:

[0005] (1) Insufficient deep mining of spatio-temporal correlation information in cardiac cine MRI images. Many current MRI reconstruction methods rely only on static single-frame images for reconstruction, ignoring the potential correlation between adjacent frames in the cine sequence; even if some methods attempt to utilize the information in the time dimension, they usually only focus on the overall spatio-temporal correlation and do not fully mine the fine-grained correlation information between adjacent frames. This limitation results in the reconstructed cardiac cine images lacking detail coherence and realism. Especially under undersampling conditions, important dynamic change information may not be fully recovered, directly affecting the diagnostic value of the images.

[0006] (2) Difficulty in modeling motion artifacts and dynamic changes in cardiac cine MRI sequences. During the actual cardiac MRI imaging process, the continuous beating of the heart leads to strong motion artifacts and deformations, especially in short-axis and long-axis cine images; due to the fast movement speed and complex three-dimensional structure of the heart, existing reconstruction methods are difficult to effectively capture the subtle motion changes in each frame of the cardiac image; for example, during the heartbeat, the contraction and relaxation of the myocardium may cause geometric deformation or image blurring between adjacent frames, making the loss or misalignment of details more serious during image reconstruction; especially under low-resolution and undersampling conditions, these dynamic changes cannot be fully recovered, resulting in the final images lacking accurate dynamic information clinically, which will affect the diagnostic accuracy of doctors in judging cardiac function, disease development, etc.

[0007] (3) Lack of feature alignment mechanism and inefficiency of spatio-temporal information transmission. When processing cardiac cine MRI sequences, due to the dynamic motion characteristics of the heart and the low-resolution characteristics of sampling, there are large deformations or motion artifacts between frames. Existing methods are difficult to accurately align the features between different frames, resulting in the inefficient utilization of supplementary information between adjacent frames. In addition, existing models lack an effective information transmission mechanism and cannot form complete feature sharing and aggregation throughout the sequence, which limits the improvement of the reconstructed image quality.

[0008] (4) Insufficient global information integration ability. Previous reconstruction methods usually rely on fixed-scale feature extraction and local convolution operations, making it difficult to capture the global long-range dependencies in cardiac cine MRI sequences. This leads to limited modeling ability of the model for long-distance spatio-temporal correlations and inability to fully integrate the global information between different frames. Especially during multi-scale feature aggregation, there are prone to feature loss or degradation phenomena, which limits the detail expressiveness of the reconstructed images. Summary of the Invention

[0009] To address the deficiencies of existing cardiac cine MRI reconstruction methods in spatio-temporal information integration, artifact removal, feature alignment, and global dependency modeling, the present invention proposes a Knowledge-guided Multi-Geometry Window Transformer (KGMgT) network, which can effectively capture the spatio-temporal correlation information between adjacent frames, enhance the feature alignment ability, and achieve global information integration, thereby significantly improving the reconstruction quality of cardiac cine MRI.

[0010] To achieve the above object, the present invention adopts the following technical solutions:

[0011] In the first aspect, the present invention provides a method for reconstructing cardiac cine MRI.

[0012] A method for reconstructing cardiac cine MRI includes the following processes:

[0013] Preprocess the acquired cardiac cine MRI data to obtain a cardiac Cine MRI sequence;

[0014] Use the current frame image and the adjacent previous frame image as the input of the Mentor network to obtain the guided features generated after being guided by the Mentor network;

[0015] Use the guided features and the adjacent next frame image as the input of the Learner network. According to the prior features in the Mentor network, the guided features, and the adjacent next frame image, after being guided by the Learner network, obtain the reconstruction result.

[0016] In a second aspect, the present invention provides a cardiac cine MRI reconstruction system.

[0017] A cardiac cine MRI reconstruction system, comprising:

[0018] A preprocessing unit configured to preprocess the acquired cardiac cine MRI data to obtain a cardiac CineMRI sequence;

[0019] A Mentor network guiding unit configured to use the current frame image and the adjacent previous frame image as inputs to the Mentor network to obtain guiding features generated after being guided by the Mentor network;

[0020] A Learner network guiding unit configured to use the guiding features and the adjacent next frame image as inputs to the Learner network, and according to the prior features in the Mentor network, the guiding features, and the adjacent next frame image, obtain a reconstruction result after being guided by the Learner network.

[0021] In a third aspect, the present invention provides a computer device, comprising: a processor and a computer-readable storage medium;

[0022] A processor adapted to execute a computer program;

[0023] A computer-readable storage medium having stored therein a computer program, which when executed by the processor, implements the cardiac cine MRI reconstruction method as described in the first aspect of the present invention.

[0024] In a fourth aspect, the present invention provides a computer-readable storage medium having stored therein a computer program, which is adapted to be loaded and executed by a processor to perform the cardiac cine MRI reconstruction method as described in the first aspect of the present invention.

[0025] Compared with the prior art, the beneficial effects of the present invention are:

[0026] 1. By introducing a knowledge guidance mechanism and drawing on the idea of knowledge distillation, the present invention transfers valuable knowledge information between adjacent frames, thereby significantly enhancing the ability to mine spatio-temporal correlation information; a shared feature encoder is used to extract deep spatial features from different frames of the video sequence, and at the same time, the spatial consistency between MR time frames is explicitly modeled through an adaptive spatio-temporal attention mechanism (ASA), ensuring the detail coherence of cardiac cine MRI images during the dynamic change process; this mechanism effectively restores the dynamic information of the image under undersampling conditions and improves the diagnostic value of the reconstructed image.

[0027] 2. Through the optimization of the overall method, the present invention designs a multi-geometric window Transformer network (KGMgT), which can capture more comprehensive spatio-temporal correlation information throughout the sequence, thus effectively solving the problems of motion artifacts and deformation. The present invention can establish precise spatio-temporal correlation and feature alignment between different frames, improve the reconstruction performance, reduce the influence of motion artifacts, and thus restore the dynamic change process of the heart. Even under complex motion conditions, the network can provide high-quality reconstruction results, ensuring the spatio-temporal consistency and detail restoration of the image.

[0028] 3. Through the adaptive spatio-temporal attention mechanism (ASA) and an efficient image registration method, the present invention accurately calculates the spatio-temporal correlation between frames, thereby realizing feature alignment between different frames, reducing the influence of motion artifacts and deformation. By effectively capturing dynamic changes, the model can accurately restore the motion information of the heart. Even under undersampling conditions, it can still ensure the effective reconstruction of the image.

[0029] 4. To solve the problem of insufficient global information integration ability, the present invention designs a Transformer-driven dynamic multi-scale feature aggregation module (TdFA). By expanding the receptive field, it enhances the model's ability to capture long-distance spatio-temporal dependence relationships. This module effectively reduces the phenomena of feature loss and degradation, improves the aggregation ability of multi-scale features, makes the final reconstructed image richer in detail performance, and has stronger expressiveness and clinical value.

[0030] Advantages of additional aspects of the present invention will be partly given in the following description, partly will become apparent from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0032] Figure 1 It is the overall architecture diagram of the cardiac cine MRI reconstruction method based on the knowledge-guided multi-geometric window Transformer (KGMgT) network provided in Embodiment 1 of the present invention;

[0033] Figure 2 It is the working principle diagram of the adaptive spatio-temporal attention mechanism provided in Embodiment 1 of the present invention;

[0034] Figure 3 It is the working principle diagram of the Transformer provided in Embodiment 1 of the present invention;

[0035] Figure 4Schematic diagram of the operation of multi-geometric window attention provided in Embodiment 1 of the present invention;

[0036] Figure 5 Comparison chart of visualization results of various methods provided in Embodiment 1 of the present invention;

[0037] Figure 6 Schematic diagram of a cardiac cine MRI reconstruction system provided in Embodiment 2 of the present invention;

[0038] Figure 7 Schematic diagram of a computer device provided in Embodiment 3 of the present invention. Detailed implementation manners

[0039] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0040] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0041] Without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0042] Embodiment 1:

[0043] This implementation proposes a cardiac cine MRI reconstruction method based on a knowledge-guided multi-geometric window Transformer (KGMgT) network, and designs a knowledge-guided multi-geometric window Transformer network. As Figure 1 shown, the framework first extracts multi-level deep features from the input cardiac Cine MRI sequence through a feature encoder; on this basis, combined with a knowledge guidance mechanism, efficient transmission of valuable information between different layers of the model is achieved. Subsequently, an adaptive spatio-temporal attention mechanism (ASA) is introduced to accurately capture the spatio-temporal correlation between the target frame and adjacent frames, and the expression ability of global and local information is enhanced through multi-scale feature aggregation; in order to further improve the effect of inter-frame information integration, the present invention designs a Transformer-driven dynamic feature aggregation module (TdFA), which realizes high-quality feature representation by expanding the receptive field and establishing long-distance dependence relationships; finally, the data consistency module ensures that the reconstruction result is consistent with the original sampled data, thereby significantly improving the quality and accuracy of the reconstructed image.

[0044] More specifically, in this implementation, the input data is first preprocessed. The input data is cardiac cine MRI data in complex form, and these data are undersampled. Undersampling usually speeds up the imaging speed by reducing the number of sampling points during the acquisition process, but this will cause information loss and artifact generation in the image.

[0045] For unified calculation, the present invention preprocesses all the reconstructed data sets. Specifically, zero-padding (512×256) is first performed on the edges of the original data set (512×204) in the frequency domain, and it is cropped to a scale of 256×256 in the image domain. Since the obtained MR image is in complex form, the present invention separates the real part and the imaginary part of the complex number and connects them into two channels, avoiding information loss or distortion during the interpolation process of complex data. The data dimension is converted from 256×256 to 256×256×2. After converting to two-channel data, the image is normalized so that the pixel values of the image are kept within a unified range ([0,1]). This step helps to reduce the influence of different MRI acquisition conditions on the image quality and improve the stability and consistency of subsequent network training.

[0046] In this implementation, in order to fully unlock the potential of MR images and better extract their respective deep spatial features from different frames of the video, the present invention designs a feature encoder, as Figure 1 shown. A continuous sequence of cardiac MR images (the present invention uses three frames τ-1, τ, τ+1 to achieve the best reconstruction effect) is input into the feature extraction part of the network; next, the deep features of the target frame and the adjacent reference frame are respectively extracted and denoted as Subsequently, the pre-trained VGG19 network is used to perform multi-scale feature extraction on Ref τ-1 to obtain deep feature representations at three different scales, denoted as The above process can be expressed as follows:

[0047]

[0048] where ψ represents deep feature extraction, and (,) is used to separate multiple variables so that the function receives different values respectively. The above is the extraction process of the Mentor network, and the extraction process in the Learner network is the same as above.

[0049] In this implementation, a knowledge guidance mechanism is designed, as Figure 1 shown. Specifically, Input τ and the previous frame reference frame Ref τ-1 are used as the input of the Mentor network, and the features generated after being guided by the Mentor network and the subsequent reference frame As the input of the Learner network, the goal of this mechanism is to make full use of the prior knowledge in the Mentor network to capture more relevant information. Therefore, the Learner network receives prior features from the Mentor network at each scale. Through the method of transfer learning, the knowledge transfer from the Mentor network to the Learner network is realized, thus maximizing the performance and generalization ability of the Learner network. The knowledge transfer process can be expressed as:

[0050]

[0051] where κ represents knowledge guidance, and its subscripts m and l represent the Mentor network and the Learner network respectively. Rec is the reconstruction result. In this process, the Mentor network focuses on providing knowledge guidance, while the Learner network is responsible for updating the weight parameters.

[0052] In this implementation, the multi-scale depth features from Input τ and Ref τ-1 where 1 ≤ i ≤ 3, and the shallow features extracted from Input τ through convolution are passed to the Mentor network together.

[0053] In this implementation, the Learner network receives the prior features guided by knowledge and uses the multi-scale features extracted from Reft t+1 where 1 ≤ j ≤ 3, and the shallow features from The Learner network has a similar structure to the Mentor network, but there is a key difference, that is, the Learner network uses the knowledge guidance mechanism to refine the knowledge from the Mentor network at each scale.

[0054] In this implementation, multi-scale feature aggregation is also performed. Taking the Learner network as an example, given the deep features and the shallow features Multi-scale feature aggregation refines the features at three different scales through cascading adaptive spatio-temporal attention and Transformer-driven dynamic feature aggregation; at the same time, using the skip connection feature of the U-net network, the multi-scale features are iterated to better capture the feature information at different scales. This process is as shown in Figure 1 multi-scale feature aggregation.

[0055] In this implementation, in order to effectively learn spatial and temporal features, the present invention introduces a new Adaptive Spatio-Temporal Attention (ASA). As shown in Figure 2 , ASA has the ability to dynamically adapt to different local receptive fields according to different scale sizes, and can effectively achieve image registration of adjacent view features to ensure their spatial alignment.

[0056] Specifically, the present invention uses the extracted from the feature encoder as the initial Query (Q), Key (K), and Value (V) respectively. By comparing the local features of Q and K, similarity information is generated, and the similarity information reflects the correlation between features at different times. The generation of the similarity information can be expressed by the following formula: The generation of the similarity information can be expressed by the following formula:

[0057]

[0058] where Corr represents correlation matching, are the correlation value and the position index respectively, pad is the padding operation, and (p1, p2, p3, p4) is a quadruple representing the padding amounts in four directions, which is set to (1, 1, 1, 1) in the present invention. Further, the displacement information is used to generate the optical flow. In order to explore the spatio-temporal correlation more deeply, the present invention uses optical flow estimation to capture the temporal correlation, and the generation process of the optical flow information is as shown in the formula:

[0059]

[0060] where mod is the modulo operation, h represents the height, w represents the width, (x, y) takes values from [0, w - 1] and [0, h - 1] respectively, and grid is used to generate the coordinates of all pixel positions. Then, based on the displacement information is generated The displacement information describes the displacement of each pixel in relative to τ The generation process of O

[0061] O τ = Shift(flow τ , (λ x , λ y )) (8);

[0062] where Shift represents the translation operation, (λ x , λ yThe values in τ and V are in the range of [0, 2], so there are a total of n = 3×3 different offset versions. Next, perform a mapping operation on flow Figure 2 and V to map the features of the target frame to the spatial positions aligned with the adjacent reference frames, so that they are aligned in the same coordinate system for accurate 3D reconstruction, as

[0063]

[0064] shown. The subsequent calculation process is as follows: represents the concatenation operation, is the mapping operation, σ is the sigmoid activation function, θ is the tanh activation function, and ⊕ represent matrix multiplication and addition operations respectively. The target frame feature is one of the original inputs of ASA, and its initial value is Using the calculated multi-view information of the front and back frames, a deformable convolutional network (DCN) is used to model the spatio-temporal correlation features:

[0065]

[0066] In this implementation, in order to aggregate spatio-temporal features at multiple scales the target frame feature and the prior feature The present invention designs a Transformer-driven dynamic feature aggregation module (TdFA) to merge the feature maps from multiple frames. As Figure 3 shown, TdFA consists of a convolutional layer and n multi-geometric window Transformers (MGwinT), and the present invention sets n = 4. Specifically, first concatenate the feature maps in the spatial dimension, then perform preliminary fusion using convolution, and then use MGwinT to obtain the target frame feature at the next scale In addition, the present invention also uses a residual connection to stabilize the training process. The specific process is as follows:

[0067]

[0068] where Conv is the convolution operation, represents MGwinT, is the concatenation operation.

[0069] In this implementation, a multi-geometric window Transformer is adopted. Some tissues in the heart area have similarities, as Figure 4The local similarities in the marked areas (I), (II), (III), and (VI), and the global similarities in the marked areas shown in (IV) and (V). This means that dependencies can be established both locally and globally, thereby capturing more similar features. However, using a square sliding window alone or a rectangular window, without adjusting the window size, it is impossible to establish good local and overall dependencies. Therefore, the present invention further considers the combination of the two, that is, using two Transformers with different windows simultaneously to further enhance the global dependencies of the network aggregation module. The present invention proposes a new multi-geometric window attention Transformer (MGwinT) that effectively establishes long- and long-distance dependencies while fusing spatio-temporal information in adjacent frames, such as Figure 3 shown. To reduce the computational complexity of window attention, the present invention executes two window attentions in parallel in the same MGwinT, namely square window multi-head self-attention (SW-MSA) and rectangular window multi-head self-attention (RW-MSA).

[0070] For the design of the square window (S-Window), a regular window partitioning strategy is used. Starting from the bottom-left pixel, the feature map is divided into α×α β×β windows ( Figure 4 as shown in which α = 2 and β = 6), and the regularly partitioned windows are cyclically shifted by Figure 4 pixels to the upper-left, as shown by the position changes in areas (I), (II), and (III) in

[0071] For the design of the rectangular window (R-Window), as shown in Figure 4 (IV), (V), and (VI). The present invention divides it into two types, vertical windows and horizontal windows, and serially uses them in RW-MSA of different scales. By alternately using the two windows, global-scale modeling can be maintained to capture more similar features without expanding the window size.

[0072] In this implementation, data consistency is introduced. In MRI reconstruction, introducing data consistency (DC) constraints in the k-space domain (frequency domain) can effectively improve the image quality. This is because the generated image may have information loss in the frequency domain. The data consistency method uses the Fourier transform to relate the image domain (Rec′, GT) to the k-space domain Specifically, the data consistency constraint ensures that the sampled area in the k-space is consistent with the original k-space data, while the unsampled area retains the values estimated from the reconstruction process. This selective fidelity helps to optimize the recovery of frequency domain information. Mathematically, the coefficients of the sampled area are constrained to be the same as the original data match, while the coefficients of the non-sampled regions are consistent with the reconstruction result remain consistent, and the process can be expressed by the following formula:

[0073]

[0074] where η≥0 represents the noise term, [m,n] represents matrix index operation, represents the k-space data after applying the DC constraint, is the sampling mask, after performing the inverse Fourier transform on its corresponding representation in the image domain can be restored, and the detailed DC process is as shown in the right region of Figure 1

[0075] To construct the training dataset, fully sampled short-axis cardiac MR dynamic images of 27 healthy volunteers (a total of 2430 images, each volunteer contains 6 groups of slices, 15 time frames in a group) were randomly selected. During the training process, the fully sampled data was retrospectively undersampled through a random Cartesian undersampling mask, combined with four different acceleration factors, to generate the initial input (Input / Ref) of the network. To evaluate the difference between the reconstruction result Rec and the gold standard GT and improve the reconstruction quality, the present invention uses three loss functions, namely L1-pixel loss, Perceptual loss, and Adversarial loss, to optimize the model.

[0076] To evaluate the performance of the network proposed by the present invention, three cardiac cine GRAPPA reconstruction datasets were used, including an in-house dataset (short-axis direction) and two CMR×Recon datasets (short-axis direction and long-axis direction). At the same time, to ensure the security and privacy during the data usage process, all data were desensitized and anonymized before use to reduce the risk of patient identity information leakage.

[0077] The method proposed by the present invention is implemented on an NVIDIA A100 GPU (40GB) workstation based on the PyTorch framework. The optimization algorithm uses Adam, and the learning rate is set to 1×10 -4 , the momentum decay coefficient is set to [0.9, 0.999], and the batch size is set to 2. In addition, two hyperparameters in the training loss: λ 1 = 1e -4 , λ 2 = 1e -6 . Considering that there is a high similarity between adjacent frames of cardiac cine, the number of input frames is set to 3, and this design helps the model better capture the temporal correlation and the similar information between frames.​

[0078] As shown in Tables 1, 2, 3, 4, 5, and 6, the present invention compares the quantitative performance of KGMgT with other reconstruction methods on in-house datasets and CMR×Recon datasets. The results show that for both short-axis cine and long-axis cine, KGMgT achieves the best results in all metrics, demonstrating significant competitiveness compared to the current state-of-the-art methods. This indicates that the method proposed in the present invention has made a breakthrough in modeling the cardiac cycle motion, maximizing the recovery of motion information through spatio-temporal perception capabilities.

[0079] Table 1: Quantitative comparison with SOTA on in-house (4×, 6×) short-axis (Sax) datasets. The best performance is highlighted in bold. Evaluation metrics include PSNR (dB), SSIM, and NMSE (×10 -1 )

[0080]

[0081] Table 2: Quantitative comparison with SOTA on in-house (8×, 10×) short-axis (Sax) datasets. The best performance is highlighted in bold. Evaluation metrics include PSNR (dB), SSIM, and NMSE (×10 -1 )

[0082]

[0083]

[0084] Table 3: Quantitative comparison with SOTA on CMR×Recon (4×, 6×) short-axis (Sax*) datasets. The best performance is highlighted in bold. Evaluation metrics include PSNR (dB), SSIM, and NMSE (×10 -1 )

[0085]

[0086] Table 4: Quantitative comparison with SOTA on CMR×Recon (8×, 10×) short-axis (Sax*) datasets. The best performance is highlighted in bold. Evaluation metrics include PSNR (dB), SSIM, and NMSE (×10 -1 )

[0087]

[0088] Table 5: Quantitative comparison with SOTA on the CMR×Recon (4×, 6×) long-axis (Lax*) dataset. The best performance is highlighted in bold. The evaluation metrics include PSNR (dB), SSIM, and NMSE (×10 -1 ).

[0089]

[0090]

[0091] Table 6: Quantitative comparison with SOTA on the CMR×Recon (8×, 10×) long-axis (Lax*) dataset. The best performance is highlighted in bold. The evaluation metrics include PSNR (dB), SSIM, and NMSE (×10 -1 ).

[0092]

[0093] Compared with the traditional method, the peak signal-to-noise ratio (PSNR) of the reconstructed image by the present invention is increased by about 0.82 - 5.83 dB, the structural similarity (SSIM) is increased by about 0.03 - 0.3, and the maximum reduction of the normalized mean square error (NMSE) reaches 3.8297 (×10 -1 ), significantly improving the clarity and structural fidelity of the image. This indicates that the model can more accurately recover the details of the cardiac cine MRI image during the reconstruction process, avoiding the blurring phenomenon in the traditional method.

[0094] Compared with the existing method, the model of the present invention can reduce the inter-frame information loss by about 10% - 15%, thereby more comprehensively integrating spatio-temporal information and improving the continuity and accuracy of the reconstructed image.

[0095] Compared with the reconstruction technology based on the traditional optimization method, the present invention can achieve higher image quality under the same acceleration factors (4×, 6×, 8×, 10×), and the calculation speed of the reconstruction process is increased by about 15%. This means that the image reconstruction task can be completed more quickly with the same computing resources, improving the efficiency of clinical diagnosis.

[0096] Due to the adoption of the ASA mechanism and the multi-geometric window strategy, the present invention can effectively reduce the artifacts caused by cardiac motion; compared with the traditional method, the artifacts are significantly reduced by about 20% - 30%, improving the usability and accuracy of the image, especially in the case of rapid cardiac motion, and can better retain the key structural information.

[0097] By reducing the acquisition time and combining it with the improvement of image quality, the present invention can achieve higher precision in cardiac MRI image reconstruction, and shorten the scanning time for each patient (taking a 10-fold acceleration as an example, the original scanning time of 10 minutes can be shortened to 1 minute), reducing the discomfort of patients that may be caused by long-term scanning. It is particularly suitable for emergency or rapid diagnosis scenarios and improves the efficiency of clinical operations.

[0098] It can be understood that in some other implementation manners, the key parameters used in the present invention can be dynamically adjusted to achieve multiple functions to meet different requirements, and the parameter settings of different models include, but are not limited to: adjustment of network depth, weight initialization method or optimization function; different data acquisition, preprocessing and postprocessing methods, such as image enhancement methods or data interpolation methods in the field of images; the present invention is targeted at a specific field (medical imaging) and can be extended to other similar fields (such as industrial inspection, remote sensing images).

[0099] In the present invention, the spatio-temporal attention mechanism (ASA) is used to accurately capture the spatio-temporal correlation between the target frame and adjacent frames. It can be understood that in some other implementation manners, if other technologies are used to implement this function, the following alternative solutions can be considered: The graph neural network models the relationship between frames by propagating information on the graph structure. Each frame is regarded as a node in the graph, and the connection weights between nodes are obtained through spatio-temporal feature learning. This method is suitable for situations where the relationship between frames in the data is relatively complex or non-linear.

[0100] It can be understood that in some other implementation manners, dilated convolution is used to achieve feature aggregation. By using dilated convolution instead of traditional convolution operations, the receptive field is extended by introducing holes in the convolution kernel, improving the ability to capture long-range dependencies. By stacking dilated convolution modules, the receptive field range can also be increased.

[0101] Embodiment 2:

[0102] As Figure 6 shown, this implementation manner provides a cardiac cine MRI reconstruction system, including:

[0103] A preprocessing unit, configured to: preprocess the acquired cardiac cine MRI data to obtain a cardiac CineMRI sequence;

[0104] A Mentor network guiding unit, configured to: use the current frame image and the adjacent previous frame image as the input of the Mentor network to obtain the guiding features generated after being guided by the Mentor network;

[0105] The Learner network guiding unit is configured to: use the guiding feature and the adjacent subsequent frame image as the input of the Learner network, and obtain a reconstruction result after being guided by the Learner network according to the prior feature in the Mentor network, the guiding feature, and the adjacent subsequent frame image.

[0106] For the specific working methods of the above units, please refer to the description in Embodiment 1 and will not be elaborated here.

[0107] It can be understood that the above units can be separately or wholly combined into one or several other units to form, or some of them can be further split into multiple smaller units in terms of function to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the system can also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0108] According to another embodiment of this application, the system described in this embodiment can be constructed by running a computer program (including program code) that can execute the steps involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), and a Read-Only Memory (ROM). The computer program can be recorded on a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0109] Embodiment 3:

[0110] As Figure 7 shown, this implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer-readable storage medium 1003. Among them, the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected through a bus or other means.

[0111] Among them, the communication interface 1002 is used for receiving and sending data. The computer-readable storage medium 1003 can be stored in the memory of the electronic device. The computer-readable storage medium 1003 is used for storing computer programs, and the computer programs include program instructions. The processor 1001 is used for executing the program instructions stored in the computer-readable storage medium 1003.

[0112] The processor 1001 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0113] The processor 1001 is configured to execute the following process:

[0114] Preprocess the acquired cardiac cine MRI data to obtain a cardiac Cine MRI sequence;

[0115] Use the current frame image and the adjacent previous frame image as the input of the Mentor network to obtain the guided features generated after being guided by the Mentor network;

[0116] Use the guided features and the adjacent next frame image as the input of the Learner network. According to the prior features in the Mentor network, the guided features, and the adjacent next frame image, after being guided by the Learner network, obtain the reconstruction result.

[0117] For the specific working method, please refer to the introduction in Embodiment 1 and will not be elaborated here.

[0118] Embodiment 4:

[0119] This implementation provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the electronic device and is used for storing programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the electronic device.

[0120] Moreover, in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0121] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the one or more instructions stored in the computer-readable storage medium are loaded and executed by a processor to implement the following process:

[0122] Preprocess the acquired cardiac cine MRI data to obtain a cardiac Cine MRI sequence;

[0123] Use the current frame image and the adjacent previous frame image as the input of the Mentor network to obtain the guided features generated after being guided by the Mentor network;

[0124] Use the guided features and the adjacent next frame image as the input of the Learner network, and according to the prior features in the Mentor network, the guided features, and the adjacent next frame image, obtain a reconstruction result after being guided by the Learner network.

[0125] For the specific working method, see the introduction in Embodiment 1 and will not be elaborated here.

[0126] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A cardiac cine MRI reconstruction method, characterized in that: The process includes: Preprocessing the acquired cardiac Cine MRI data to obtain a cardiac Cine MRI sequence; The current frame image and the adjacent previous frame image are used as the input of the Mentor network to obtain the guided features generated after being guided by the Mentor network; The guiding features and the adjacent next frame image are used as inputs of the Learner network, and a reconstruction result is obtained after being guided by the Learner network according to the prior features in the Mentor network, the guiding features and the adjacent next frame image.

2. The cardiac cine MRI reconstruction method according to claim 1, characterized in that: In the Mentor network, a feature encoder is used to extract multi-scale features of the current frame image and the adjacent previous frame image, and the multi-scale features are used as the query vector, key vector and value vector of adaptive spatiotemporal attention, combined with the target frame features of the current frame image, to obtain the spatiotemporal correlation features.

3. The cardiac cine MRI reconstruction method according to claim 2, characterized in that: In the Mentor network, a Transformer-driven dynamic feature aggregation module is used to aggregate spatiotemporal correlation features and target frame features at multiple scales to obtain target frame features at the next scale. After being processed by a denoising encoder and data consistency, guided features are obtained.

4. The cardiac cine MRI reconstruction method according to claim 1, characterized in that: In the Learner network, a feature encoder is used to extract multi-scale features of the current frame image and the adjacent next frame image, and the multi-scale features are used as the query vector, key vector and value vector of adaptive spatiotemporal attention, combined with the target frame features of the current frame image, to obtain spatiotemporal correlation features.

5. The cardiac cine MRI reconstruction method according to claim 4, characterized in that: In the Learner network, a Transformer-driven dynamic feature aggregation module is used to aggregate spatiotemporal correlation features, target frame features and prior features of the Mentor network at multiple scales to obtain target frame features at the next scale. After denoising encoder and data consistency processing, a reconstructed image is obtained.

6. The cardiac cine MRI reconstruction method according to claim 3 or 5, characterized in that: The Transformer-driven dynamic feature aggregation module includes convolutional layers and multiple multi-geometric window Transformers, which connect the spatiotemporal correlation features at multiple scales, target frame features, and the prior features of the Mentor network in the spatial dimension, use convolutional layers for preliminary fusion, and use multiple multi-geometric window Transformers to obtain the target frame features of the next scale.

7. The cardiac cine MRI reconstruction method according to claim 2 or 4, characterized in that: The multi-scale features are used as query vectors, key vectors and value vectors of adaptive spatiotemporal attention, combined with the target frame features of the current frame image, to obtain spatiotemporal correlation features, including: By comparing the local features of the query vector and the key vector, similarity information is generated. The similarity information reflects the correlation between the features at different times. The generation of is represented as: Among them, Corr represents correlation matching, are the correlation value and position index respectively, pad(·) is the padding operation, (p1, p2, p3, p4) is a four-tuple indicating the number of paddings in four directions, Q and K are the query vector and key vector respectively. With the help of optical flow estimation, we can capture the temporal correlation and obtain the optical flow information flow τ , based on optical flow information flow τ Generate displacement information O τ , for optical flow information flow τ The sum vector V uses a mapping operation to map the features of the target frame to a spatial position aligned with the adjacent reference frames so that they are aligned in the same coordinate system, and utilizes the calculated multi-view information of the previous and next frames to model the spatiotemporal correlation features by using a deformable convolutional network.

8. A cardiac cine MRI reconstruction system, characterized in that: include: The preprocessing unit is configured to: preprocess the acquired cardiac Cine MRI data to obtain a cardiac Cine MRI sequence; The Mentor network guidance unit is configured as follows: the current frame image and the adjacent previous frame image are used as inputs of the Mentor network, and the guidance features generated after being guided by the Mentor network are obtained; The Learner network guiding unit is configured to: use the guiding features and the adjacent next frame image as inputs of the Learner network, and obtain a reconstruction result after being guided by the Learner network according to the prior features in the Mentor network, the guiding features and the adjacent next frame image.

9. A computer device, characterized in that: include: a processor and a computer readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the cardiac cine MRI reconstruction method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the cardiac cine MRI reconstruction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video super-resolution reconstruction method and system based on multi-scale local self-attention

    CN115082308A

  • Lightweight video super-resolution reconstruction method based on hybrid space-time convolution

    CN117830095A

  • Joint forecasting of feature and feature motion

    US20220180133A1