Cardiac cine MRI reconstruction method, system, medium, and computer device
Through the knowledge-guided multi-geometric window Transformer network, the problems of spatiotemporal information integration and motion artifacts in cardiac cine MRI reconstruction are solved, high-quality image reconstruction is achieved, and diagnostic accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202510218445.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing cardiac cine MRI reconstruction methods have deficiencies in spatiotemporal information integration, artifact removal, feature alignment, and global dependency modeling, resulting in a lack of detail coherence in the reconstructed images, incomplete recovery of dynamic information, severe motion artifacts, and feature loss, which affects diagnostic accuracy.
The knowledge-guided multi-geometric window Transformer (KGMgT) network is adopted. By introducing the knowledge distillation mechanism and the adaptive spatiotemporal attention mechanism (ASA), combined with the Transformer-driven dynamic multi-scale feature aggregation module (TdFA), the spatiotemporal correlation information mining and global information integration between adjacent frames are realized, thereby improving the feature alignment capability.
It significantly improves the reconstruction quality of cardiac cine MRI images, enhances the coherence of dynamically changing details and the spatiotemporal consistency of images, reduces motion artifacts, and improves the diagnostic value and clinical application value of reconstructed images.
Smart Images

Figure CN120070640B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image processing, in particular to a cardiac cine MRI reconstruction method, a cardiac cine MRI reconstruction system, a computer device and a computer readable storage medium. BACKGROUND
[0002] The statements in this section merely provide background technology related to the present application and do not necessarily constitute prior art.
[0003] Cardiac short-axis cine and long-axis cine play an important role in cardiac MRI reconstruction. Short-axis cine provides dynamic images of the transverse section of the heart, which can accurately observe the contraction and relaxation of the myocardium during the heartbeat, providing an important perspective for clinicians to understand the function and disease of the heart. Long-axis cine shows the perspective of the heart from the base to the top, allowing doctors to intuitively see the shape and functional state of the main chambers of the heart. Due to the continuous beating of the heart, CMR imaging faces special challenges. Therefore, professional reconstruction techniques play a crucial role in eliminating blurring and distortion caused by motion.
[0004] Current cardiac cine MRI reconstruction strategies still have the following problems:
[0005] (1) Insufficient depth mining of spatio-temporal correlation information in cardiac cine MRI images. Many current MRI reconstruction methods rely only on static single-frame images for reconstruction, ignoring the potential correlation between adjacent frames in the cine sequence. Even if some methods try to use information in the time dimension, they usually only focus on the overall spatio-temporal correlation, without fully mining the fine-grained correlation information between adjacent frames. This limitation results in a lack of detail and coherence in the reconstructed cardiac cine images, and especially under undersampling conditions, important dynamic change information may not be fully recovered, directly affecting the diagnostic value of the images.
[0006] (2) Difficulty in modeling motion artifacts and dynamic changes in cardiac cine MRI sequences. In actual cardiac MRI imaging, the continuous beating of the heart leads to strong motion artifacts and deformation, especially in short-axis and long-axis cine images. Due to the fast motion of the heart and its complex three-dimensional structure, existing reconstruction methods have difficulty in effectively capturing the subtle motion changes in each frame of the heart image. For example, during the beating of the heart, the contraction and relaxation of the myocardium can cause geometric deformation or image blurring between adjacent frames, making the details missing or misplaced more serious during image reconstruction. Especially under low resolution and undersampling conditions, these dynamic changes cannot be fully recovered, resulting in a lack of accurate dynamic information in the final images, which affects the diagnostic accuracy of doctors in judging the function of the heart, disease development, etc.
[0007] (3) Lack of feature alignment mechanism and inefficiency of spatio-temporal information transmission. In processing cardiac cine MRI sequences, due to the dynamic motion characteristics of the heart and the low resolution characteristics of the sampling, there are large deformations or motion artifacts between frames. The existing methods are difficult to accurately align the features between different frames, resulting in the inability to efficiently utilize the complementary information of adjacent frames. In addition, the existing models lack effective information transmission mechanisms, and cannot form a complete feature sharing and aggregation in the entire sequence, which limits the improvement of the quality of the reconstructed image.
[0008] (4) Insufficient global information integration capability. Previous reconstruction methods usually rely on fixed-scale feature extraction and local convolution operations, which are difficult to capture global long-distance dependencies in cardiac cine MRI sequences. This results in limited modeling capability of the model for long-distance spatio-temporal associations, and the inability to fully integrate global information between different frames, especially when multi-scale feature aggregation occurs, which can lead to feature loss or degradation, limiting the detail expressiveness of the reconstructed image. SUMMARY
[0009] In order to solve the problems of the existing cardiac cine MRI reconstruction methods in spatio-temporal information integration, artifact removal, feature alignment and global dependency modeling, the present application proposes a knowledge-guided multi-geometry window Transformer (KGMgT) network, which can effectively capture the spatio-temporal association information of adjacent frames, enhance the feature alignment capability, and realize the global information integration, thereby significantly improving the reconstruction quality of the cardiac cine MRI.
[0010] In order to achieve the above purpose, the present application adopts the following technical solutions:
[0011] In a first aspect, the present application provides a cardiac cine MRI reconstruction method.
[0012] A cardiac cine MRI reconstruction method, comprising the following processes:
[0013] Pretreating the acquired cardiac cine MRI data to obtain a cardiac cine MRI sequence;
[0014] The current frame image and the adjacent previous frame image are input into the Mentor network to obtain guided features generated after being guided by the Mentor network;
[0015] The guided features and the adjacent next frame image are input into the Learner network, and the reconstruction result is obtained after being guided by the Learner network according to the prior features in the Mentor network, the guided features and the adjacent next frame image.
[0016] In a second aspect, the present application provides a cardiac cine MRI reconstruction system.
[0017] A cardiac cine MRI reconstruction system comprises:
[0018] A preprocessing unit is configured to preprocess acquired cardiac cine MRI data to obtain a cardiac Cine MRI sequence;
[0019] A Mentor network guiding unit is configured to take a current frame image and an adjacent previous frame image as inputs of a Mentor network to obtain guided features generated after Mentor network guidance;
[0020] A Learner network guiding unit is configured to take the guided features and an adjacent next frame image as inputs of a Learner network to obtain a reconstruction result after Learner network guidance according to prior features in the Mentor network, the guided features and the adjacent next frame image.
[0021] In a third aspect, the present application provides a computer device comprising a processor and a computer readable storage medium.
[0022] The processor is adapted to execute a computer program.
[0023] The computer readable storage medium has a computer program stored therein, and the computer program, when executed by the processor, implements the cardiac cine MRI reconstruction method according to the first aspect of the present application.
[0024] In a fourth aspect, the present application provides a computer readable storage medium having a computer program stored therein, and the computer program is adapted to be loaded and executed by a processor to implement the cardiac cine MRI reconstruction method according to the first aspect of the present application.
[0025] Compared with the prior art, the present application has the following beneficial effects:
[0026] 1. The present application introduces a knowledge guiding mechanism, draws on the idea of knowledge distillation, and transfers valuable knowledge information between adjacent frames, thereby significantly improving the mining ability of spatio-temporal correlation information; a shared feature encoder is used to extract deep spatial features from different frames of a video sequence, and an adaptive spatio-temporal attention mechanism (ASA) is used to explicitly model the spatial consistency between MR time frames, thereby ensuring the detail coherence of the cardiac cine MRI image in the dynamic change process; this mechanism effectively restores the dynamic information of the image under the condition of undersampling, and improves the diagnostic value of the reconstructed image.
[0027] 2、The application optimizes the overall method and designs a multi-geometry window Transformer network (KGMgT), which can capture more comprehensive spatio-temporal correlation information in the entire sequence, thereby effectively solving the motion artifact and deformation problem, the application can establish accurate spatio-temporal correlation and feature alignment between different frames, improve the reconstruction performance, reduce the influence of motion artifacts, and restore the dynamic change process of the heart, even under complex motion conditions, the network can provide high-quality reconstruction results, and ensure the spatio-temporal consistency and detail restoration of the image.
[0028] 3、The application accurately calculates the inter-frame spatio-temporal correlation through an adaptive spatio-temporal attention mechanism (ASA) and an efficient image registration method, thereby realizing feature alignment between different frames, reducing the influence of motion artifacts and deformation, and accurately restoring the motion information of the heart by effectively capturing dynamic changes, even under undersampling conditions, the image can still be effectively reconstructed.
[0029] 4、In order to solve the problem of insufficient global information integration capability, the application designs a Transformer-driven dynamic multi-scale feature aggregation module (TdFA), which expands the receptive field and enhances the model's ability to capture long-range spatio-temporal dependencies, effectively reduces feature loss and degradation, improves multi-scale feature aggregation capability, and makes the final reconstructed image more detailed and expressive.
[0030] The advantages of the additional aspects of the application will be partially given in the following description, partially will become obvious from the following description, or will be understood by the practice of the application. BRIEF DESCRIPTION OF DRAWINGS
[0031] The drawings accompanying the specification of this application form a part thereof, serve to provide further understanding of the application, and together with the description of the exemplary embodiments of the application and their description serve to explain the application, and do not constitute an improper limitation of the application.
[0032] Figure 1 The overall architecture diagram of the heart cine MRI reconstruction method based on the knowledge-guided multi-geometry window Transformer (KGMgT) network provided for embodiment 1 of the application;
[0033] Figure 2 The working principle diagram of the adaptive spatio-temporal attention mechanism provided for embodiment 1 of the application;
[0034] Figure 3 The working principle diagram of the Transformer provided for embodiment 1 of the application;
[0035] Figure 4A running schematic diagram of multi-geometry window attention provided for the embodiment 1 of the present application;
[0036] Figure 5 A comparison chart of multi-method visualization results provided for the embodiment 1 of the present application;
[0037] Figure 6 A schematic diagram of a cardiac cine MRI reconstruction system provided for the embodiment 2 of the present application;
[0038] Figure 7 A schematic diagram of a computer device provided for the embodiment 3 of the present application. DETAILED DESCRIPTION
[0039] The present application will be further described below in conjunction with the accompanying drawings and embodiments.
[0040] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0041] The embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0042] Embodiment 1:
[0043] The present implementation proposes a cardiac cine MRI reconstruction method based on knowledge-guided multi-geometry window Transformer (KGMgT) network, and designs a knowledge-guided multi-geometry window Transformer network, as shown in the following formula: Figure 1 As shown in the formula, the framework first extracts multi-level deep features from the input cardiac Cine MRI sequence through a feature encoder; on this basis, combined with a knowledge-guided mechanism, efficient transmission of valuable information between different layers of the model is realized. Subsequently, an adaptive spatio-temporal attention mechanism (ASA) is introduced to accurately capture the spatio-temporal correlation between the target frame and the adjacent frame, and through multi-scale feature aggregation, the expression ability of global and local information is enhanced; in order to further improve the effect of inter-frame information integration, the present application designs a Transformer-driven dynamic feature aggregation module (TdFA), which realizes high-quality feature representation by expanding the receptive field and establishing long-distance dependency relationship; finally, a data consistency module ensures that the reconstruction result is consistent with the original sampling data, thereby significantly improving the quality and accuracy of the reconstructed image.
[0044] More specifically, the present implementation first preprocesses the input data, which is complex cardiac cine MRI data, and these data are undersampled, which is usually used to speed up the imaging process by reducing the number of sampling points in the acquisition process, but this will cause information loss and artifacts in the image.
[0045] In order to perform unified calculation, the present application preprocesses all the reconstructed data sets. Specifically, the original data set (512x204) is zero-padded (512x256) in the frequency domain, and is cropped to a scale of 256x256 in the image domain. Since the obtained MR image is complex, the present application separates the real part and the imaginary part of the complex number, and connects them as two channels, avoiding information loss or distortion in the interpolation process. The data dimension is converted from 256x256 to 256x256x2. After converting into double-channel data, the image is normalized to keep the pixel value of the image within a uniform range ([0, 1]), which helps to reduce the impact of different MRI acquisition conditions on image quality and improve the stability and consistency of subsequent network training.
[0046] In the present implementation, in order to fully unlock the potential of MR images and better extract respective deep spatial features from different frames of videos, the present application designs a feature encoder, as shown in Figure 1 The continuous sequence of cardiac MR images (the present application uses three frames τ-1, τ, τ+1 to achieve the best reconstruction effect) is input into the feature extraction part of the network; next, the deep features of the target frame and the adjacent reference frame are extracted, denoted as Subsequently, the pre-trained VGG19 network is used to extract multi-scale features of Ref τ-1 , and three deep feature representations at different scales are obtained, denoted as The above process can be represented as follows:
[0047]
[0048] Where ψ represents deep feature extraction, (,) is used to separate multiple variables, so that the function receives different values, and the above is the extraction process of the Mentor network. The extraction process in the Learner network is the same as above.
[0049] In the present implementation, a knowledge guiding mechanism is designed, as shown in Figure 1 Specifically, Input τ and the previous frame reference frame Ref τ-1 are input into the Mentor network, and the features generated after being guided by the Mentor network and the subsequent reference frame As the input of the Learner network, the goal of this mechanism is to make full use of the prior knowledge in the Mentor network to capture more relevant information. Therefore, the Learner network receives the prior features from the Mentor network at each scale. Through the transfer learning method, the knowledge transfer from the Mentor network to the Learner network is realized, thereby maximizing the performance and generalization ability of the Learner network. The knowledge transfer process can be expressed as:
[0050]
[0051] Among them, κ represents knowledge guidance, its subscripts m and l represent the Mentor network and Learner network respectively, and Rec is the reconstruction result. In this process, the Mentor network focuses on providing knowledge guidance, while the Learner network is responsible for updating the weight parameters.
[0052] In this implementation, the τ and Ref τ-1 Multi-scale deep features Where 1≤i≤3, and Input τ Shallow features extracted by convolution Passed together to the Mentor network.
[0053] In this implementation, the Learner network receives the prior features guided by knowledge and use With Reft t+1 Extracted multi-scale features where 1≤j≤3, and from Shallow features The Learner network has a similar structure to the Mentor network, but there is a key difference: the Learner network uses a knowledge guidance mechanism to refine the knowledge from the Mentor network at each scale.
[0054] In this implementation, multi-scale feature aggregation is also performed. Taking the Learner network as an example, given the deep features and shallow features Multi-scale feature aggregation refines features at three different scales by cascading adaptive spatiotemporal attention and Transformer-driven dynamic feature aggregation; at the same time, the skip connection characteristics of the U-net network are used to iterate the multi-scale features, thereby better capturing feature information at different scales. This process is as follows Figure 1 As shown in the multi-scale feature aggregation.
[0055] In the present implementation, in order to effectively learn the spatial and temporal related features, the present application introduces a new adaptive spatio-temporal attention (ASA), as shown in Figure 2 The ASA has the ability to dynamically adapt to different local receptive fields according to different scale sizes, and can effectively realize image registration of adjacent view features, ensuring that they are spatially aligned.
[0056] Specifically, the present application utilizes the respectively as the initial Query (Q), Key (K), Value (V). By comparing the local features of Q and K, similarity information is generated, which reflects the relevance between features at different times. The generation of similarity information can be expressed as the following formula:
[0057]
[0058] wherein Corr represents the correlation matching, respectively as the correlation value and the position index, pad is the padding operation, (p1, p2, p3, p4) is a four-tuple representing the padding number in four directions, which is set to (1, 1, 1, 1) in the present application. Further, displacement information is utilized to generate optical flow. In order to further explore the spatio-temporal correlation, the present application captures the temporal correlation by means of optical flow estimation, and the generation process of optical flow information is as shown in the formula:
[0059]
[0060] wherein mod is the modulo operation, h represents the height, w represents the width, (x, y) respectively takes values from [0, w-1] and [0, h-1], and grid is used to generate the coordinates of all pixel positions. Then, based on displacement information is generated which describes the displacement of each pixel in relative to τ The generation process of O
[0061] is as follows: τ τ x y
[0062] wherein Shift represents the translation operation, (λ x ,λ y ) are in the range [0, 2], so there are n = 3 x 3 different offset versions. Next, the flow τ and V are mapped to align the features of the target frame to the spatial positions of the neighboring reference frames, so that they are aligned in the same coordinate system for accurate 3D reconstruction, as shown in Figure 2 The subsequent calculation process is as follows:
[0063]
[0064] where, denotes the concatenation operation, is a mapping operation, σ is a sigmoid activation function, and θ is a tanh activation function, and ⊕ represent matrix multiplication and addition operations, respectively. The target frame feature is one of the original inputs of the ASA, and its initial value is Using the calculated multi-view information of the front and rear frames, the spatio-temporal correlation features are modeled by using a deformable convolution network (DCN):
[0065]
[0066] In this implementation, in order to aggregate spatio-temporal features at multiple scales the target frame feature and the prior feature The present application designs a Transformer-driven dynamic feature aggregation module (TdFA) to combine feature maps from multiple frames. As shown in Figure 3 TdFA is composed of a convolution layer and n multi-geometry window Transformers (MGwinT), and the present application sets n = 4. Specifically, the feature maps are first concatenated in the spatial dimension, then a convolution is used for preliminary fusion, and then MGwinT is used to obtain the target frame feature In addition, the present application also uses a residual connection to stabilize the training process, and the specific process is as follows:
[0067]
[0068] where, Conv is a convolution operation, denotes MGwinT, is a connection operation.
[0069] In this implementation, multi-geometry window Transformers are used, and some tissues in the heart region have similarities, such as Figure 4The local similarity in the annotated areas shown in (I), (II), (III), and (VI), and the global similarity in the annotated areas shown in (IV) and (V). This means that dependencies can be established both locally and globally, thereby capturing more similar features. However, using a square sliding window or a rectangular window alone cannot establish good local and global dependencies without adjusting the window size. Therefore, the present invention further considers the combination of the two, that is, using two Transformers with different windows at the same time to further enhance the global dependencies of the network aggregation module. The present invention proposes a new multi-geometric window attention Transformer (MGwinT) that effectively establishes long-range and long-distance dependencies, while fusing the spatiotemporal information in adjacent frames, such as Figure 3 In order to reduce the computational complexity of window attention, the present invention performs two window attentions in parallel in the same MGwinT, namely square window multi-head self-attention (SW-MSA) and rectangular window multi-head self-attention (RW-MSA).
[0070] For the design of the square window (S-Window), a regular window partitioning strategy is used. Starting from the lower left corner pixel, the feature map is divided into α×α β×β windows ( Figure 4 As shown in α=2,β=6), let the regular partition window shift cyclically to the upper left. pixels, such as Figure 4 The position changes of regions (I), (II), and (III) are shown in the figure.
[0071] For the design of rectangular window (R-Window), such as Figure 4 As shown in (IV), (V), and (VI) of the figure, this paper divides these into two types: vertical windows and horizontal windows, and applies them serially to RW-MSA at different scales. By alternating between the two types of windows, global modeling can be maintained to capture more similar features without expanding the window size.
[0072] In this implementation, data consistency is introduced. In MRI reconstruction, the image quality can be effectively improved by introducing data consistency (DC) constraints in the k-space domain (frequency domain). This is because the generated image may have information loss in the frequency domain. The data consistency method uses Fourier transform to separate the image domain (Rec′, GT) from the k-space domain. Specifically, the data consistency constraint ensures that the sampled regions in k-space remain consistent with the original k-space data, while the unsampled regions retain the values estimated from the reconstruction process. This selective fidelity helps optimize the recovery of frequency domain information. From a mathematical point of view, the coefficients of the sampled regions are constrained to be consistent with the original data. match, while the coefficients of the unsampled region match the reconstruction result remain consistent, the process can be represented as the following formula:
[0073]
[0074] where η≥0 represents the noise term, [m,n] represents the matrix index operation, denotes the k-space data after applying the DC constraint, is the sampling mask, and after performing the inverse Fourier transform, the corresponding representation in the image domain can be recovered, and the detailed DC process is shown on the right side of . Figure 1
[0075] In order to construct the training data set, 27 healthy volunteers (a total of 2430 images, each volunteer contains 6 groups of slices, 15 time frames for a group) were randomly selected for full-sampled short-axis cardiac MR dynamic images. In the training process, the full-sampled data was retrospectively undersampled by a random Cartesian undersampling mask combined with four different acceleration multiples to generate the initial input (Input / Ref) of the network. In order to evaluate the difference between the reconstruction result Rec and the gold standard GT and improve the reconstruction quality, the present application uses three loss functions, L1-pixel loss, perceptual loss and adversarial loss, to optimize the model.
[0076] In order to evaluate the performance of the network proposed in the present application, three cardiac cine GRAPPA reconstruction data sets were used, including one in-house data set (short-axis direction) and two CMR x Recon data sets (short-axis direction and long-axis direction). At the same time, in order to ensure the safety and privacy during data use, all data were desensitized and anonymized before use to reduce the risk of patient identity information leakage.
[0077] The method proposed in the present application is implemented on an NVIDIA A100 GPU (40GB) workstation based on the PyTorch framework. The optimization algorithm uses Adam, and the learning rate is set to 1x10 -4 , the momentum attenuation coefficient is set to [0.9, 0.999], and the batch size is set to 2. In addition, two hyperparameters in the training loss: λ1=1e -4 , λ2=1e -6 . Considering that the cardiac cine has high similarity between adjacent frames, the input frame number is set to 3, which helps the model better capture the time correlation and inter-frame similarity information.
[0078] As shown in Tables 1, 2, 3, 4, 5, 6, the application compares the quantitative performance of KGMgT with other reconstruction methods on in-house data sets and CMR x Recon data sets. The results show that KGMgT achieves the best performance on all indicators, whether short-axis cine or long-axis cine, and shows significant competitiveness compared with the current most advanced method. This shows that the method proposed in the application has made a breakthrough in modeling the heart cycle movement of the heart, and has maximized the recovery of motion information through spatial and temporal perception capabilities.
[0079] Table 1: Quantitative comparison with SOTA on in-house (4x, 6x) short-axis (Sax) data sets, the best performance is highlighted in bold, and the evaluation indicators include PSNR (dB), SSIM and NMSE (x10 -1 ).
[0080]
[0081] Table 2: Quantitative comparison with SOTA on in-house (8x, 10x) short-axis (Sax) data sets, the best performance is highlighted in bold, and the evaluation indicators include PSNR (dB), SSIM and NMSE (x10 -1 ).
[0082]
[0083]
[0084] Table 3: Quantitative comparison with SOTA on CMR x Recon (4x, 6x) short-axis (Sax*) data sets, the best performance is highlighted in bold, and the evaluation indicators include PSNR (dB), SSIM and NMSE (x10 -1 ).
[0085]
[0086] Table 4: Quantitative comparison with SOTA on CMR x Recon (8x, 10x) short-axis (Sax*) data sets, the best performance is highlighted in bold, and the evaluation indicators include PSNR (dB), SSIM and NMSE (x10 -1 ).
[0087]
[0088] Table 5: Quantitative comparison with SOTA on CMR x Recon (4x, 6x) long-axis (Lax*) data sets, the best performance is highlighted in bold, and the evaluation indicators include PSNR (dB), SSIM and NMSE (x10-1 )。
[0089]
[0090]
[0091] Table 6: Quantitative comparison with SOTA on CMR x Recon (8x, 10x) long axis (Lax*) data sets, best performance highlighted in bold, evaluation metrics include PSNR (dB), SSIM and NMSE (x10 -1 )。
[0092]
[0093] Compared with the traditional method, the peak signal-to-noise ratio (PSNR) of the reconstructed image is increased by about 0.82-5.83dB, the structural similarity (SSIM) is increased by about 0.03-0.3, and the maximum reduction of the normalized mean square error (NMSE) reaches 3.8297 (x10 -1 ), which significantly improves the clarity and structural fidelity of the image. This indicates that the model can more accurately recover the details of the cardiac cine MRI image during the reconstruction process, avoiding the blurring phenomenon in the traditional method.
[0094] Compared with the existing method, the model can reduce about 10%-15% of the inter-frame information loss, thereby more comprehensively integrating the spatio-temporal information and improving the continuity and accuracy of the reconstructed image.
[0095] Compared with the reconstruction technology based on the traditional optimization method, the present application can achieve higher image quality under the same acceleration factor (4x, 6x, 8x, 10x), and the calculation speed of the reconstruction process is increased by about 15%. This means that the image reconstruction task can be completed more quickly under the same computing resources, improving the efficiency of clinical diagnosis.
[0096] Due to the adoption of the ASA mechanism and the multi-geometry window strategy, the present application can effectively reduce the artifacts caused by cardiac motion; compared with the traditional method, the artifacts are significantly reduced by about 20%-30%, improving the usability and accuracy of the image, especially in the case of rapid cardiac motion, the key structural information can be well preserved.
[0097] By reducing the acquisition time and improving the image quality, the present application can achieve higher accuracy in cardiac MRI image reconstruction, and shorten the scanning time for each patient (for example, with 10 times acceleration, the original scanning time of 10 minutes can be shortened to 1 minute), reducing the discomfort of the patient caused by long-time scanning, especially suitable for emergency or rapid diagnosis scenarios, improving the efficiency of clinical operation.
[0098] It can be understood that in other implementations, the key parameters used in the present application can be dynamically adjusted to achieve various functions to adapt to different requirements, parameter settings of different models, including but not limited to: adjustment of network depth, weight initialization method or optimization function; different data acquisition, preprocessing and post-processing methods, such as image enhancement methods or data interpolation methods in the image field; the present application is aimed at a specific field (medical imaging), which can be extended to other similar fields (such as industrial detection, remote sensing images).
[0099] In the present application, the spatio-temporal attention mechanism (ASA) is used to accurately capture the spatio-temporal correlation between the target frame and the adjacent frame. It can be understood that in other implementations, if other technologies are used to achieve this function, the following alternatives can be considered: the graph neural network models the relationship between frames by propagating information on the graph structure, each frame is regarded as a node in the graph, and the connection weight between nodes is obtained through spatio-temporal feature learning. This method is suitable for cases where the relationship between frames in the data is complex or nonlinear.
[0100] It can be understood that in other implementations, dilated convolution (Dilated Convolution) is used to implement feature aggregation. Instead of using traditional convolution operations, dilated convolution is used to expand the receptive field by introducing a dilated expansion in the convolution kernel, which improves the ability to capture long-range dependencies. Through the stacking of dilated convolution modules, the receptive field range can also be improved.
[0101] Embodiment 2:
[0102] As shown in Figure 6 The present implementation provides a cardiac cine MRI reconstruction system, which comprises:
[0103] A preprocessing unit configured to preprocess the acquired cardiac cine MRI data to obtain a cardiac cine MRI sequence;
[0104] A Mentor network guiding unit configured to take the current frame image and the adjacent previous frame image as inputs of the Mentor network to obtain guided features generated after Mentor network guidance;
[0105] A Learner network guiding unit configured to take the guided features and the adjacent next frame image as inputs of the Learner network, and obtain a reconstruction result after Learner network guidance based on the prior features in the Mentor network, the guided features and the adjacent next frame image.
[0106] The specific working methods of the above units are described in Embodiment 1, which will not be repeated here.
[0107] It can be understood that each of the above units can be combined into one or several other units to constitute, or some of the units can be further split into a plurality of units with smaller functions to constitute, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In actual application, the function of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the system can also include other units. In actual application, these functions can also be assisted by other units, and can be implemented by multiple units in cooperation.
[0108] According to another embodiment of the present application, the system described in the embodiment can be constructed and the method of the embodiment 1 of the present application can be implemented by running a computer program (including program codes) capable of performing each step involved in the corresponding method described in the embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM), etc., the computer program can be recorded on a computer readable recording medium, loaded into the above computing device through the computer readable recording medium, and run therein.
[0109] Embodiment 3:
[0110] As shown in Figure 7 The present implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer readable storage medium 1003. The processor 1001, the communication interface 1002, and the computer readable storage medium 1003 can be connected through a bus or other means.
[0111] The communication interface 1002 is configured to receive and send data, the computer readable storage medium 1003 can be stored in the memory of the electronic device, the computer readable storage medium 1003 is configured to store a computer program, the computer program includes program instructions, and the processor 1001 is configured to execute the program instructions stored in the computer readable storage medium 1003.
[0112] The processor 1001 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device, which is suitable for implementing one or more instructions, and is particularly suitable for loading and executing one or more instructions to implement a corresponding method flow or a corresponding function.
[0113] The processor 1001 is configured to perform the following process:
[0114] The acquired cardiac cine MRI data is preprocessed to obtain a cardiac Cine MRI sequence;
[0115] The current frame image and the adjacent previous frame image are input into the Mentor network to obtain guided features generated after being guided by the Mentor network;
[0116] The guided features and the adjacent next frame image are input into the Learner network, and the reconstruction result is obtained after being guided by the Learner network according to the prior features in the Mentor network, the guided features, and the adjacent next frame image.
[0117] The specific working method is described in Embodiment 1, which will not be repeated here.
[0118] Embodiment 4:
[0119] The present implementation provides a computer readable storage medium (Memory), which is a memory device in an electronic device, used to store programs and data. It can be understood that the computer readable storage medium herein can include a built-in storage medium in the electronic device, and of course can also include an expansion storage medium supported by the electronic device. The computer readable storage medium provides a storage space, which stores a processing system of the electronic device.
[0120] And in the storage space, there is also one or more instructions suitable for being loaded and executed by the processor, which can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium herein can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory; optionally, it can also be at least one computer readable storage medium located away from the aforementioned processor.
[0121] In one embodiment, the computer readable storage medium stores one or more instructions; the processor loads and executes the one or more instructions stored in the computer readable storage medium to implement the following process:
[0122] The acquired cardiac cine MRI data is preprocessed to obtain a cardiac Cine MRI sequence;
[0123] The current frame image and the adjacent previous frame image are input into the Mentor network to obtain guided features generated after being guided by the Mentor network;
[0124] The guide feature and an adjacent next frame image are taken as inputs of a Learner network, and a reconstruction result is obtained through the Learner network according to the prior feature in the Mentor network, the guide feature and the adjacent next frame image.
[0125] The specific working method is described in Example 1, which will not be repeated here.
[0126] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method of cardiac cine MRI reconstruction, characterized by, The method comprises the following steps: preprocessing the acquired cardiac cine MRI data to obtain a cardiac Cine MRI sequence; the current frame image and the adjacent previous frame image are input into the Mentor network to obtain guided features generated after being guided by the Mentor network; the guided features and the adjacent next frame image are input into the Learner network, and the reconstruction result is obtained after being guided by the Learner network according to the prior features in the Mentor network, the guided features and the adjacent next frame image. In the Mentor network, a feature encoder is used to extract multi-scale features of the current frame image and the adjacent previous frame image, the multi-scale features are used as query vectors, key vectors and value vectors of adaptive spatio-temporal attention, and the spatio-temporal correlation features are obtained by combining the target frame features of the current frame image. In the Learner network, a feature encoder is used to extract multi-scale features of the current frame image and the adjacent next frame image, the multi-scale features are used as query vectors, key vectors and value vectors of adaptive spatio-temporal attention, and the spatio-temporal correlation features are obtained by combining the target frame features of the current frame image. The multi-scale features are used as query vectors, key vectors and value vectors of adaptive spatio-temporal attention, and the spatio-temporal correlation features are obtained by combining the target frame features of the current frame image, comprising: By comparing the local features of the query vector and the key vector, similarity information is generated, which reflects the relevance between features at different times. The generation of the similarity information is represented as: , wherein, represents a relevance match, , are a relevance value and a position index, respectively, is a padding operation, is a four-tuple representing the number of padding in four directions, and are the query vector and the key vector, respectively. Temporal correlation is captured by means of optical flow estimation, resulting in optical flow information Displacement information is generated based on the optical flow information The optical flow information And the value vector Using a mapping operation, the features of the target frame are mapped to spatial positions aligned with the adjacent reference frame, so that they are aligned in the same coordinate system, and the spatio-temporal correlation features are modeled by using a deformable convolution network with the calculated multi-view information of the front and rear frames. 2. The cardiac cine MRI reconstruction method of claim 1, wherein In the Mentor network, a Transformer-driven dynamic feature aggregation module is used to aggregate the spatio-temporal correlation features and the target frame features at multiple scales to obtain target frame features at the next scale, and the guided features are obtained after denoising encoder and data consistency processing.
3. The cardiac cine MRI reconstruction method of claim 1, wherein In the Learner network, a Transformer-driven dynamic feature aggregation module is used to aggregate the spatio-temporal correlation features, the target frame features and the prior features of the Mentor network at multiple scales to obtain target frame features at the next scale, and the reconstruction image is obtained after denoising encoder and data consistency processing.
4. The cardiac cine MRI reconstruction method of claim 2 or 3, wherein In the Transformer-driven dynamic feature aggregation module, a convolution layer and a plurality of multi-geometry window Transformers are included, the spatio-temporal correlation features, the target frame features and the prior features of the Mentor network at multiple scales are concatenated in the spatial dimension, the convolution layer is used for preliminary fusion, and the multi-geometry window Transformers are used to obtain target frame features at the next scale.
5. A cardiac cine MRI reconstruction system, characterized by, The cardiac cine MRI reconstruction method of any one of claims 1-4 comprises: a preprocessing unit configured to preprocess the acquired cardiac cine MRI data to obtain a cardiac Cine MRI sequence; The Mentor network guiding unit is configured to take a current frame image and a neighboring previous frame image as inputs of the Mentor network, and obtain guided features generated after Mentor network guidance; The Learner network guiding unit is configured to take the guided features and a neighboring next frame image as inputs of the Learner network, and obtain a reconstruction result after Learner network guidance according to prior features in the Mentor network, the guided features and the neighboring next frame image.
6. A computer device, comprising: Comprise: a processor and a computer readable storage medium; a processor adapted to execute a computer program; a computer readable storage medium having stored therein a computer program, which, when executed by the processor, implements the cardiac cine MRI reconstruction method according to any one of claims 1 to 4.
7. A computer readable storage medium characterized by The computer readable storage medium stores a computer program, which is adapted to be loaded and executed by the processor to implement the cardiac cine MRI reconstruction method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Video super-resolution reconstruction method and system based on multi-scale local self-attention
CN115082308A
Lightweight video super-resolution reconstruction method based on hybrid space-time convolution
CN117830095A