Heart motion estimation method and equipment based on motion decomposition
By constructing a dual-branch feature encoder and feature hybrid decoder network and combining a multi-head cross-attention mechanism for cardiac motion estimation, the problems of poor adaptability and single deformation representation of existing methods in low signal-to-noise ratio image data are solved, and high accuracy and robustness of cardiac motion estimation are achieved.
Patent Information
- Application Number
- CN202610097504.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-15
AI Technical Summary
Existing cardiac medical image motion estimation methods have poor adaptability to low signal-to-noise ratio image data and low consistency of results across different modalities, leading to inaccurate cardiac motion estimation. Furthermore, existing deep learning-based methods only use a single deformation representation, making it difficult to accurately model cardiac motion information.
A cardiac motion estimation method based on motion decomposition is adopted. By constructing a dual-branch feature encoder network and a feature fusion decoder network, stationary and moving images are processed asynchronously. A feature fusion transformer module and a multi-head cross-attention mechanism are used for deformation decomposition and adaptive learning. A total loss function is constructed for training.
It improves the accuracy and robustness of cardiac motion estimation, can better handle complex motion information, enhances the ability to encode cardiac motion information, and improves image coding performance.
Smart Images

Figure CN122048980A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing technology, and more specifically relates to a method and device for cardiac motion estimation based on motion decomposition. Background Technology
[0002] Existing motion estimation methods for cardiac medical images generally suffer from technical drawbacks such as poor adaptability to low signal-to-noise ratio image data and low consistency of results across different modalities, leading to inaccurate cardiac motion estimation. Therefore, there is an urgent need to develop a high-precision, robust, and multimodal-adaptive motion estimation method for cardiac medical images to overcome these technical bottlenecks.
[0003] In recent years, deep learning-based registration methods have demonstrated excellent performance in addressing the aforementioned problems. After data-driven training, deep learning-based registration methods can predict deformations in new data, achieving both speed and high accuracy. Among these, unsupervised registration algorithms, because they do not rely on manually labeled anatomical structures, can save significant manpower and resources and are applicable to medical image data that is difficult to label, showing great promise for applications in cardiac motion estimation.
[0004] The aforementioned methods all rely on a single deformation representation, establishing a single mapping relationship between image features and the predicted deformation field. This approach has limitations when processing cardiac medical images with complex motion information. Specifically, these methods use only a single vector to represent displacement at a specific spatial location, neglecting the fact that cardiac motion is composed of multiple deformations with different amplitudes, directions, and sources. Simply representing cardiac motion makes it difficult to accurately model cardiac motion information. Therefore, a registration method specifically for cardiac motion estimation is urgently needed to improve its accuracy. Summary of the Invention
[0005] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a cardiac motion estimation method and device based on motion decomposition, the purpose of which is to improve the accuracy of motion estimation of cardiac medical images.
[0006] To achieve the above objectives, the present invention provides a method and device for cardiac motion estimation based on motion decomposition.
[0007] The above-mentioned objectives of the present invention are achieved through the following technical means:
[0008] A cardiac motion estimation method based on motion decomposition includes the following steps:
[0009] Step S1: Acquire paired moving images and fixed image As samples, training and test sets are constructed based on the samples;
[0010] Step S2: Construct a cardiac motion estimation network, which includes a dual-branch feature encoder network and a feature fusion decoder network.
[0011] In a dual-branch feature encoder network, the fixed image is processed asynchronously. and moving images Processing is performed to obtain a fixed image. The corresponding first-level fixed image coding features of the pyramid structure ~Level 5 Fixed Image Coding Features and obtaining moving images Corresponding Level 1 moving image coding features ~Level 5 moving image coding features ;
[0012] In the feature-mixing decoder network, the first-level fixed image encoding features are... ~Level 5 Fixed Image Coding Features and Level 1 moving image coding features ~Level 5 moving image coding features The deformation field is fused and output. ;
[0013] Step S3: Construct the total loss function;
[0014] Step S4: Train the cardiac motion estimation network based on the training set;
[0015] Step S5: Input the pair of moving and stationary images to be predicted into the trained cardiac motion estimation network to obtain the estimated deformation field.
[0016] As described above, the dual-branch feature encoder network includes a 3×3 convolutional module, a first single-step encoding module, a second single-step encoding module, a downsampling module, and a third single-step encoding module; fixed image and moving images Asynchronous input features are fed into a 3×3 convolutional module. The output features of the 3×3 convolutional module are fed into the first single-step encoding module and the downsampling module, respectively. The output features of the first single-step encoding module are fed into the second single-step encoding module, and the output features of the downsampling module are fed into the third single-step encoding module.
[0017] Fixed image When input into the dual-branch feature encoder network: the output features of the 3×3 convolution module, the output features of the downsampling module, the output features of the first single-step encoding module, the output features of the third single-step encoding module, and the output features of the second single-step encoding module are respectively used as the first-level fixed image encoding features. Level 2 fixed image coding features Level 3 fixed image coding features Level 4 fixed image coding features and Level 5 fixed image coding features ;
[0018] Moving images When input into the dual-branch feature encoder network: the output features of the 3×3 convolutional module, the output features of the downsampling module, the output features of the first single-step encoding module, the output features of the third single-step encoding module, and the output features of the second single-step encoding module are sequentially used as the first-level motion image encoding features. Level 2 moving image coding features Level 3 moving image coding features Level 4 moving image coding features and Level 5 moving image coding features .
[0019] As described above, the first, second, and third single-step encoding modules each include a block convolution module, an instance regularization layer, a convolutional layer, an instance regularization layer, and multiple Swin transformer modules. The image features input to the first, second, and third single-step encoding modules are first processed by the block convolution module to obtain image-reduced block codes. The image-reduced block codes are then processed sequentially by the instance regularization layer, the convolutional layer, and the instance regularization layer. After being flattened, they are processed sequentially by multiple Swin transformer modules to obtain the encoded features as output features.
[0020] The stride of the block convolution modules in the first, second, and third single-step encoding modules, as described above, is denoted as... The kernel size of the block convolution module Filling of block convolution modules , To round down, the step size of the downsampling module is set to... .
[0021] As described above, the feature mixing decoder network includes a first-level feature mixing transformer module, a second-level feature mixing transformer module, a third-level feature mixing transformer module, a fourth-level feature mixing transformer module, and a fifth-level feature mixing transformer module;
[0022] In the feature fusion decoder network:
[0023] Level 5 Fixed Image Coding Features and Level 5 moving image coding features Hybrid features are obtained after channel splicing. Mixed features After channel dimensionality reduction using the corresponding 3×3 convolutional module, the hybrid features are obtained. Mixed features Level 5 fixed image coding features and Level 5 moving image coding features The input is fed into the 5th-level feature mixing transformer module, which outputs the intermediate deformation field. and mixed features ,
[0024] definition For intermediate deformation fields Upsampling is performed to obtain the upsampled deformation field. Using upsampled deformation field For the first Hierarchical moving image coding features Features are obtained by twisting ,feature Mixed features and the Hierarchical fixed image coding features After channel concatenation, the features are obtained by channel dimensionality reduction using a corresponding 3×3 convolutional module. ,feature , No. Hierarchical fixed image coding features and characteristics Enter the number Hierarchical Feature Hybrid Transformer Module, No. The hierarchical feature hybrid transformer module predicts intermediate temporary deformation fields. Intermediate temporary deformation field and upsampling deformation field By merging, an intermediate deformation field can be obtained. , No. The hierarchical feature blending transformer module updates the output blended features. , The deformation field is the final output of the feature mixing decoder network. .
[0025] As described above, in the 5th level feature mixing transformer module: the 5th level is a fixed image coding feature. and Level 5 moving image coding features The query matrix is obtained by expanding each matrix separately and then performing linear projection on each matrix separately. Bond matrix Mixed features After being processed by a parameter-regularized convolutional head and unfolded into a value matrix, Query matrix Key matrix and value matrix constitute matrix, The matrix is input to each windowed multi-head cross-attention mechanism unit for computation, resulting in the sub-deformation field of the deformation component corresponding to each windowed multi-head cross-attention mechanism unit. , The sequence number of the windowed multi-head cross-attention mechanism unit, sub-deformation field The input sub-component weighted summation module obtains weights corresponding one-to-one with the deformation components through adaptive learning using a corresponding multilayer perceptron. Sub-deformation field and weight The intermediate temporary deformation field is calculated by weighted summation. Mixed features The image is obtained by upsampling using the image patch expansion module. ;
[0026] In the 4th to 1st level feature hybrid transformer module: Hierarchical fixed image coding features and characteristics The query matrix is obtained by expanding each matrix separately and then performing linear projection on each matrix separately. Bond matrix Mixed features After being processed by a parameter-regularized convolutional head and unfolded into a value matrix, Query matrix Key matrix and value matrix constitute matrix, The matrix is input to each windowed multi-head cross-attention mechanism unit for computation, resulting in the sub-deformation field of the deformation component corresponding to each windowed multi-head cross-attention mechanism unit. , The sequence number of the windowed multi-head cross-attention mechanism unit, sub-deformation field The input sub-component weighted summation module obtains weights corresponding one-to-one with the deformation components through adaptive learning using a corresponding multilayer perceptron. Sub-deformation field and weight The intermediate temporary deformation field is calculated by weighted summation. Mixed features The image is obtained by upsampling using the image patch expansion module. .
[0027] Constructing the total loss function as described above involves the following steps:
[0028] Using deformation field For moving images Distorted image vs. fixed image Calculate similarity loss based on deformation field The smoothness loss is calculated using the L2 norm of the gradient, and the weighted sum of the similarity loss and the smoothness loss is used as the total loss function.
[0029] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the above-described cardiac motion estimation method based on motion decomposition.
[0030] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described cardiac motion estimation method based on motion decomposition.
[0031] A computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described motion decomposition-based cardiac motion estimation method.
[0032] In summary, the above-described technical solutions conceived in this invention can achieve the following beneficial effects:
[0033] (1) The dual-branch coding structure designed in this invention re-divides image features at multiple different image sizes, which helps the network learn anatomical structures in medical images from multiple different scales. This avoids the problem that the fixed division method in the general block coding process restricts the network from processing multi-level semantic information and recognizing specific anatomical structures, thereby strengthening the network's ability to encode cardiac motion information and improving the accuracy of cardiac motion estimation.
[0034] (2) The specially designed block convolutional unit covers the entire image block region through a sufficiently large convolutional kernel, realizing information interaction within the image block. This avoids the problem that conventional transformer-based methods only consider global semantic correlation and ignore information interaction within the image block, thus strengthening the network's ability to model local and global semantic information of the image and improving the overall image coding performance.
[0035] (3) Feature Hybrid Transformer Module: Based on the cross-attention mechanism, the fixed and moving images are used as weights, and the fused features of the two are used as the initial deformation field. The initial deformation field is optimized by using the weight coefficients. At the same time, a multi-head mechanism is introduced to realize the decomposition and adaptive learning of deformation. Attached Figure Description
[0036] Figure 1 This is a flowchart of the present invention;
[0037] Figure 2 This is a schematic diagram of the structure of the dual-branch feature encoder network of the cardiac motion estimation network in an embodiment of the present invention.
[0038] Figure 3 This is a schematic diagram of the feature hybrid decoder network of the cardiac motion estimation network in an embodiment of the present invention;
[0039] Figure 4 The following are visualization results of the method and other image registration-based methods in the embodiments of the present invention on the public dataset ACDC; (a1) represents the fixed image; (a2) to (a9) represent the superimposed images after the moving image and moving image label are distorted by the deformation fields predicted by the present invention, SyN, VoxelMorph, TransMorph, TransMatch, NICE-Trans, ModeT, and AutoFuse-Trans respectively; (b1) represents the moving image; (b2) to (b9) represent the deformation fields predicted by the present invention, SyN, VoxelMorph, TransMorph, TransMatch, NICE-Trans, ModeT, and AutoFuse-Trans respectively, wherein the displacements in the x, y, and z directions are mapped to the three channels of the RGB image respectively; (c1) represents the difference map between the moving image label and the fixed image label; (c2) to (c9) represent the difference map between the label after the moving image label is distorted by the deformation fields predicted by each method and the label of the fixed image.
[0040] Figure 5 The following are visualization results of the method and other image registration-based methods in the embodiments of the present invention on the public dataset CMRxMotion. (a1) represents the fixed image; (a2) to (a9) represent the superimposed images after the moving image and moving image label are distorted by the deformation fields predicted by the present invention, SyN, VoxelMorph, TransMorph, TransMatch, NICE-Trans, ModeT, and AutoFuse-Trans, respectively; (b1) represents the moving image; (b2) to (b9) represent the deformation fields predicted by the present invention, SyN, VoxelMorph, TransMorph, TransMatch, NICE-Trans, ModeT, and AutoFuse-Trans, respectively, wherein the displacements in the x, y, and z directions are mapped to the three channels of the RGB image, respectively; (c1) represents the difference map between the moving image label and the fixed image label; (c2) to (c9) represent the difference map between the label after the deformation fields predicted by each method distort the moving image label and the fixed image label. Detailed Implementation
[0041] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0042] Example 1
[0043] like Figure 1 As shown in the figure, this embodiment of the invention provides a cardiac motion estimation method based on motion decomposition, which mainly includes the following steps:
[0044] Step S1: Obtain case data. Each case data set contains a pair of images, one of which is used as a fixed image. Another image as a moving image Moving images in pairs and fixed image As samples, training and test sets are constructed based on the samples;
[0045] Step S2: Construct a cardiac motion estimation network based on motion decomposition and adaptive learning. The cardiac motion estimation network includes a dual-feet encoder network (DFEnc) and a feature mixer decoder network (FMDec).
[0046] The Dual-feet Encoder (DFEnc) network consists of a 3×3 convolutional module, a first branch, and a second branch. The first branch comprises a first single-step encoding module and a second single-step encoding module connected in sequence. The second branch comprises a downsampling module and a third single-step encoding module connected in sequence.
[0047] The first through third single-step encoding modules all include setting the step size. kernel size ,filling The block convolution module, the stride of the downsampling module is set to .
[0048] In a dual-branch feature encoder network, the fixed image is processed asynchronously. and moving images Processing is performed to obtain a fixed image. The corresponding first-level fixed image coding features of the pyramid structure ~Level 5 Fixed Image Coding Features and obtaining moving images Corresponding Level 1 moving image coding features ~Level 5 moving image coding features .
[0049] Fixed image When inputting into a dual-branch feature encoder network: fixed image The input features are fed into a 3×3 convolutional module, and the output features of the 3×3 convolutional module are fed into the first and second branches, respectively. In the first branch: the output features of the 3×3 convolutional module are fed into the first single-step encoding module, and the output features of the first single-step encoding module are fed into the second single-step encoding module. In the second branch: the output features of the 3×3 convolutional module are fed into a downsampling module, and the output features of the downsampling module are fed into the third single-step encoding module. The output features of the 3×3 convolutional module, the downsampling module, the first single-step encoding module, the third single-step encoding module, and the second single-step encoding module are sequentially used as the first-level fixed image encoding features. Level 2 fixed image coding features Level 3 fixed image coding features Level 4 fixed image coding features and Level 5 fixed image coding features .
[0050] Moving images When input into a dual-branch feature encoder network: moving image The input features are fed into a 3×3 convolutional module, and the output features of the 3×3 convolutional module are fed into the first and second branches, respectively. In the first branch: the output features of the 3×3 convolutional module are fed into the first single-step encoding module, and the output features of the first single-step encoding module are fed into the second single-step encoding module. In the second branch: the output features of the 3×3 convolutional module are fed into a downsampling module, and the output features of the downsampling module are fed into the third single-step encoding module. The output features of the 3×3 convolutional module, the downsampling module, the first single-step encoding module, the third single-step encoding module, and the second single-step encoding module are sequentially used as the first-level motion image encoding features. Level 2 moving image coding features Level 3 moving image coding features Level 4 moving image coding features and Level 5 moving image coding features .
[0051] This allows the image coding features generated in the first and second branches to complement each other, avoiding the problem of sparse feature size. Combining the image coding features generated in the first and second branches can yield complete and continuous resolution image coding features.
[0052] The first, second, and third single-step encoding modules all include setting the step size. kernel size ,filling Block convolution module, To round down, the input image features are re-divided into image blocks to obtain block image features; each also includes a convolutional kernel with a size of [missing value]. The convolutional layers are used to adjust the number of channels in the block image features; each also includes several Swing transformer modules. The image features input to the single-step encoding module first pass through the block convolution module to obtain block codes with reduced image size. The block codes with reduced image size then pass through the instance regularization layer, convolutional layer, and instance regularization layer in sequence to adjust the number of channels and data regularization. After being flattened, they are used as the input to the subsequent Swing transformer modules. After being processed by multiple Swing transformer modules in sequence, the encoded features are obtained as the output features of the single-step encoding module.
[0053] The feature fusion decoder network consists of five layers of feature fusion transformer modules: the first layer, the second layer, the third layer, the fourth layer, and the fifth layer. By fusing image features of different sizes, it achieves deformation prediction from coarse to fine.
[0054] In the feature fusion decoder network:
[0055] For the 5th level feature mixing transformer module: 5th level fixed image coding features and Level 5 moving image coding features Hybrid features are obtained after channel splicing. Mixed features After channel dimensionality reduction using the corresponding 3×3 convolutional module, the hybrid features are obtained. Mixed features Level 5 fixed image coding features and Level 5 moving image coding features The input is fed into the 5th-level feature mixing transformer module, which outputs the intermediate deformation field. and mixed features ,
[0056] For the feature hybrid transformer module of levels 4-1: Definition For intermediate deformation fields Upsampling (quadratic linear interpolation makes the intermediate deformation field The resolution is doubled, and the value is multiplied by 2 to obtain the upsampled deformation field. Using upsampled deformation field For the first Hierarchical moving image coding features Features are obtained by twisting ,feature Mixed features and the Hierarchical fixed image coding features After channel concatenation, the features are obtained by channel dimensionality reduction using a corresponding 3×3 convolutional module. ,feature , No. Hierarchical fixed image coding features and characteristics Enter the number Hierarchical Feature Hybrid Transformer Module, No. The hierarchical feature hybrid transformer module predicts intermediate temporary deformation fields. Intermediate temporary deformation field and upsampling deformation field By merging, an intermediate deformation field can be obtained. , No. The hierarchical feature blending transformer module updates the output blended features. . The deformation field is the final output of the feature mixing decoder network. .
[0057] Each level of feature hybrid transformer module includes a windowed multi-head cross-attention mechanism unit, a sub-component weighted summation unit, and an image patch expansion unit.
[0058] Among them, the windowed multi-head cross-attention mechanism unit is used to perform deformation decomposition. When calculating the cross-attention mechanism based on the window, the multi-head features are preserved, and a single head is used to represent a class of motion modes in order to obtain the sub-deformation field after motion decomposition.
[0059] In the 5th level feature hybrid transformer module: 5th level fixed image coding features and Level 5 moving image coding features The query matrix is obtained by expanding each matrix separately and then performing linear projection on each matrix separately. Bond matrix Mixed features After being processed by a parameter-regularized convolutional head and unfolded into a value matrix, Query matrix Key matrix and value matrix constitute matrix, The matrix is input to each windowed multi-head cross-attention mechanism unit for computation, resulting in the sub-deformation field of the deformation component corresponding to each windowed multi-head cross-attention mechanism unit. , The sequence number of the windowed multi-head cross-attention mechanism unit, sub-deformation field The input sub-component weighted summation module obtains weights corresponding one-to-one with the deformation components through adaptive learning using a corresponding multilayer perceptron. Sub-deformation field and weight The intermediate temporary deformation field is calculated by weighted summation. Mixed features The image is obtained by upsampling using the image patch expansion module. ;
[0060] In the 4th to 1st level feature mixing transformer module: Define , No. Hierarchical fixed image coding features and characteristics The query matrix is obtained by expanding each matrix separately and then performing linear projection on each matrix separately. Bond matrix Mixed features After being processed by a parameter-regularized convolutional head and unfolded into a value matrix, Query matrix Key matrix and value matrix constitute matrix, The matrix is input to each windowed multi-head cross-attention mechanism unit for computation, resulting in the sub-deformation field of the deformation component corresponding to each windowed multi-head cross-attention mechanism unit. , The sequence number of the windowed multi-head cross-attention mechanism unit, sub-deformation field The input sub-component weighted summation module obtains weights corresponding one-to-one with the deformation components through adaptive learning using a corresponding multilayer perceptron. Sub-deformation field and weight The intermediate temporary deformation field is calculated by weighted summation. Mixed features The image is obtained by upsampling using the image patch expansion module. .
[0061] The sub-component weighted summation unit is used to perform adaptive sub-deformation learning. It uses a multi-head perceptron to adaptively learn the influence of each component of the sub-deformation field on the actual deformation, assigns corresponding weights to each component, and obtains the deformation field through adaptive weighted summation.
[0062] The image block expansion unit is used to perform image feature upsampling.
[0063] This enables an explicit mapping between image features and the predicted deformation field.
[0064] Step S3: Construct the total loss function using the deformation field. For moving images Distorted image vs. fixed image Calculate similarity loss based on deformation field The smoothness loss is calculated using the L2 norm of the gradient, and the total loss function is a weighted sum of the similarity loss and the smoothness loss.
[0065] Step S4: Train the cardiac motion estimation network based on the training set. In this embodiment, the Adam optimizer is used with a learning rate of 1×10⁻⁶. -4 The training consisted of 200 rounds, with batch sizes of 8 and 4 for the ACDC and CMRxMotion datasets, respectively. The weight of the smoothness loss in the loss function was set to 0.01.
[0066] Step S5: Input the pair of moving and stationary images to be predicted into the trained cardiac motion estimation network to obtain the estimated deformation field.
[0067] Example 2
[0068] This invention provides a cardiac motion estimation method based on motion decomposition, comprising:
[0069] A pair of cardiac medical images to be estimated (a moving image and a fixed image, both with manually segmented labels, defined as moving image label and fixed image label, respectively) are input into a trained cardiac motion estimation network to obtain the estimated deformation field. This trained cardiac motion estimation network is the one trained according to Example 1.
[0070] The relevant solutions are described in Example 1 and will not be repeated here.
[0071] To verify the effectiveness of the method of the present invention, in this embodiment, the case data used are the publicly available cardiac cine magnetic resonance imaging dataset ACDC (hereinafter referred to as the ACDC dataset) and the CMRxMotion dataset. The ACDC dataset contains case data of 150 patients, with each patient having 12-35 frames of image data, including manually segmented labels for end-systolic and end-diastolic phases. After resampling and cropping, the image size of each frame is 128×128×32, with a resolution of 1.5×1.5×3.15 mm. 3The CMRxMotion dataset contains image data from 39 volunteers. Each volunteer had image data collected for four types of motion, along with corresponding hand-segmented labels. The image data was cropped to a size of 256×256×16.
[0072] The cardiac motion estimation network constructed based on the above embodiments was compared with other state-of-the-art methods on the CDC and CMRxMotion datasets. Quantitative comparisons were performed using the Dice Similarity Coefficient (DSC) and the Jacobian Determinant (Jac). To verify the effectiveness of this method, it was compared with seven commonly used state-of-the-art methods; the experimental results are shown in Tables 1 and 2. Figure 4 and Figure 5 As shown in Tables 1 and 2, seven commonly used advanced methods are Symmetric Normalization (SyN), VoxelMorph, TransMorph, TransMatch, NICE-Trans (Non-iterative Coarse-to-fine Transformer), ModeT (Motion Decomposition Transformer), and AutoFuse-Trans (Automatic Fusion Transformer). In Tables 1 and 2, RV, Myo, LV, and Avg. represent the DSC results and average values for the right ventricle, myocardium, and left ventricle, respectively. This method represents the method proposed in the embodiments of this invention. Bold black text indicates optimal performance under this evaluation metric. Figure 4 and Figure 5In the diagram, (a1) represents the fixed image; (a2) to (a9) represent the superimposed images after the deformation fields predicted by the present invention, SyN, VoxelMorph, TransMorph, TransMatch, NICE-Trans, ModeT, and AutoFuse-Trans are distorted respectively; (b1) represents the moving image; (b2) to (b9) represent the deformation fields predicted by the present invention, SyN, VoxelMorph, TransMorph, TransMatch, NICE-Trans, ModeT, and AutoFuse-Trans respectively, wherein the displacements in the x, y, and z directions are mapped to the three channels of the RGB image respectively; (c1) represents the difference map between the moving image label and the fixed image label; (c2) to (c9) represent the difference maps between the labels after the deformation fields predicted by each method are distorted and the labels in the moving image label and the fixed image label. The cleaner the difference map, the more accurate the cardiac motion estimation.
[0073] Table 1. Results of each method on the ACDC dataset.
[0074]
[0075] Table 2 Results of each method on the CMRxMotion dataset
[0076]
[0077] Example 3
[0078] This invention provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method in Embodiment 1 or Embodiment 2 described above.
[0079] The relevant technical solutions are the same as above, and will not be repeated here.
[0080] Example 4
[0081] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method in Embodiment 1 or Embodiment 2 described above.
[0082] Specifically, the memory may include high-speed random access memory, as well as non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital (SD) cards, flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.
[0083] The relevant technical solutions are the same as above, and will not be repeated here.
[0084] Example 5
[0085] This application provides a computer program product, including a computer program that, when run on a computer, causes the computer to perform the steps of the method in Embodiment 1 or Embodiment 2 described above.
[0086] The relevant technical solutions are the same as above, and will not be repeated here.
[0087] It should be noted that the specific embodiments described in this invention are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains can make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. A cardiac motion estimation method based on motion decomposition, characterized in that, Includes the following steps: Step S1: Acquire paired moving images and fixed image As samples, training and test sets are constructed based on the samples; Step S2: Construct a cardiac motion estimation network, which includes a dual-branch feature encoder network and a feature fusion decoder network. In a dual-branch feature encoder network, the fixed image is processed asynchronously. and moving images Processing is performed to obtain a fixed image. The corresponding first-level fixed image coding features of the pyramid structure ~Level 5 Fixed Image Coding Features and obtaining moving images Corresponding Level 1 moving image coding features ~Level 5 moving image coding features ; In the feature-mixing decoder network, the first-level fixed image encoding features are... ~Level 5 Fixed Image Coding Features and Level 1 moving image coding features ~Level 5 moving image coding features The deformation field is fused and output. ; Step S3: Construct the total loss function; Step S4: Train the cardiac motion estimation network based on the training set; Step S5: Input the pair of moving and stationary images to be predicted into the trained cardiac motion estimation network to obtain the estimated deformation field.
2. The cardiac motion estimation method based on motion decomposition according to claim 1, characterized in that, The dual-branch feature encoder network includes a 3×3 convolutional module, a first single-step encoding module, a second single-step encoding module, a downsampling module, and a third single-step encoding module; fixed image and moving images Asynchronous input features are fed into a 3×3 convolutional module. The output features of the 3×3 convolutional module are fed into the first single-step encoding module and the downsampling module, respectively. The output features of the first single-step encoding module are fed into the second single-step encoding module, and the output features of the downsampling module are fed into the third single-step encoding module. Fixed image When input into the dual-branch feature encoder network: the output features of the 3×3 convolution module, the output features of the downsampling module, the output features of the first single-step encoding module, the output features of the third single-step encoding module, and the output features of the second single-step encoding module are respectively used as the first-level fixed image encoding features. Level 2 fixed image coding features Level 3 fixed image coding features Level 4 fixed image coding features and Level 5 fixed image coding features ; Moving images When input into the dual-branch feature encoder network: the output features of the 3×3 convolutional module, the output features of the downsampling module, the output features of the first single-step encoding module, the output features of the third single-step encoding module, and the output features of the second single-step encoding module are sequentially used as the first-level motion image encoding features. Level 2 moving image coding features Level 3 moving image coding features Level 4 moving image coding features and Level 5 moving image coding features .
3. The cardiac motion estimation method based on motion decomposition according to claim 2, characterized in that, The first, second, and third single-step encoding modules each include a block convolution module, an instance regularization layer, a convolutional layer, an instance regularization layer, and multiple Swin transformer modules. The image features input to the first, second, and third single-step encoding modules are first processed by the block convolution module to obtain block codes with reduced image size. The block codes with reduced image size are then processed sequentially by the instance regularization layer, the convolutional layer, and the instance regularization layer. After being flattened, they are processed sequentially by multiple Swin transformer modules to obtain the encoded features as output features.
4. The cardiac motion estimation method based on motion decomposition according to claim 3, characterized in that, The stride of the block convolution modules in the first, second, and third single-step encoding modules is denoted as... The kernel size of the block convolution module Filling of block convolution modules , To round down, the step size of the downsampling module is set to... .
5. The cardiac motion estimation method based on motion decomposition according to claim 1, characterized in that, The feature mixing decoder network includes a first-level feature mixing transformer module, a second-level feature mixing transformer module, a third-level feature mixing transformer module, a fourth-level feature mixing transformer module, and a fifth-level feature mixing transformer module; In the feature fusion decoder network: Level 5 Fixed Image Coding Features and Level 5 moving image coding features Hybrid features are obtained after channel splicing. Mixed features After channel dimensionality reduction using the corresponding 3×3 convolutional module, the hybrid features are obtained. Mixed features Level 5 fixed image coding features and Level 5 moving image coding features The input is fed into the 5th-level feature mixing transformer module, which outputs the intermediate deformation field. and mixed features , definition For intermediate deformation fields Upsampling is performed to obtain the upsampled deformation field. Using upsampled deformation field For the first Hierarchical moving image coding features Features are obtained by twisting ,feature Mixed features and the Hierarchical fixed image coding features After channel concatenation, the features are obtained by channel dimensionality reduction using a corresponding 3×3 convolutional module. ,feature , No. Hierarchical fixed image coding features and characteristics Enter the number Hierarchical Feature Hybrid Transformer Module, No. The hierarchical feature hybrid transformer module predicts intermediate temporary deformation fields. Intermediate temporary deformation field and upsampling deformation field By merging, an intermediate deformation field can be obtained. , No. The hierarchical feature blending transformer module updates the output blended features. , The deformation field is the final output of the feature mixing decoder network. .
6. The cardiac motion estimation method based on motion decomposition according to claim 5, characterized in that, In the fifth-level feature hybrid transformer module: the fifth level has fixed image coding features. and Level 5 moving image coding features The query matrix is obtained by expanding each matrix separately and then performing linear projection on each matrix separately. Bond matrix Mixed features After being processed by a parameter-regularized convolutional head and unfolded into a value matrix, Query matrix Key matrix and value matrix constitute matrix, The matrix is input to each windowed multi-head cross-attention mechanism unit for computation, resulting in the sub-deformation field of the deformation component corresponding to each windowed multi-head cross-attention mechanism unit. , The sequence number of the windowed multi-head cross-attention mechanism unit, sub-deformation field The input sub-component weighted summation module obtains weights corresponding one-to-one with the deformation components through adaptive learning using a corresponding multilayer perceptron. Sub-deformation field and weight The intermediate temporary deformation field is calculated by weighted summation. Mixed features The image is obtained by upsampling using the image patch expansion module. ; In the 4th to 1st level feature hybrid transformer module: Hierarchical fixed image coding features and characteristics The query matrix is obtained by expanding each matrix separately and then performing linear projection on each matrix separately. Bond matrix Mixed features After being processed by a parameter-regularized convolutional head and unfolded into a value matrix, Query matrix Key matrix and value matrix constitute matrix, The matrix is input to each windowed multi-head cross-attention mechanism unit for computation, resulting in the sub-deformation field of the deformation component corresponding to each windowed multi-head cross-attention mechanism unit. , The sequence number of the windowed multi-head cross-attention mechanism unit, sub-deformation field The input sub-component weighted summation module obtains weights corresponding one-to-one with the deformation components through adaptive learning using a corresponding multilayer perceptron. Sub-deformation field and weight The intermediate temporary deformation field is calculated by weighted summation. Mixed features The image is obtained by upsampling using the image patch expansion module. .
7. The cardiac motion estimation method based on motion decomposition according to claim 5, characterized in that, The construction of the total loss function includes the following steps: Using deformation field For moving images Distorted image vs. fixed image Calculate similarity loss based on deformation field The smoothness loss is calculated using the L2 norm of the gradient, and the weighted sum of the similarity loss and the smoothness loss is used as the total loss function.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the cardiac motion estimation method based on motion decomposition as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the cardiac motion estimation method based on motion decomposition as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the cardiac motion estimation method based on motion decomposition as described in any one of claims 1 to 7.