Heart T1 mapping motion correction method based on Swin Transform
By using an encoder-decoder module based on Swin Transformer and a spatial transformation network, the problem that existing cardiac T1 quantitative registration algorithms cannot effectively capture long-distance dependencies is solved, achieving efficient registration of cardiac T1-weighted images and improving image correction performance.
Patent Information
- Application Number
- CN202410586892.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2025-11-14
AI Technical Summary
Existing convolutional neural network-based cardiac T1 quantitative registration algorithms struggle to effectively capture long-range dependencies in images, resulting in insufficient T1-weighted image registration performance for the heart. In particular, they are unable to achieve effective image registration when faced with significant contrast differences and differences in the position and shape of anatomical structures caused by motion.
A quantitative registration method for cardiac T1 images based on the Swin Transformer is adopted. By constructing a multi-level hierarchical Swin Transformer encoder-decoder module and skip connections, the deformation velocity field is predicted and the scaling square layer integral is used. Combined with the spatial transformation network, image distortion operation is performed to achieve effective registration of cardiac T1 weighted images.
It improves the performance of cardiac T1 quantitative registration, effectively corrects contrast differences caused by acquisition parameters and differences in the position and shape of anatomical structures caused by motion, and improves image registration results.
Smart Images

Figure CN120953326A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and more specifically, to a cardiac T1 mapping motion correction method based on Swin Transformer. Background Technology
[0002] Longitudinal relaxation time (T1) of the myocardium is used to diagnose various myocardial diseases, such as myocardial infarction and diffuse myocardial fibrosis. Clinically, sequences such as MOLLI (Modified Look-Locker inversion recovery) and STONE (slice-interleaved T1) are commonly used to acquire multiple T1-weighted images of the heart at different inversion times within the same slice. These images are then fitted pixel-by-pixel using a three-parameter fitting model to obtain the tissue's T1 value. However, a patient's respiration and cardiac motion will cause differences in heart position and shape between T1 images within the same slice, leading to incorrect T1 value estimation. Therefore, image registration of T1-weighted images of the heart within the same slice is necessary to align their corresponding anatomical structures and correct for motion artifacts. However, the biggest challenge in quantitative T1 registration of the heart is the significant contrast differences and variations between T1-weighted images acquired at different inversion times. This makes it difficult to effectively register these images using general intensity-based unimodal image registration and feature-based multimodal image registration methods.
[0003] In existing technologies, deep learning-based T1 quantitative registration algorithms for the heart rely on convolutional neural networks (CNNs) to extract features. However, CNNs, due to their fixed receptive fields, limit their ability to capture long-range dependencies in images, and their kernel weights cannot adaptively adjust based on the input, thus limiting their ability to analyze and process complex images. To circumvent the inherent limitations of CNNs, many computer vision tasks have introduced Transformer models in recent years. Transformer models can encode long-range dependencies and obtain their effective representations. Their attention mechanism helps to acquire global information and map it to multiple spaces, thereby improving the model's expressive power.
[0004] Currently, deep learning-based medical image registration algorithms exhibit excellent registration performance and efficient processing capabilities, making them highly influential in the field of medical image processing. For example, the VoxelMorph medical image registration framework proposed by Balakrishnan et al. has become an effective tool for medical image analysis and is widely used. To address the problem of quantifying specific significant contrast changes in cardiac T1 imaging, researchers have also proposed several registration network variants based on the VoxelMorph architecture. These registration networks all use convolutional neural networks similar to the UNet architecture to extract image features and possess certain registration performance. However, due to the relatively limited expressive power of convolutional neural networks, the performance of current cardiac T1 quantitative registration algorithms still has room for improvement.
[0005] Analysis reveals that existing deep learning-based quantitative T1 registration algorithms for the heart typically rely on the UNet architecture for feature extraction, followed by deformation field prediction to perform registration. However, UNet-based registration networks usually encode and infer from low-level local features without considering the global correlation between images. The actual receptive field of UNet is much smaller than its theoretical receptive field, thus limiting its ability to model long-distance dependencies and making it difficult to handle complex image analysis. To address this issue, some studies have introduced dilated convolutions to expand the network's receptive field; however, their performance on medical image registration tasks has been limited. In contrast, the Transformer model deploys a self-attention mechanism to determine key content in the input image based on contextual information. Furthermore, its self-attention mechanism possesses a considerably large effective receptive field, enabling it to capture long-distance spatial information. In medical image registration tasks, the Transformer model, with its self-attention mechanism and large receptive field, can quickly focus on the region requiring deformation and more effectively match corresponding anatomical structures that are far apart in moving and reference images. However, effective motion correction methods for cardiac T1 mapping based on the Transformer model are currently lacking. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a cardiac T1 mapping motion correction method based on the Swing Transformer. This method includes the following steps:
[0007] Preprocess the experimental dataset (T1-weighted images of the heart);
[0008] The cardiac T1 image is input into a trained motion correction network to obtain a corrected image;
[0009] The motion correction network has an encoder structure and a decoder structure. The multi-level hierarchical Swing Transformer module contained in the encoder structure has skip connections with the corresponding Swing Transformer module in the decoder structure. The motion correction network is used to predict the deformation velocity field from the input image, and then uses a scaling square layer to integrate the velocity field to obtain the deformation field. Finally, a spatial transformation network is used to perform a distortion operation on the moving image to obtain the deregistration result image.
[0010] Compared with the prior art, the advantages of this invention are that, in order to solve the problem that the motion correction network based on CNN has poor correction performance due to its inability to capture long-distance dependencies, a cardiac T1 quantitative registration method based on Swin Transformer is proposed. By constructing a motion correction network based on Swin Transformer, more effective feature extraction and alignment can be performed on the cardiac region (large misalignment region) in cardiac T1 weighted images with complex structure and contrast changes, thereby improving the cardiac T1 quantitative registration performance.
[0011] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.
[0013] Figure 1 This is a flowchart of a cardiac T1 mapping motion correction method based on a Swing Transformer according to an embodiment of the present invention;
[0014] Figure 2 This is an architecture diagram of a cardiac T1 quantitative motion correction network based on a Swing Transformer according to an embodiment of the present invention.
[0015] Figure 3 This is a schematic diagram of the Swing Transformer module structure according to an embodiment of the present invention;
[0016] Figure 4 This is a schematic diagram of the overall process of a cardiac T1 mapping motion correction method based on a Swing Transformer according to an embodiment of the present invention. Detailed Implementation
[0017] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the invention.
[0018] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.
[0019] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0020] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0021] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0022] This invention proposes a quantitative motion correction network for cardiac T1 based on the Swin Transformer model, or Trans-MOCO. First, the cardiac T1 dataset undergoes preprocessing steps such as cropping and affine registration to initially align the registration pairs. Then, it is input into the Swin Transformer-based motion correction network Trans-MOCO for training. Trans-MOCO utilizes a multi-level (e.g., four-level) hierarchical Swin Transformer encoder-decoder module with skip connections to predict the deformation velocity field from the input image. A scaling-squared layer is then used to integrate the velocity field to obtain the deformation field. Finally, a spatial transformation network is used to perform a warping operation on the moving image to solve for the registered image.
[0023] Specifically, see Figure 1 As shown, the provided cardiac T1 mapping motion correction method based on Swing Transformer includes the following steps:
[0024] Step S110: Construct a dataset containing multiple samples of cardiac T1 quantitative imaging and corresponding myocardial segmentation labels.
[0025] For example, both network training and testing used the T1Dataset210 dataset for quantitative cardiac T1 imaging from the Harvard University database (Harvard Data Space). This dataset was acquired using a 1.5T Philips Achieva system and a 32-channel cardiac coil. The dataset contains cardiac T1-mapped scan data from 210 patients (134 males, aged 57 ± 14 years) with known or suspected cardiovascular disease. Each patient's dataset consists of five short-axial slices. The myocardial and endocardial contours of all images (N = 11550) in the dataset were manually delineated for quantitative results analysis.
[0026] Step S120: Construct a Transformer-based motion correction network.
[0027] In one embodiment, a Transformer-based motion correction network, Trans-MOCO, is constructed. The workflow of this network is as follows: Figure 2 As shown. I m and I f These are the moving image and the fixed image, respectively. The Trans-MOCO backbone is similar to the UNet architecture, consisting of an encoder section comprised of a series of Swing Transformer modules and patch merging layers, a decoder section comprised of a series of Swing Transformer modules and patch expansion layers, and skip connections. The structure of the Swing Transformer module is as follows: Figure 3 As shown, W-MSA refers to the multi-head self-attention module with a regular window, and SW-MSA refers to the multi-head self-attention module with a shifted window configuration. The Swin Transformer module is used to model the correlation between high-dimensional features in key regions, thereby achieving effective matching of long-distance related regions. Compared with existing convolutional neural networks, the motion correction network based on Swin Transformer designed in this invention can more effectively register T1-weighted images of the heart with severe anatomical structural position and shape misalignment. The output of the decoder is the predicted velocity field. Therefore, a scaling squared layer is subsequently added to integrate this velocity field to obtain a differentiable and reversible deformation field, thereby achieving topology-preserving registration.
[0028] In one embodiment, the motion correction network aims to optimize the objective function of equation (1), where, This represents the dense deformation field to be optimized. I m and I f It is a registration pair, consisting of the moving image and the reference image. This indicates the use of a dense deformation field Φ on the moving image I mPerform the corresponding deformation to obtain the registered image (Moved Image): L sim This network is used to calculate the image similarity between the registered image and the reference image. Mutual Information (MI) and Modality Independent Neighbourhood Descriptor (MIND) can be used to compute image similarity in this network. The purpose of applying the regularization constraint to the dense deformation field is to ensure the smoothness of the dense deformation field Φ.
[0029]
[0030] Step S130: Train the motion correction network using the constructed dataset to meet the set loss function criteria.
[0031] This invention treats the registration of cardiac T1-weighted images as a multimodal registration task. Therefore, a comprehensive similarity measure (WSM) for evaluating the similarity of cardiac T1-weighted images is proposed. WSM combines mutual information (MI) and modality-independent neighborhood descriptors (MIND), enabling a comprehensive measurement of the intensity and structure similarity between any two cardiac T1-weighted images. WSM is given by equation (2), where λ1 and λ2 are the weights of MI and MIND, respectively, and I... f Represents a reference image, and This represents the image showing the registration result.
[0032]
[0033] In one embodiment, the loss function for training the motion correction network comprises the following components:
[0034] (1) The image similarity loss function is expressed as:
[0035]
[0036] (2) The deformation field smoothness loss function, see formula (4), is calculated by the square of the L2 norm of the gradient of the deformation field at point p.
[0037]
[0038] Therefore, the overall network loss function for training Trans-MOCO is expressed as:
[0039] L Trans-MOCO =L sim +λ·L smooth (5)
[0040] In actual training, the comprehensive network loss function L is used. Trans-MOCO The optimization objective is to obtain optimized network parameters by optimizing the digestion of the most efficient components.
[0041] Once the motion correction network is trained, it can be used for actual cardiac T1 mapping motion correction. Overall, the roadmap of this invention can be found in [reference needed]. Figure 4 As shown.
[0042] In summary, compared with the prior art, the present invention has the following advantages:
[0043] 1) This invention modifies the basic architecture of the cardiac T1 quantitative motion correction network to a network structure based on the Swing Transformer module, which enables the network to extract image features and long-distance feature dependencies more effectively, thereby improving motion correction performance.
[0044] 2) This invention proposes using a motion correction network based on the Swing Transformer to more effectively model the long-distance dependencies of input image pairs by adding a self-attention mechanism and increasing the receptive field. Experimental results demonstrate that this invention addresses the registration challenges in cardiac T1-weighted images caused by contrast differences due to acquisition parameters and differences in the position and shape of anatomical structures due to motion, thereby improving the registered image quality.
[0045] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0046] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0047] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0048] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.
[0049] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0050] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0051] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0052] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.
[0053] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.
Claims
1. A cardiac T1 mapping motion correction method based on Swing Transformer, comprising the following steps: Obtain T1 images of the heart of the target object; The cardiac T1 image is input into a trained motion correction network to obtain a corrected image; The motion correction network has an encoder structure and a decoder structure. The multi-level hierarchical Swing Transformer module contained in the encoder structure has skip connections with the corresponding Swing Transformer module in the decoder structure. The motion correction network is used to predict the deformation velocity field from the input image, and then uses a scaling square layer to integrate the velocity field to obtain the deformation field. Finally, a spatial transformation network is used to perform a distortion operation on the moving image to obtain the deregistration result image.
2. The method according to claim 1, characterized in that, The loss function for training the motion correction network is set as follows: L Trans-MOCO =L sim +λ·L smooth Among them, L Trans-MOCO L is the loss value of the motion correction network. sim It is the image similarity loss function, L smooth λ is the deformation field smoothness loss function, and λ is the set weight parameter.
3. The method according to claim 2, characterized in that, The image similarity loss function is set as follows: Wherein, WSM represents a similarity measure for assessing the similarity of T1-weighted images of the heart, I f Represents a reference image, I m This represents a moving image, where Φ represents a dense deformation field. This indicates the use of a dense deformation field Φ on the moving image I m Perform the corresponding transformation.
4. The method according to claim 2, characterized in that, The deformation field smoothness loss function is expressed as: in, This represents the gradient of the deformation field at point p.
5. The method according to claim 3, characterized in that, The similarity measure used to assess the similarity of T1-weighted cardiac images is expressed as follows: Where MI represents mutual information, MIND represents modality-independent neighborhood descriptor, and λ1 and λ2 are the weight coefficients of the corresponding terms.
6. The method according to claim 3, characterized in that, The motion correction network is used to optimize the following objective function. in, This represents the dense deformation field to be optimized. This indicates the use of a dense deformation field Φ on the moving image I m Perform the corresponding deformation to obtain the registration result image.
7. The method according to claim 1, characterized in that, The Swin Transformer module includes a multilayer perceptron, a layer normalization layer, a multi-head self-attention module with a fixed window configuration, and a multi-head self-attention module with a shifted window configuration.
8. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
9. A computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.