Space-time medical image registration method and system
By using a hybrid recursive network architecture and wavelet transform loss function, the accuracy and stability problems of traditional medical image registration methods under extreme deformation or complex temporal changes are solved, and high-precision spatiotemporal medical image registration is achieved.
Patent Information
- Application Number
- CN202510853115.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-31
AI Technical Summary
Existing medical image registration methods suffer from insufficient accuracy and stability when dealing with large deformations or complex temporal changes. In particular, they cannot effectively capture the trend of change in the time dimension in lung registration. Furthermore, traditional methods have a limited receptive field, resulting in localized uneven deformation fields that affect the accuracy and stability of registration.
A hybrid recursive network architecture is adopted, combining ConvLSTM and SwinLSTM. Multi-scale spatiotemporal attention aggregation is used to assist skip connections, and wavelet transform loss function is introduced to explicitly model global spatial features and optimize high-frequency information alignment, thereby improving the accuracy and robustness of image registration.
It significantly improves the network's ability to capture local spatiotemporal details and global spatiotemporal dependencies in temporal medical image registration, and enhances registration accuracy and stability, especially maintaining high-precision alignment when dealing with large-scale or complex deformations.
Smart Images

Figure CN120876557A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image registration technology, specifically relating to a spatiotemporal medical image registration method and system. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Deformable medical image registration is a key technology in medical image processing, used to establish spatial correspondences between images to achieve precise alignment of anatomical structures. It plays a crucial role in medical applications such as disease diagnosis, radiotherapy planning, and surgical navigation.
[0004] In practical applications, due to the complexity of organ deformation, tumor growth, and pathological changes during treatment, the registration of time-series medical images has become an extremely challenging task.
[0005] Traditional medical image registration relies on a combination of transformation models, loss functions, and optimization algorithms to iteratively update the deformation field of image pairs. However, this method is computationally expensive and slow in practice. With the development of deep learning technology, learning-based registration methods are considered an effective alternative to traditional methods in terms of accuracy and computational cost.
[0006] Many researchers have provided valuable insights into avoiding reliance on real deformation fields, overcoming large-scale deformations and nonlinear deformations caused by complex movements of soft tissue organs during physiological activities such as respiration, and solving long-distance spatial relationship problems in image registration. While these deep learning-based methods have demonstrated superior performance in many medical image registration tasks, they still have limitations when facing large deformations or complex temporal changes, especially in applications requiring accurate temporal change modeling. For example, in lung registration, traditional pairwise registration methods typically focus on registering two images with large deformations, considering only the relative deformation information between specific frames while ignoring the continuity between phases, resulting in insufficient modeling of the temporal continuity of the deformation field. This can lead to the model failing to effectively capture the temporal trend of lung deformation during dynamic respiration, resulting in locally inconsistent or uneven deformation fields, thus affecting the accuracy and stability of registration.
[0007] On the other hand, existing technologies still rely on local convolution operations to address the challenges of temporal relationships in medical image registration. The size of the convolution kernel limits the receptive field, which limits these methods in capturing global spatial features in images. This local operation may lead to a failure to fully perceive long-range dependencies between anatomical structures when dealing with large-scale or complex deformations, resulting in local unevenness or even irrationality in the registration deformation field, thus affecting the stability and accuracy of registration.
[0008] In addition, while some existing methods have successfully enhanced image alignment using wavelet transform, a dedicated wavelet transform loss function has not yet been introduced during training to further optimize the alignment of high-frequency information. Furthermore, traditional spatial domain-based similarity measures and mutual information (MI) are not sensitive enough to the gradient of high-frequency information, leading to errors in the registration of fine structures. Summary of the Invention
[0009] To address the aforementioned problems, this invention proposes a spatiotemporal medical image registration method and system. This invention provides a hybrid recursive network architecture that can explicitly model global spatial features, thereby overcoming the limitations of traditional convolutional operations in terms of receptive field. It also utilizes multi-scale spatiotemporal attention aggregation to assist skip connections, thus better integrating spatiotemporal features at different scales. Furthermore, this invention designs a wavelet transform-based regularized loss function to further improve the registration results by aligning high-frequency information in the frequency domain.
[0010] According to some embodiments, the present invention adopts the following technical solution: A spatiotemporal medical image registration method includes the following steps: The image sequence to be registered is obtained, and a predetermined time phase is used as the moving image, while the remaining time phases are used as fixed images. The trained hybrid recursive spatiotemporal registration network is used to process the image sequence to obtain the deformation field based on the moving image, and then it is registered to other images. The hybrid recursive spatiotemporal registration network includes an encoder, a bottleneck layer, and a decoder. In the encoding stage of the encoder, a ConvLSTM module is introduced to model the temporal dependence in the image sequence, enabling the network to extract local dynamic features frame by frame and form a stable temporal representation on a spatial scale. The bottleneck layer includes two consecutive core spatiotemporal modeling units. The structure of the core spatiotemporal modeling units incorporates long short-term memory and self-attention mechanisms to model long-term dependencies and global spatiotemporal relationships in image sequences. The decoder includes a multi-level decoding module. A multi-scale spatiotemporal attention fusion module is introduced into the feature fusion path of the hybrid recursive spatiotemporal registration network. Each level of the decoding module fuses with the output features of the corresponding scale of the multi-scale spatiotemporal attention fusion module of the encoder through a skip connection to compensate for the spatial detail loss of high-level semantic features. The multi-scale spatiotemporal attention fusion module includes a spatial aggregation path and a channel enhancement path, which are used to model the spatial semantic relationships between local regions and the response importance between channels, respectively, and achieve cross-scale semantic integration through dual-branch feature enhancement.
[0011] As an alternative implementation, the encoder employs multi-layer 3D convolution and spatial downsampling operations, where each layer halves the spatial size and doubles the number of channels. The input image is first processed through a 3D convolutional layer for feature mapping. Then, the feature maps are fed into the convolutional long short-term memory module and spatial downsampling in chronological order to achieve local modeling of temporal information. In addition to the feature maps, information from the previous time step image is passed to the network in the form of unit states and hidden states. At least two levels of the encoding path introduce ConvLSTM modules, and the output tensors of each ConvLSTM module have different sizes.
[0012] As an alternative implementation, the core spatiotemporal modeling unit includes a Patch Embedding layer, a SwinLSTM unit, and an attention block. The Patch Embedding layer divides the input into multiple small blocks, and each small block is transformed into a low-dimensional space suitable for processing by the SwinLSTM unit through a linear mapping. The SwinLSTM unit combines the temporal modeling capability of LSTM with the self-attention mechanism of Swin Transformer. At each time step, the input features and the hidden state of the previous time step work together to update the current hidden state and the unit state through the SwinLSTM unit. The attention block employs a multi-head window self-attention mechanism and a moving window self-attention mechanism to model global spatial dependencies in 3D images. It captures long-range spatial relationships by calculating the similarity between all locations in the image through self-attention.
[0013] As an alternative implementation, the spatial aggregation path is used to capture the spatial context relationship between different time frames, including three convolutional kernels of different scales to extract local and global spatial information, and to form spatial enhancement features through multi-scale fusion. A spatial attention weight map is constructed using pooling operations and a sigmoid activation function. The spatial attention weight map weights the original input features element-wise, thereby highlighting regional features in the spatiotemporal dimension.
[0014] As an alternative implementation, the channel enhancement path uses global average pooling to obtain global channel statistics for each time frame, and forms a channel attention weight map through two channel-wise convolutional layers and ReLU activation. The channel attention weight map is then applied to the spatially enhanced feature map.
[0015] As an alternative implementation, the weighted outputs of the spatial aggregation path and the channel enhancement path are fused by element-wise multiplication and addition to obtain the feature-enhanced fused output.
[0016] As an alternative implementation, during the training process of the hybrid recursive spatiotemporal registration network, the loss function includes an image similarity loss function, a deformation field smoothness loss, a wavelet transform loss function, and a Jacobian matrix determinant regularization term, which jointly optimize the network parameters.
[0017] As an alternative implementation, the wavelet transform loss function is based on the multi-scale decomposition characteristics of discrete wavelet transform. The input image is decomposed into low-frequency and high-frequency sub-bands, where the low-frequency sub-band is used to characterize the global anatomical structure and the high-frequency sub-band is used to capture multi-directional edge and texture details. The wavelet transform loss function is applied only to the last scale of the hybrid recursive spatiotemporal registration network. The goal of the wavelet transform loss function is to minimize the frequency domain difference between the registered image and the fixed image.
[0018] A spatiotemporal medical image registration system, comprising: The image acquisition module is configured to acquire the image sequence to be registered, using a predetermined time phase as the moving image and the remaining time phases as the fixed image; The registration module is configured to process the image sequence using a trained hybrid recursive spatiotemporal registration network to obtain a deformation field based on the moving image, and then register it to other images. The hybrid recursive spatiotemporal registration network includes an encoder, a bottleneck layer, and a decoder. In the encoding stage of the encoder, a ConvLSTM module is introduced to model the temporal dependence in the image sequence, enabling the network to extract local dynamic features frame by frame and form a stable temporal representation on a spatial scale. The bottleneck layer includes two consecutive core spatiotemporal modeling units. The structure of the core spatiotemporal modeling units incorporates long short-term memory and self-attention mechanisms to model long-term dependencies and global spatiotemporal relationships in image sequences. The decoder includes a multi-level decoding module. A multi-scale spatiotemporal attention fusion module is introduced into the feature fusion path of the hybrid recursive spatiotemporal registration network. Each level of the decoding module fuses with the output features of the corresponding scale of the multi-scale spatiotemporal attention fusion module of the encoder through a skip connection to compensate for the spatial detail loss of high-level semantic features. The multi-scale spatiotemporal attention fusion module includes a spatial aggregation path and a channel enhancement path, which are used to model the spatial semantic relationships between local regions and the response importance between channels, respectively, and achieve cross-scale semantic integration through dual-branch feature enhancement.
[0019] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps in the method described above.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention deeply integrates ConvLSTM and SwinLSTM, significantly improving the network's ability to simultaneously capture local spatiotemporal details and global spatiotemporal dependencies in temporal medical image registration, and solving the problem of local non-smoothness of deformation field caused by local operations in traditional ConvLSTM.
[0021] This invention introduces a multi-scale spatiotemporal attention aggregation module, which enhances the integration capability of spatiotemporal features at different scales, thereby further improving the accuracy and robustness of registration.
[0022] This invention introduces wavelet transform loss into the loss function. By evaluating image differences in the frequency domain, wavelet transform loss can more effectively align local details, edges, and high-frequency information in the image. This enables the proposed method to maintain high registration accuracy when processing time-series images with large deformations, especially in the alignment at the detail level. It achieves constraints on the high-frequency components of the image and provides a new optimization approach for high-precision alignment of time-series image registration.
[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0024] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0025] Figure 1 This is a flowchart of a registration method according to one embodiment; Figure 2 This is a schematic diagram of a hybrid recursive spatiotemporal network registration network framework according to one embodiment; Figure 3 This is a schematic diagram of a multi-scale spatiotemporal attention fusion module according to one embodiment; Figure 4 This is a schematic diagram of wavelet transform loss in one embodiment; Figure 5 This is a set of marker points for Case 8 in the DIR-Lab dataset before and after registration, as shown in one embodiment. (a) is before registration and (b) is after registration. All axes are in millimeters. Figure 6 This is the registration result of Case 2 in the Creatis dataset of one embodiment. The first column is the fixed image, the second column is the moving image, the third column is the composite image of the fixed image and the moving image before registration, the fourth column is the composite image of the fixed image and the moving image after registration, and the sixth column is the registered image. Figure 7 This is a scatter plot of each case in the Creatis dataset of one embodiment. The horizontal axis represents the position of the marker along the z-axis (mm), and the vertical axis represents the TRE size at the corresponding position (mm). In the figure, blue dots represent TRE before registration, and orange dots represent TRE after registration. The histogram on the horizontal axis shows the distribution of TRE before registration along the z-axis, and the histogram on the vertical axis shows the distribution of TRE before registration (blue) and after registration (orange). Figure 8 This is an example of intensity difference maps for Case 8 in the Dir-Lab dataset. The first row shows the intensity difference between the fixed and moving images before registration; the second row shows the intensity difference between the fixed and distorted images after W-STFNet registration; the third row shows the intensity difference between the fixed and distorted images after STFNet registration; the fourth row shows the intensity difference between the fixed and distorted images after W-TFNet registration; and the fifth row shows the intensity difference between the fixed and distorted images after W-STNet registration. Figure 9 This is a bar chart showing the ablation experiment results of W-STFNet on the Dir-Lab dataset, based on one embodiment. Detailed Implementation
[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0027] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0028] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0029] Where there is no conflict, the embodiments and features described in this application may be combined with each other.
[0030] Example 1 A spatiotemporal medical image registration method includes the following steps: The image sequence to be registered is obtained, and a predetermined time phase is used as the moving image, while the remaining time phases are used as fixed images. The trained hybrid recursive spatiotemporal registration network is used to process the image sequence to obtain the deformation field based on the moving image, and then it is registered to other images. The hybrid recursive spatiotemporal registration network includes an encoder, a bottleneck layer, and a decoder. In the encoding stage of the encoder, a ConvLSTM module is introduced to model the temporal dependence in the image sequence, enabling the network to extract local dynamic features frame by frame and form a stable temporal representation on a spatial scale. The bottleneck layer includes two consecutive core spatiotemporal modeling units. The structure of the core spatiotemporal modeling units incorporates long short-term memory and self-attention mechanisms to model long-term dependencies and global spatiotemporal relationships in image sequences. The decoder includes a multi-level decoding module, and a multi-scale spatiotemporal attention fusion module is introduced into the feature fusion path of the hybrid recursive spatiotemporal registration network. Each level of the decoding module fuses with the output features of the corresponding scale of the multi-scale spatiotemporal attention fusion module of the encoder through a skip connection to compensate for the spatial detail loss of high-level semantic features. The multi-scale spatiotemporal attention fusion module includes a spatial aggregation path and a channel enhancement path, which are used to model the spatial semantic relationships between local regions and the response importance between channels, respectively, and achieve cross-scale semantic integration through dual-branch feature enhancement.
[0031] The following is a detailed introduction.
[0032] like Figure 1 As shown, this embodiment proposes a registration method, the main process of which is as follows: Let... = This represents a sequence of images to be registered.
[0033] This embodiment uses lung image registration as an example for illustration.
[0034] A complete 4D-CT sequence of the lungs includes two phases: the expiratory phase (…). - ) and inhalation phase ( - In this embodiment, to simplify the model training and inference process, a single motion state sequence is used for registration, selecting the maximum inhalation phase. The input image sequence is treated as a moving image, while the remaining time phases are treated as stationary images. After processing by a sequence registration network, the output is... The deformation field is used as a reference, and then... Registration is performed on other images. The DVF (Deformation Vector Field) output by this registration network is... .
[0035] As shown in formula (1), STN (Spatial Transformer Network) operates frame by frame. This achieves stepwise registration of images within a sequence. Ultimately, the resulting distorted image sequence is: = To improve registration accuracy, this embodiment introduces an image similarity loss function, a deformation field smoothness loss, a wavelet transform loss function, and a Jacobian matrix determinant regularization term to jointly optimize network parameters.
[0036] (1) like Figure 2 As shown in (a), the hybrid recursive spatiotemporal registration network architecture proposed in this embodiment consists of three parts: an encoder, a bottleneck layer, and a decoder. It aims to integrate local dynamic perception and global spatiotemporal modeling capabilities to improve the accuracy and consistency of 4D-CT medical image registration. This framework combines convolutional long short-term memory (LSTM) modules and Swin LSTM modules to enhance the ability to capture local dynamic changes and global context. The network input is a sequence of images with a time dimension, denoted as: (2) Where T is the number of time frames (sequence length). These represent the image's height, width, and number of slices (i.e., spatial dimensions), respectively. The initial number of input channels is set to 1 in this embodiment. Different colored blocks in the tensor structure represent different time channels, and the width represents the number of spatial channels.
[0037] The encoding path employs multi-layer 3D convolution (Conv3D) and spatial downsampling operations, with each layer compressing the spatial size by half and doubling the number of channels. To enhance the modeling capability in the temporal dimension, ConvLSTM modules are introduced in multiple encoding stages to model the temporal dependencies in the image sequence, enabling the network to extract local dynamic features frame by frame and form a stable temporal representation at the spatial scale. This design not only enhances the encoder's ability to perceive subtle deformations between adjacent temporal phases but also compensates for the temporal representation deficiencies of ordinary convolution, forming a stable and temporally consistent local feature representation. The input image is first processed by a convolution kernel with a kernel size of... The convolutional layer performs feature mapping, and the output size is A tensor is used, where C is set to 16. The feature maps are then fed into a convolutional long short-term memory module and spatial downsampling in chronological order to achieve local modeling of temporal information. In addition to the feature maps, information from the previous time-step image is passed to the network in the form of unit states and hidden states (when processing...). When mapping the feature maps, the initial cell state and hidden state are set to 0. This achieves full utilization of historical information and is consistent with the task of generating deformation fields for sequential image registration. The first-stage ConvLSTM output size is... The tensor, the output size of the second-stage ConvLSTM is The tensor.
[0038] The bottleneck layer, located between the encoder and decoder, is designed as two consecutive core spatiotemporal modeling units. Its structure integrates Long Short-Term Memory (LSTM) and a 3D SwinTransformer self-attention mechanism to model long-term dependencies and global spatiotemporal relationships in image sequences. The output of the bottleneck layer is the deepest feature map, with a first-level output size of [size missing]. The tensor, the second-stage output size is In this embodiment, C is set to 16 for the tensor. This bottleneck structure expands the network's effective receptive field and improves its ability to model large-scale deformations, global consistency, and temporal structures.
[0039] The decoding path gradually restores spatial resolution through multiple upsampling layers and two consecutive convolutional layers. Furthermore, to achieve information interaction and fusion between features at different time steps and different resolutions, this embodiment introduces a multi-scale spatiotemporal attention fusion module (MSTAF) in the feature fusion path of the backbone network. Spatiotemporal features at different resolutions are upsampled or downsampled to the target resolution, concatenated, and used as input to the MSTAF. Each decoding module fuses with the output features of the corresponding scale of the encoder's MSTAF through skip connections to compensate for the loss of spatial details in high-level semantic features.
[0040] The decoding module output is gradually restored to the original resolution. The tensor. Finally, through a convolution kernel of size... The convolutional layers generate output sequence deformation fields (DVFs) to drive image registration.
[0041] This network design introduces a local dynamic modeling mechanism (ConvLSTM) at the encoding end and introduces a window attention combined with a recursive 3D SwinLSTM structure at the bottleneck layer. For the first time, it achieves multi-level spatiotemporal consistency modeling from local to global in the registration task, which significantly improves the accuracy, stability and semantic understanding of 4D-CT registration.
[0042] SwinLSTM Block exist Figure 2 In the network architecture shown, the SwinLSTM Block serves as the core module, combining the advantages of SwinTransformer and LSTM to model global spatiotemporal dependencies in medical image sequences. To adapt to 4D-CT medical image registration, this embodiment extends the traditional two-dimensional SwinLSTM module to a three-dimensional SwinLSTM and applies it to medical image sequence registration for the first time. The detailed structure of the SwinLSTM module is as follows... Figure 2 As shown in (b), the module consists of three main parts: Patch Embedding, SwinLSTM Cell, and Swin Transformer Block (STB).
[0043] First, the Patch Embedding layer divides the input into multiple small patches, each with a dimension of [missing information]. (p=4), these small blocks are transformed into a low-dimensional space suitable for the Swing Transformer to process through a linear mapping.
[0044] like Figure 2 As shown in (c), the 3D SwinLSTM Cell combines the temporal modeling capabilities of LSTM with the self-attention mechanism of SwinTransformer. At each time step, the input features and the hidden state from the previous time step work together to update the current hidden state through the SwinLSTM Cell. and unit state Its key equation is shown in equation (3): (3) Where STB represents the Swing Transformer block, and LP represents the linear projection. It is in gated state. and These are the hidden state and the unit state from the previous moment, respectively.
[0045] like Figure 2 As shown in (d), STB employs a multi-head self-attention (W-MSA) and a moving window self-attention (SW-MSA) mechanism to model global spatial dependencies in 3D images. STB captures long-range spatial relationships by calculating the similarity between all locations in the image through self-attention. The key equation of STB is shown in equation (4).
[0046] (4) In medical image sequence registration tasks, accurately modeling spatiotemporal dependencies is crucial. By introducing a 3D SwinLSTM module, the network can effectively capture global spatiotemporal information when processing image sequences, which is particularly important for registering sequences with complex deformations, thereby improving the accuracy and stability of registration.
[0047] The following section details the Multi-scale Spatio-Temporal Attention Fusion (MSTAF) module.
[0048] like Figure 3 As shown, this module consists of a spatial aggregation path and a channel enhancement path, which respectively model the spatial semantic relationships between local regions and the response importance between channels. Through dual-branch feature enhancement, it achieves fine semantic integration across scales.
[0049] Spatial aggregation paths aim to capture the spatial contextual relationships between different time frames. First, the input feature map... , After three convolutional kernels of different scales ( , Local and global spatial information is extracted and spatially enhanced features are formed through multi-scale fusion. Subsequently, a spatial attention weight map is constructed using pooling operations and a sigmoid activation function. ,in This weighted graph applies element-wise weights to the original input features, thereby highlighting regional features in the spatiotemporal dimensions. The channel enhancement path employs global average pooling to obtain global channel statistics for each time frame. Then, it uses two channel-wise... Convolutional layers and ReLU activation form a channel attention weight map This improves the discriminative channel response. The attention map is applied to the spatially enhanced feature map to further perform weighted adjustments on the channel dimensions.
[0050] Finally, the weighted outputs of the spatial aggregation path and the channel enhancement path are fused through element-wise multiplication and addition to obtain the feature-enhanced fused output. Ultimately, this achieves refined integration and enhancement of features.
[0051] In the medical image sequence registration method of this embodiment, we designed and used a variety of loss functions and regularization terms to ensure that the network can achieve efficient and stable registration results. As shown in formula (5), the design of the loss function considers multiple aspects such as image registration accuracy, deformation field smoothness and physical constraints, specifically including image similarity loss, deformation field smoothness loss, Jacobian loss and wavelet transform loss.
[0052] (5) in , , It is a hyperparameter used to balance the weights of various loss terms in the overall optimization process.
[0053] (1) Similarity loss To measure the similarity between the registered image and the fixed image, this embodiment uses normalized cross-correlation (NCC) as the similarity loss function. Specifically, Let T represent the local intensity of image T at position P, as shown in equation (6): (6) in, This indicates that the size around voxel p is... The image is looped on the cube; in the experimental setup of this embodiment, q=5. Specifically, the image is fixed. With the registered image Similarity loss function between Defined as shown in equation (7): (7) in, Represents the image domain. Within the range [-1, 1], The smaller the value, the higher the similarity of the images.
[0054] (2) Smoothing regularization loss To ensure the smoothness of the deformation field during registration, we introduce a gradient smoothing loss. The specific calculation is as follows: (8) in It is the spatial gradient operator.
[0055] (3) Jacobi loss In the medical image registration process, to ensure the physical consistency of anatomical structures and prevent unreasonable folding deformations from occurring in the deformation field, this embodiment introduces Jacobian loss. Its calculation formula is shown below: (9) Where M represents The total number of elements in the array. Represents the deformation field The determinant of the Jacobian matrix at position p.
[0056] Specifically, when the determinant of the Jacobian matrix is negative (i.e. This indicates the presence of folding deformation. By penalizing negative Jacobian determinants, the physical consistency of tissues and organs is ensured, avoiding folding deformation that does not conform to the deformation laws of tissues and organs. The definition of the Jacobian matrix is as follows: (10) in, , , These represent the components of the deformation field in the x, y, and z directions, respectively.
[0057] (4) Wavelet transform loss Wavelet transform can differentiate and optimize the gradient features of high-frequency subbands through frequency domain decomposition, thereby improving edge alignment accuracy. Inspired by this, this embodiment proposes a wavelet transform loss function and introduces it into the medical image registration task for the first time. By constraining the deformation sensitivity of high-frequency subbands in the frequency domain, the registration results are further optimized.
[0058] The wavelet transform loss function designed in this embodiment Based on the multi-scale decomposition properties of discrete wavelet transform. For example... Figure 4 As shown, the input image is decomposed into low-frequency (LLL) and high-frequency subbands (LLH, LHL, etc.), where the low-frequency subband is used to characterize the global anatomical structure, and the high-frequency subband is used to capture multi-directional edge and texture details. Furthermore, this embodiment applies this only to the last scale of the network. This allows the model to focus on fine-tuning high-frequency details after coarse-grained registration, avoiding local optima caused by premature convergence of shallow features. Specifically, The registration result is optimized by calculating the difference between the target image and the registered image in the frequency domain. Equation (11) shows this process, calculating the square of the maximum difference of the two images in each frequency band k in the wavelet domain: (11) Where DWT represents Discrete Wavelet Transform, T n0 This represents the target image in frame N. Indicates passing through a deformation field The transformed distorted image, where K is the set of wavelet transform frequency bands considered.
[0059] like Figure 4 As shown, the design goal of this loss function is to minimize the frequency domain difference between the registered image and the stationary image. By calculating the maximum difference in each high-frequency band (denoted by k), this loss performs fine-tuning of the image at different frequencies, thereby improving registration accuracy. Specifically, this loss function helps the network capture and align subtle structural changes in images, especially in the high-frequency regions of the image, during temporal image registration.
[0060] Compared to traditional spatial domain loss functions, wavelet transform loss, by evaluating image differences in the frequency domain, can more effectively align local details, edges, and high-frequency information in images. This enables the method proposed in this embodiment to maintain high registration accuracy, especially at the level of detail, when processing temporal images with significant deformation.
[0061] The proposed W-STFNet employs a one-time learning strategy, eliminating the need for large amounts of data, and its convergence criterion follows the same settings as GroupRegNet. The algorithm is implemented in the PyTorch framework and optimized using the Adam optimizer with a learning rate of 0.001. The total number of trainable parameters in the network is 4.22 million. Regularization term... , and Based on experience, the values were set to 1×10−3, 1×10−2, and 1×10−2 respectively. All experiments were conducted on NVIDIA A100 GPUs.
[0062] This embodiment evaluates the accuracy of registration by calculating the average Euclidean distance (i.e., target registration error) between the marker points in the fixed image and the corresponding marker points in the distorted image, as shown in Equation (12): (12) in and These represent the corresponding marker points in the moving and stationary images, respectively. ( ) represents the estimated deformation field.
[0063] Furthermore, the present invention also calculates The rationality of the deformation field is assessed by the proportion of voxels. Specifically, the determinant of the Jacobian matrix of the deformation field describes the change in deformation volume it represents. When it has a negative value, it indicates that the tissue or organ has undergone unreasonable folding deformation. To quantitatively evaluate the effectiveness of W-STFNet, this embodiment uses the Dir-Lab dataset from the University of Texas MD Anderson Cancer Center and the Creatis dataset from the Léon Bérard Cancer Center in Lyon, France. The Dir-Lab dataset contains 4D-CT scans of the chest from 10 patients, with each scan set including 10 phases. 3D-CT images of [the data set] are provided. The slice thickness is 2.5 mm, and the in-plane resolution ranges from 0.97 mm to 1.16 mm. This dataset provides the end-inspiratory phase ([the data set is missing here]). ) and end-expiratory phase ( The dataset contains 300 pairs of markers corresponding to the two extreme phases. Furthermore, this dataset includes images from the expiratory phase (…). The Creatis dataset contains 75 sets of markers depicted on six images. These markers were repeatedly identified and aligned by multiple thoracic imaging experts with reference to anatomical features, and this set of markers serves as a reference for evaluating the registration accuracy of lung images in each case. The Creatis dataset contains 4D-CT scans of six patients with malignant lung tumors, with each scan set including 10 phases. The dataset contains 3D-CT images with a slice thickness of 2 mm and an in-plane resolution ranging from 0.78 mm to 1.17 mm. In this dataset, the maximum inspiratory phase for Case 1 is T10, and the maximum expiratory phase is T60; for the remaining cases, the maximum inspiratory and expiratory phases are T00 and T50. Approximately 100 expert markers are provided for each maximum inspiratory and expiratory phase.
[0064] During preprocessing, the original image is cropped to the lung region based on the distribution of reference markers. First, the 5% and 95% quantiles of the values in the image sequence are calculated to limit the image values to this range, so as to avoid artifacts that are too bright or too dark from adversely affecting the value range. Then, the image intensity is normalized to [0,1].
[0065] Evaluation of the algorithm on the Dir-Lab dataset Figure 5 This demonstrates the landmark 300 registration performance of the proposed method W-STFNet on chest images of Case 8 patients in the Dir-Lab dataset. Figure 5 As can be seen in (a), before registration, there is a significant deviation between the moving markers (green) and the fixed markers (red), especially in the periphery of the chest tissue, where the misalignment of the markers is more severe, indicating significant displacement and deformation. After W-STFNet registration, as shown... Figure 5 As shown in (b), the moving markers significantly overlap with the fixed markers, the spatial difference between the two sets of markers is greatly reduced, and the overall color saturation is improved, indicating that the method can effectively align the markers. Experimental results show that the average marker distance between the maximum inhalation (end of inspiration) and maximum exhalation (end of exhalation) phases in the dataset decreased from 8.46 mm to 1.13 mm.
[0066] To quantitatively analyze the performance of the W-STFNet registration network, this embodiment evaluates it on the widely used Dir-Lab dataset and compares it with traditional methods such as PTVreg, supervised learning methods (e.g., Hering et al.), and current mainstream unsupervised registration methods based on deep learning (e.g., Fechter et al.'s One-Shot Learning method, GroupRegNet's group registration method, Lung-CRNet's convolutional recurrent network method, PLOSL method, ORRN's ODE-based method, C-LSTMNet's convolutional long short memory network method, HPRN's hybrid multi-scale method, LungRegNet's dual-network method, and VoxelMorph's general fast framework) (see Table 1). The Dir-Lab dataset contains significant breathing motion, especially in Case 6, Case 7, and Case 8, where the deformation is more severe.
[0067] As shown in Table 1, the initial TRE values of each case were high before registration, indicating significant displacement and deformation. The W-STFNet proposed in this embodiment demonstrates a significant advantage in average target registration error (TRE), achieving an average TRE of 1.13 ± 0.72 mm on landmark 300. Especially in cases with large deformations, such as Case 6 (1.13 ± 0.68 mm), Case 7 (1.09 ± 0.58 mm), and Case 8 (1.23 ± 1.02 mm), W-STFNet exhibits more robust registration performance and higher registration accuracy compared to other methods.
[0068] Specifically, W-STFNet performs comparably to the traditional PTVreg method. Compared to Fechter et al., although both methods rely on single-step learning for registration, the method in this embodiment captures more comprehensive temporal and spatial information through a hybrid spatiotemporal network structure, thus achieving higher registration accuracy. Furthermore, compared to GroupRegNet, W-STFNet explicitly captures spatiotemporal dynamic information through a hybrid spatiotemporal network, enabling more precise handling of complex lung motion and significant anatomical changes. While Lung-CRNet, ORRN, and C-LSTMNet employ recursive network structures, their performance is still inferior to the method in this embodiment when dealing with large deformations. HPRN proposes a multi-scale learning method, but it falls slightly short in preserving fine-scale details. Although VoxelMorph performs well in terms of efficiency, its simple convolutional structure makes it difficult to effectively capture large-scale and complex temporal information. The PLOSL method introduces a loss function for lung tissue volume and vascular constraints, demonstrating outstanding performance in microstructure registration, but it still slightly lags behind the method in this embodiment in terms of TRE.
[0069] In summary, the experimental results of W-STFNet on the Dir-Lab dataset clearly demonstrate its advantages in accuracy and robustness, especially in the registration of complex, highly deformed sequential medical images, which fully verifies the effectiveness of the wavelet transform regularized hybrid recurrent network proposed in this embodiment.
[0070] Table 1 Comparison of TREs (mean) std in mm): Evaluate W-STFNe with other learning-based and traditional DIR methods using landmark 300 on the Dir-Lab dataset.
[0071] To further verify the robustness and effectiveness of the proposed W-STFNet on images of different respiratory phases, we performed a refined evaluation on landmark75 provided by the Dir-Lab dataset for different target phases (T10, T20, T30, T40, T50) and compared it with the traditional methods pTVreg and GroupRegNet (see Table 2).
[0072] Table 2 shows the TRE results for each patient case (Case 1 to Case 10) at different target phases. Compared with traditional methods pTVreg and GroupRegNet, W-STFNet exhibits better registration accuracy at T50, the phase of drastic respiratory motion changes. This indicates that W-STFNet, by combining local and global spatiotemporal information captured by Conv-LSTM and SwinLSTM, more effectively adapts to deformation changes between different respiratory phases. Furthermore, the TREs at T10 and T50 are generally smaller than those at other phases, possibly because T10 has smaller deformation, while T50 is more stable than intermediate phases.
[0073] The TRE results from each case show that W-STFNet has significant advantages in overall accuracy and stability. In particular, in cases such as Case 6, Case 7 and Case 8, which are usually considered difficult to register, W-STFNet still maintained good TRE, demonstrating its superiority in capturing complex motion patterns and maintaining registration accuracy.
[0074] Table 2 Comparison of TREs (mean) std in mm): TREs (mean ± standard deviation, mm) of W-STFNet on different target phase images were compared using landmark 75 of the Dir-Lab dataset. Results for the T50 phase of pTVreg and GroupRegNet are included for reference.
[0075] In summary, the W-STFNet proposed in this embodiment not only demonstrates significant advantages in overall accuracy when processing medical image sequence registration tasks, but also exhibits good robustness and adaptability when dealing with drastic motion and large deformations. This further verifies the effectiveness and advancement of the hybrid recursive spatiotemporal network structure employed by W-STFNet.
[0076] Algorithm evaluation on the Creatis dataset To further verify the applicability and robustness of the proposed W-STFNet in more complex medical image sequence registration scenarios, we conducted quantitative evaluation and comparative analysis on the publicly available Creatis 4D-CT dataset. As shown in Table 3, this experiment compared the W-STFNet method with the widely cited classic method Vandemeulebroucke on the Creatis dataset, as well as several deep learning-based methods, including GDL-FIR, Fechter et al., GroupRegNet, and ORRN.
[0077] Table 3. Comparison of TREs (mean ± standard value, mm) between W-STFNet and other learning-based and traditional methods on the Creatis dataset.
[0078] Experimental results show that W-STFNet achieves the best registration accuracy on the Creatis dataset, with an average TRE of 0.87±0.56 mm, significantly outperforming the traditional method Vandemeulebroucke (1.46±1.65 mm). This result demonstrates that W-STFNet can more accurately model lung tissue deformation in complex respiratory motion scenarios.
[0079] Compared with deep learning-based registration methods such as GDL-FIRE (1.74±1.26 mm), Fechter et al. (1.49±1.59 mm), GroupRegNet (1.03±0.64 mm), and ORRN (1.00±0.87 mm), W-STFNet demonstrates significant advantages in both registration accuracy and stability, fully reflecting its unique ability to capture global and local spatiotemporal relationships. Particularly in Case 2, a patient with intense respiratory motion and significant deformation, the W-STFNet method exhibits better robustness and accuracy, further demonstrating the model advantages brought by the fusion of ConvLSTM and SwinLSTM.
[0080] Figure 6 The registration results for Case 2 in the Creatis dataset are presented. As can be seen from the figures, before registration, the green and purple structures in the composite image are severely misaligned, indicating a large initial positional deviation. After registration, the structural alignment in the image is significantly enhanced, with clearer edges, especially maintaining anatomical consistency in the hilum and vascular regions, further validating the robustness and accuracy of W-STFNet under large deformation conditions.
[0081] To more intuitively demonstrate the registration effect of W-STFNet, Figure 7 The scatter plot shows the joint distribution of each case in the Creatis dataset. As can be seen from the plot, the TRE after registration is significantly lower than before registration, indicating that W-STFNet can effectively reduce the registration error of landmarks across the entire lung space. The marginal distribution plots above and to the right of the scatter plot also visually demonstrate that the distribution of TRE after registration is concentrated in a lower numerical range, further validating the advantages of W-STFNet in terms of accuracy and robustness, and showcasing the model's excellent performance in handling spatially complex and significantly deformed lung motion scenarios.
[0082] ablation experiment To verify the effectiveness of the proposed W-STFNet, we conducted a series of ablation experiments using the publicly available Dir-Lab dataset to analyze the impact of different modules on model performance. As shown in Table 4, we removed different modules during the ablation experiments, including wavelet transform loss, MSTAF module, and SwinLSTM module, and named the models after removal as STFNet, W-STNet, and W-TFNet. The purpose of these ablation experiments was to evaluate the impact of each module on model registration performance.
[0083] Table 4 Ablation experiment results of W-STFNet on the Dir-Lab dataset
[0084] Experimental results show that removing the wavelet transform loss improves the TRE of STFNet, especially in handling complex deformation scenes (Case 6, Case 7, and Case 8), where registration accuracy drops significantly. This result verifies the important role of wavelet transform loss in capturing details. Removing the MSTAF module significantly reduces W-STNet's ability to fuse spatiotemporal information, particularly in cases with complex motion scenes, where registration accuracy drops markedly. Removing the bottleneck layer SwinLSTM module reduces W-TFNet's ability to capture complex spatiotemporal dependencies. In Case 7 and Case 8, the SwinLSTM module plays a particularly prominent role, effectively modeling spatiotemporal features and improving the model's global registration performance. Figure 8 It is evident that W-STFNet significantly outperforms STFNet in aligning details of complex deformations, particularly in fine-grained alignment of key information. This strategy not only improves registration quality but also reduces the risk of overfitting, as the wavelet transform loss is applied only to the most relevant high-frequency components, thereby enhancing the network's alignment capability. The detail intensity difference in W-STFNet is significantly lower than that in W-STNet, demonstrating that the MSTAF module plays a crucial role in the effective fusion of multi-scale spatiotemporal features and detail preservation.
[0085] As shown in Figure 9, W-STFNet achieves the best performance in registration accuracy, especially in complex deformation scenes, where it has the lowest registration error, with an average TRE of 1.13 ± 0.72 mm. Compared to other models that remove certain modules, W-STFNet is superior in fine registration and detail processing capabilities for complex spatiotemporal images. Furthermore, we evaluate the regularity of the deformation field by calculating the percentage of negative values in the Jacobian matrix determinant. As shown in Table IV, W-STFNet reduces the percentage of negative values in the Jacobian matrix determinant by 19% compared to W-TFNet. This indicates that W-STFNet has more stable control over the entire deformation field, avoiding unreasonable deformation folding phenomena, especially in more complex registration tasks. This result verifies the advantages of the proposed hybrid recursive spatiotemporal network in temporal dependency modeling. It effectively combines local detail feature extraction with global contextual information modeling, reducing the local unsmoothness of the deformation field, thereby significantly improving the stability and accuracy of registration.
[0086] In summary, the proposed Hybrid Recurrent Spatiotemporal Network (W-STFNet) based on wavelet transform regularization demonstrates excellent performance in medical image registration tasks, especially in handling complex spatiotemporal dependencies and scenes with large deformations. By combining ConvLSTM and SwinLSTM, W-STFNet can effectively capture local dynamic changes and global spatiotemporal dependencies, overcoming the uneven deformation field caused by local operations in ConvLSTM. In particular, W-STFNet optimizes the registration results in the frequency domain by introducing wavelet transform loss, thereby more sensitively capturing subtle differences in images (especially in high-frequency details), further improving registration accuracy. It performs exceptionally well in spatiotemporal medical image registration of dynamic organs such as the lungs.
[0087] Experimental results show that W-STFNet outperforms current mainstream deep learning methods and is comparable to traditional methods on the Dir-Lab and Creatis public datasets. On the Dir-Lab dataset, W-STFNet demonstrates significant robustness in registering lung images with large deformations, achieving a mean TRE of 1.13 ± 0.72 mm, which is superior to unsupervised registration methods based on deep learning. Especially when handling complex deformation scenarios (such as Case 6, Case 7, and Case 8), W-STFNet not only maintains high accuracy but also ensures the smoothness of the deformation field. Furthermore, W-STFNet performs exceptionally well in registering images at different respiratory phases, particularly at the T50 end-expiratory stage, with registration accuracy exceeding that of the compared methods. This indicates that W-STFNet can effectively adapt to deformations between different respiratory phases, enhancing the model's global adaptability and spatiotemporal information modeling capabilities.
[0088] Ablation experiments further validated the key roles of wavelet transform loss, the MSTAF module, and the SwinLSTM module in improving model performance. Compared to W-STFNet, STFNet showed a significant decrease in registration accuracy in complex deformation scenarios, verifying the important role of the wavelet transform loss function in detail processing. Removing the MSTAF module significantly reduced the registration accuracy of W-STFNet, indicating that the MSTAF module plays a crucial role in the effective fusion of multi-scale spatiotemporal features and detail preservation. Removing the SwinLSTM module significantly weakened W-STFNet's ability to capture global spatiotemporal features. Furthermore, W-STFNet reduced the percentage of negative values in the Jacobian matrix determinant by 19%, further demonstrating the advantages of hybrid recursive spatiotemporal networks in temporal dependency modeling. It can effectively combine local detail feature extraction with global contextual information modeling, reducing the local non-smoothness of the deformation field.
[0089] Example 2 A spatiotemporal medical image registration system, comprising: The image acquisition module is configured to acquire the image sequence to be registered, using a predetermined time phase as the moving image and the remaining time phases as the fixed image; The registration module is configured to process the image sequence using a trained hybrid recursive spatiotemporal registration network to obtain a deformation field based on the moving image, and then register it to other images. The hybrid recursive spatiotemporal registration network includes an encoder, a bottleneck layer, and a decoder. In the encoding stage of the encoder, a ConvLSTM module is introduced to model the temporal dependence in the image sequence, enabling the network to extract local dynamic features frame by frame and form a stable temporal representation on a spatial scale. The bottleneck layer includes two consecutive core spatiotemporal modeling units. The structure of the core spatiotemporal modeling units incorporates long short-term memory and self-attention mechanisms to model long-term dependencies and global spatiotemporal relationships in image sequences. The decoder includes a multi-level decoding module. A multi-scale spatiotemporal attention fusion module is introduced into the feature fusion path of the hybrid recursive spatiotemporal registration network. Each level of the decoding module fuses with the output features of the corresponding scale of the multi-scale spatiotemporal attention fusion module of the encoder through a skip connection to compensate for the spatial detail loss of high-level semantic features. The multi-scale spatiotemporal attention fusion module includes a spatial aggregation path and a channel enhancement path, which are used to model the spatial semantic relationships between local regions and the response importance between channels, respectively, and achieve cross-scale semantic integration through dual-branch feature enhancement.
[0090] Example 3 An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the steps in the method provided in Embodiment 1.
[0091] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of one or more computer-usable storage media (including, but not limited to, disk storage, etc.) containing computer-usable program code. CD - ROM It takes the form of a computer program product implemented on (such as optical memory, etc.).
[0092] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0093] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0094] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0095] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made by those skilled in the art without creative effort within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A spatiotemporal medical image registration method, characterized in that, Includes the following steps: The image sequence to be registered is obtained, and a predetermined time phase is used as the moving image, while the remaining time phases are used as fixed images. The trained hybrid recursive spatiotemporal registration network is used to process the image sequence to obtain the deformation field based on the moving image, and then it is registered to other images. The hybrid recursive spatiotemporal registration network includes an encoder, a bottleneck layer, and a decoder. In the encoding stage of the encoder, a ConvLSTM module is introduced to model the temporal dependence in the image sequence, enabling the network to extract local dynamic features frame by frame and form a stable temporal representation on a spatial scale. The bottleneck layer includes two consecutive core spatiotemporal modeling units. The structure of the core spatiotemporal modeling units incorporates long short-term memory and self-attention mechanisms to model long-term dependencies and global spatiotemporal relationships in image sequences. The decoder includes a multi-level decoding module. A multi-scale spatiotemporal attention fusion module is introduced into the feature fusion path of the hybrid recursive spatiotemporal registration network. Each level of the decoding module fuses with the output features of the corresponding scale of the multi-scale spatiotemporal attention fusion module of the encoder through a skip connection to compensate for the spatial detail loss of high-level semantic features. The multi-scale spatiotemporal attention fusion module includes a spatial aggregation path and a channel enhancement path, which are used to model the spatial semantic relationships between local regions and the response importance between channels, respectively, and achieve cross-scale semantic integration through dual-branch feature enhancement.
2. The spatiotemporal medical image registration method as described in claim 1, characterized in that, The encoder uses multi-layer 3D convolution and spatial downsampling operations, each layer compressing the spatial size by half and doubling the number of channels; The input image is first processed through a 3D convolutional layer for feature mapping. Then, the feature maps are fed into the convolutional long short-term memory module and spatial downsampling in chronological order to achieve local modeling of temporal information. In addition to the feature maps, information from the previous time step image is passed to the network in the form of unit states and hidden states. At least two levels of the encoding path introduce ConvLSTM modules, and the output tensors of each ConvLSTM module have different sizes.
3. The spatiotemporal medical image registration method as described in claim 1, characterized in that, The core spatiotemporal modeling unit includes a Patch Embedding layer, a SwinLSTM unit, and an attention block. The Patch Embedding layer divides the input into multiple small blocks, and each small block is transformed into a low-dimensional space suitable for processing by the SwinLSTM unit through a linear mapping. The SwinLSTM unit combines the temporal modeling capability of LSTM with the self-attention mechanism of Swin Transformer. At each time step, the input features and the hidden state of the previous time step work together to update the current hidden state and the unit state through the SwinLSTM unit. The attention block employs a multi-head window self-attention mechanism and a moving window self-attention mechanism to model global spatial dependencies in 3D images. It captures long-range spatial relationships by calculating the similarity between all locations in the image through self-attention.
4. The spatiotemporal medical image registration method as described in claim 1, characterized in that, The spatial aggregation path is used to capture the spatial context relationship between different time frames. It includes three convolutional kernels of different scales to extract local and global spatial information and form spatial enhancement features through multi-scale fusion. A spatial attention weight map is constructed using pooling operations and the Sigmoid activation function. The spatial attention weight map weights the original input features element by element, thereby highlighting regional features in the spatiotemporal dimension.
5. The spatiotemporal medical image registration method as described in claim 1, characterized in that, The channel enhancement path uses global average pooling to obtain global channel statistics for each time frame, and forms a channel attention weight map through two channel-wise convolutional layers and ReLU activation. The channel attention weight map is then applied to the spatially enhanced feature map.
6. The spatiotemporal medical image registration method as described in claim 1, characterized in that, spatial... The weighted output of the aggregation path and the channel enhancement path is fused through element-wise multiplication and addition. The fused output with enhanced features is obtained.
7. The spatiotemporal medical image registration method as described in claim 1, characterized in that, During the training process of the hybrid recursive spatiotemporal registration network, the loss function includes image similarity loss function, deformation field smoothness loss, wavelet transform loss function, and Jacobian matrix determinant regularization term, which jointly optimize the network parameters.
8. The spatiotemporal medical image registration method as described in claim 7, characterized in that, The wavelet transform loss function is based on the multi-scale decomposition characteristics of discrete wavelet transform. The input image is decomposed into low-frequency and high-frequency sub-bands, where the low-frequency sub-band is used to characterize the global anatomical structure and the high-frequency sub-band is used to capture multi-directional edge and texture details. The wavelet transform loss function is applied only to the last scale of the hybrid recursive spatiotemporal registration network. The goal of the wavelet transform loss function is to minimize the frequency domain difference between the registered image and the fixed image.
9. A spatiotemporal medical image registration system, characterized in that, include: The image acquisition module is configured to acquire the image sequence to be registered, using a predetermined time phase as the moving image and the remaining time phases as the fixed image; The registration module is configured to process the image sequence using a trained hybrid recursive spatiotemporal registration network to obtain a deformation field based on the moving image, and then register it to other images. The hybrid recursive spatiotemporal registration network includes an encoder, a bottleneck layer, and a decoder. In the encoding stage of the encoder, a ConvLSTM module is introduced to model the temporal dependence in the image sequence, enabling the network to extract local dynamic features frame by frame and form a stable temporal representation on a spatial scale. The bottleneck layer includes two consecutive core spatiotemporal modeling units. The structure of the core spatiotemporal modeling units incorporates long short-term memory and self-attention mechanisms to model long-term dependencies and global spatiotemporal relationships in image sequences. The decoder includes a multi-level decoding module. A multi-scale spatiotemporal attention fusion module is introduced into the feature fusion path of the hybrid recursive spatiotemporal registration network. Each level of the decoding module fuses with the output features of the corresponding scale of the multi-scale spatiotemporal attention fusion module of the encoder through a skip connection to compensate for the spatial detail loss of high-level semantic features. The multi-scale spatiotemporal attention fusion module includes a spatial aggregation path and a channel enhancement path, which are used to model the spatial semantic relationships between local regions and the response importance between channels, respectively, and achieve cross-scale semantic integration through dual-branch feature enhancement.
10. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the steps of the method according to any one of claims 1-8.