Multi-scale implicit-convolution image registration method and system with enhanced temporal features

CN122453890BActive Publication Date: 2026-09-04SHANDONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610930610.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-04
Estimated Expiration
2046-06-26

AI Technical Summary

Technical Problem

虽然成对方法在处理两帧图像间的形变时表现尚可,但在处理4D序列时,往往会导致生成的形变场在时间维度上不平滑,且难以准确预测中间时刻的复杂运动轨迹,无法满足高精度4D剂量累积与动态肿瘤追踪的临床需求

Benefits of technology

构建双重时空编码的四维连续建模机制。针对肺部呼吸运动的周期性与非均匀采样特性,在隐式神经表示中引入周期性时间编码与帧间距离编码。该机制通过将一维时间标量映射至高维特征空间,赋予模型对呼吸相位及时间跨度的敏锐感知力,有效弥补了传统离散时间采样在捕捉连续动态变化上的不足。通过建立帧间任意时刻的连续映射关系,显著缓解了稀疏采样导致的时间混叠与轨迹突变问题,实现了高分辨率的时序运动预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122453890B_ABST
    Figure CN122453890B_ABST
Patent Text Reader

Abstract

The application discloses a multi-scale implicit-convolution image registration method and system for enhancing time sequence features, relates to the technical field of image registration, and comprises the following steps: for a 4D CT image sequence, an image registration network is used to calculate the instantaneous deformation field between image pairs at adjacent time points; the cumulative deformation field from the initial time point to the current time point is obtained according to the cascade composition of all the instantaneous deformation fields between the initial time point and the current time point; the multi-scale implicit-convolution registration network for enhancing time sequence features is composed of a multi-scale time fusion network, and the multi-scale time fusion network adopts a hybrid architecture combining explicit and implicit, and deeply integrates time sequence information at each link of feature extraction and deformation prediction. By introducing the time dimension, the space-time correlation in the 4D CT image sequence is fully mined, and high-precision and space-time continuous modeling of the lung respiratory motion is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image registration technology, and in particular to a multi-scale implicit-convolutional image registration method and system for enhancing temporal features. Background Technology

[0002] In radiotherapy planning and efficacy evaluation, four-dimensional computed tomography (4D CT) has become the standard imaging modality for capturing lung respiratory motion. Unlike static three-dimensional computed tomography (3D CT), 4D CT provides a series of continuous images that change over time, fully recording the dynamic trajectory of lung tissue throughout the entire respiratory cycle. This temporal data not only contains spatial anatomical information but also implies the motion patterns in the temporal dimension.

[0003] However, most existing deep learning-based registration methods employ a pairwise registration strategy, which decomposes 4D CT sequences into multiple independent fixed-image-moving-image pairs for separate processing. This approach severs the intrinsic temporal connection of respiratory motion, ignoring the continuity and periodicity of organ motion. While pairwise methods perform reasonably well in handling deformation between two frames, when processing 4D sequences, they often result in a non-smooth deformation field in the temporal dimension and difficulty in accurately predicting complex motion trajectories at intermediate moments, failing to meet the clinical needs for high-precision 4D dose accumulation and dynamic tumor tracking.

[0004] To overcome the temporal defects of pairwise registration, some studies have introduced ConvGRU (Convolutional Gated Recurrent Unit), which captures the long-range dependencies of lung respiration by passing the hidden layer state over time, thereby smoothing the deformation prediction between frames; however, it is prone to problems such as gradient vanishing or excessive computation when processing long sequences.

[0005] Some studies have abandoned the limitation of a single reference frame and adopted a group single-sample learning strategy to implicitly constrain global temporal consistency by jointly optimizing all images in the sequence; however, group registration often struggles to find the optimal balance between preserving high-frequency details and maintaining the global average shape.

[0006] Some studies, based on the theory of Ordinary Differential Equations (ODE), have transformed discrete deformation modeling into continuous velocity field integrals, achieving physical-driven prediction of breathing trajectories. Although the trajectories are smooth, the ability to capture instantaneous changes in complex topology still needs improvement.

[0007] Furthermore, while current multi-scale implicit-explicit hybrid registration networks have achieved breakthroughs in 3D large-deformation registration, they have not yet effectively integrated the aforementioned temporal modeling mechanisms when dealing with 4D data. Existing networks mainly rely on spatial features for frame-by-frame independent registration, ignoring the highly regular temporal correlation of respiratory movements. This results in the deformation state of the previous moment not being able to effectively guide subsequent moments, limiting further improvements in accuracy. At the same time, the lack of trajectory consistency constraints makes it easy for non-physical abrupt changes to occur at anatomical points in the time series, making it difficult to maintain the smoothness and topological consistency of the deformation field in the temporal dimension. Summary of the Invention

[0008] To address the aforementioned issues, this invention proposes a multi-scale implicit-convolutional image registration method and system that enhances temporal features. By introducing a time dimension, it fully explores the spatiotemporal correlations in 4D CT image sequences, achieving high-precision and spatiotemporally continuous modeling of lung respiratory motion.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a multi-scale implicit-convolutional image registration method for enhancing temporal features, comprising: For the acquired 4D CT image sequence of the lungs, a trained registration network is used to calculate the instantaneous deformation field between image pairs at adjacent time points; The registration network's processing includes: encoding the time scalar and frame spacing separately to obtain a hybrid time-coded vector; and concatenating the image spatial coordinates with the hybrid time-coded vector to obtain an implicit global deformation field. After convolution processing of 4D CT images of the lungs, multi-scale feature maps are obtained. Pooling operation is performed on the multi-scale feature maps to obtain channel descriptors. The channel descriptors are fused with the hybrid temporal coding vector, and temporally aware channel weights are generated through a fully connected layer. The multi-scale feature maps are then weighted channel by channel to obtain an explicit residual deformation field. The instantaneous deformation field is obtained by superimposing the implicit global deformation field and the explicit residual deformation field. The cumulative deformation field from the initial moment to the current moment is obtained by cascading the instantaneous deformation fields of all intermediate moments between the initial moment and the current moment, thereby completing image registration.

[0010] As an alternative implementation, the time scalar is encoded as follows: periodic phase encoding is used, and a sinusoidal position encoding mechanism is employed to normalize the time scalar. The mapping is to a high-dimensional feature vector, and the encoding formula is: ; in, For a respiratory half-cycle, L is the maximum number of frequency bands set, and k is the frequency index.

[0011] As an alternative implementation method, the frame spacing... The encoding is: ; in, This is a half-cycle of breathing.

[0012] As an alternative implementation, the image spatial coordinates and the hybrid temporal coding vector are concatenated as features and then mapped by a multilayer perceptron to obtain the implicit global deformation field.

[0013] As an alternative implementation, the total loss function of the registration network during training... for: ; ; in, , Representing the first The image similarity loss function, smoothness loss function, and Jacobian determinant regularization term for each iteration; Let this be the cycle consistency loss function; For the first The weight of the next iteration; I is the total number of iterations; This represents the total number of pixels in the image. Image spatial coordinates; For positive cumulative deformation field, For inverse regression deformation field; The loss function hyperparameters are common to the smoothness loss function and the Jacobian determinant regularization term; is the hyperparameter of the cycle consistency loss function.

[0014] As an alternative implementation method, the weight update process is as follows: Calculate the dynamic decision threshold for the i-th iteration. ; Weights Adjustments will be made: ; Based on normalized inter-frame distance Calculate the distance influence coefficient ; The adjusted weights are updated as follows: ; in, This represents the dynamic decision threshold for the (i-1)th iteration; Let represent the average similarity loss in the i-th iteration; The adjusted weights; The updated weights; The maximum weight is set; The minimum weight is set; The maximum inter-frame distance; This is the minimum frame spacing.

[0015] Secondly, the present invention provides a multi-scale implicit-convolutional image registration system for enhancing temporal features, comprising: The image processing module is configured to use a trained registration network to calculate the instantaneous deformation field between adjacent image pairs on the acquired 4D CT image sequence of the lungs. The registration network's processing includes: encoding the time scalar and frame spacing separately to obtain a hybrid time-coded vector; and concatenating the image spatial coordinates with the hybrid time-coded vector to obtain an implicit global deformation field. After convolution processing of 4D CT images of the lungs, multi-scale feature maps are obtained. Pooling operation is performed on the multi-scale feature maps to obtain channel descriptors. The channel descriptors are fused with the hybrid temporal coding vector, and temporally aware channel weights are generated through a fully connected layer. The multi-scale feature maps are then weighted channel by channel to obtain an explicit residual deformation field. The instantaneous deformation field is obtained by superimposing the implicit global deformation field and the explicit residual deformation field. The registration module is configured to obtain the cumulative deformation field from the initial time to the current time by cascading the instantaneous deformation fields of all intermediate times between the initial time and the current time, thereby completing the image registration.

[0016] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.

[0017] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.

[0018] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: A four-dimensional continuous modeling mechanism with dual spatiotemporal coding is constructed. To address the periodic and non-uniform sampling characteristics of lung respiratory motion, periodic temporal coding and inter-frame distance coding are introduced into the implicit neural representation. This mechanism maps a one-dimensional time scalar to a high-dimensional feature space, endowing the model with a keen perception of respiratory phase and time span, effectively compensating for the shortcomings of traditional discrete-time sampling in capturing continuous dynamic changes. By establishing a continuous mapping relationship between any two frames, the temporal aliasing and trajectory abrupt changes caused by sparse sampling are significantly alleviated, achieving high-resolution temporal motion prediction.

[0020] A multi-scale implicit-convolutional hybrid architecture for spatiotemporal joint optimization is proposed. The spatial multi-scale strategy is extended to the spatiotemporal dimension, constructing a synergistic mechanism of implicit global guidance and explicit local correction: at the coarse scale, implicit neural representations (INRs) with dual encoding are used to capture global motion trends throughout the respiratory cycle as robust temporal priors; at the fine scale, an explicit convolutional neural network (CNN) is combined to extract instantaneous high-frequency textures for residual compensation. This achieves the fusion of long-range temporal memory and short-range spatial details, improving the smoothness and coherence of the deformation field in the temporal dimension.

[0021] An adaptive loss function optimization strategy based on dynamic weights is designed. Addressing the challenge of balancing image similarity accuracy and spatiotemporal topological rationality in 4D registration, a strategy is proposed that adaptively adjusts the weights of regularization and data terms according to the training progress. This strategy guides the model to prioritize the alignment of large anatomical contours in the early stages of training, and automatically strengthens the constraint on spatiotemporal topology as convergence progresses. This effectively avoids the model getting trapped in local optima prematurely, ensuring rapid convergence while optimizing the physical rationality of lung motion trajectories and significantly reducing the proportion of negative Jacobian determinants (i.e., folding rate) in the deformation field.

[0022] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of a 4D CT image sequence of the lungs provided in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the inter-frame-cross-frame composite registration strategy provided in Embodiment 1 of the present invention; Figure 3 This is a flowchart of the multi-scale implicit-convolutional registration network for enhancing temporal features provided in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of a multi-scale temporal fusion network provided in Embodiment 1 of the present invention; Figure 5 The intensity difference map of Case 8 in the DIR-Lab dataset provided in Embodiment 1 of the present invention before registration at times T00 and T50; Figure 6 The intensity difference map of Case 8 in the DIR-Lab dataset provided in Embodiment 1 of the present invention after registration using a multi-scale implicit-convolutional registration network at times T00 and T50. Figure 7 The intensity difference map of Case 8 in the DIR-Lab dataset provided in Embodiment 1 of the present invention after registration with the MSTI-NoTemporal Encoding variant at times T00 and T50; Figure 8 The intensity difference map of Case 8 in the DIR-Lab dataset provided in Embodiment 1 of the present invention after MSTI-Pairwise variant registration at times T00 and T50; Figure 9 This is a multi-starting-point closed-loop trajectory diagram of Case 1 in the DIR-Lab dataset provided in Embodiment 1 of the present invention at times T00 and T50; Figure 10 This is the intensity difference map of Case 8 in the POPI dataset provided in Embodiment 1 of the present invention before registration at times T00 and T50; Figure 11 The intensity difference map of Case 8 in the POPI dataset provided in Embodiment 1 of the present invention after registration using a multi-scale implicit-convolutional registration network at times T00 and T50. Figure 12 The intensity difference map of Case 8 in the POPI dataset provided in Embodiment 1 of the present invention after registration with the MSTI-NoTemporal Encoding variant at times T00 and T50; Figure 13 The intensity difference map of Case 8 in the POPI dataset provided in Embodiment 1 of the present invention after registration with MSTI-Pairwise variant at times T00 and T50. Detailed Implementation

[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0026] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0027] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. Furthermore, it should be understood that the terms “comprising” and “including”, and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0029] Example 1 The human lungs are a highly dynamic organ, and their shape, position, and texture change significantly due to respiratory movements. Traditional 3D CT can only provide a static anatomical snapshot at a certain moment, making it difficult to capture this dynamic process and easily leading to deviations in target delineation during radiotherapy planning.

[0030] To address this issue, four-dimensional computed tomography (4D CT) technology was developed. Essentially, 4D CT introduces a time dimension into conventional three-dimensional tomography. As the fourth dimension, it coordinates the scan with the patient's respiratory cycle through respiratory gating technology, forming a set of dynamic volume sequences with temporal correlation.

[0031] A standard 4D CT sequence of the lungs typically includes 10 phases (T00 to T90) covering the complete respiratory cycle. T00 usually represents the maximum inspiratory phase, and T50 represents the maximum expiratory phase, with the largest deformation between these two phases containing the most significant motion information of the lung anatomy. Considering the regularity of lung respiratory motion and the challenges of large deformation registration, this embodiment focuses primarily on the critical half-cycle sequence from T00 to T50.

[0032] Specifically, this embodiment selects six key time phases—T00, T10, T20, T30, T40, and T50—as the research object. For example... Figure 1 As shown, Figure 1This example visually demonstrates the dynamic evolution of lung morphology with respiratory phase in the selected sequence. Figure 1 The first row shows the overall changes in the three-dimensional volume of the lungs. It is clear that from T00 (maximum inspiration) to T50 (maximum expiration), the total lung volume gradually decreases, and the chest contour undergoes significant inward contraction. The second row shows two-dimensional sagittal section images at the same time point, further revealing the displacement details of the internal anatomical structures. The sagittal view clearly shows that as time progresses, the diaphragm gradually rises, causing a significant nonlinear compression of the lung base tissue.

[0033] This sequence fully records the unidirectional, continuous, large-scale deformation of the lung from its maximum to minimum volume. Compared to processing adjacent frames with minor deformations, accurately modeling this long-span, large-amplitude continuous deformation process has greater clinical significance for evaluating the robustness of registration algorithms in extreme scenarios. Furthermore, the unidirectional cumulative deformation characteristics inherent in this sequence provide an ideal data foundation for subsequently constructing a temporally closed-loop composite registration strategy.

[0034] Unlike independent static image pairs, 4D CT sequences contain rich spatiotemporal contextual information. A thorough understanding and utilization of these characteristics is a prerequisite for designing high-precision temporal registration algorithms.

[0035] Based on relevant literature review, this embodiment summarizes three core characteristics of this data in the registration task: 1. Strong Spatiotemporal Correlation: In the continuous sequence from T00 to T50, the images are not a random stack, but rather generated by sampling at equal time intervals. This means that the anatomical changes between adjacent time frames (e.g., T10 and T20) are small and smooth, exhibiting a very strong temporal dependence. This correlation indicates that the changes in anatomical structure between the previous time frame (T00 and T50) are subtle and smooth, with a very strong temporal dependence. The anatomical state and deformation trend of ) and the current moment ( Highly coupled. In algorithm design, effectively extracting this prior information from historical moments will help reduce the uncertainty of registration at the current moment.

[0036] 2. Smoothness of the motion trajectory: At the physiological level, the movement of lung tissue points (such as vascular bifurcation points and tumor centers) over time follows biomechanical laws, forming a continuous and smooth motion trajectory without instantaneous positional jumps. This means that in an ideal registration result, the derivative of the deformation field in the time dimension (i.e., the velocity field) should change continuously. This characteristic suggests the necessity of introducing trajectory constraints or temporal smoothing terms into the registration model.

[0037] 3. Quasi-periodicity and reversibility of respiratory movements: Respiratory movements are essentially a cyclical process. Although this embodiment selects a unidirectional process from T00 to T50, physically, the deformation path from inhalation to exhalation and the path from exhalation back to inhalation are macroscopically reversible. That is, after half a cycle of deformation, the lung anatomy theoretically possesses the physical potential to return to its initial state. This inherent characteristic provides a theoretical basis for designing composite registration strategies with loop closure functionality, making it possible to utilize cyclic consistency to self-monitor the accuracy of long-term registration.

[0038] In summary, 4D CT images not only provide high-resolution spatial anatomical information, but their inherent spatiotemporal correlation, trajectory smoothness, and motion reversibility also offer new opportunities to overcome the limitations of traditional pairwise registration. Therefore, this embodiment proposes a composite registration strategy and corresponding network architecture that can rationally utilize these temporal features.

[0039] The main steps include the following: For the acquired 4D CT image sequence of the lungs, a trained registration network is used to calculate the instantaneous deformation field between image pairs at adjacent time points; The registration network's processing includes: encoding the time scalar and frame spacing separately to obtain a hybrid time-coded vector; and concatenating the image spatial coordinates with the hybrid time-coded vector to obtain an implicit global deformation field. After convolution processing of 4D CT images of the lungs, multi-scale feature maps are obtained. Pooling operation is performed on the multi-scale feature maps to obtain channel descriptors. The channel descriptors are fused with the hybrid temporal coding vector, and temporally aware channel weights are generated through a fully connected layer. The multi-scale feature maps are then weighted channel by channel to obtain an explicit residual deformation field. The instantaneous deformation field is obtained by superimposing the implicit global deformation field and the explicit residual deformation field. The cumulative deformation field from the initial moment to the current moment is obtained by cascading the instantaneous deformation fields of all intermediate moments between the initial moment and the current moment, thereby completing image registration.

[0040] In the above method, in view of the continuity and quasi-periodicity of respiratory motion in 4D CT image sequences, this embodiment proposes a multi-scale implicit-convolutional registration network to enhance temporal features. Instead of treating 4D registration as an independent pairwise optimization problem, it models it as a temporal closed-loop process with physical constraints.

[0041] To effectively address the large-amplitude nonlinear deformations accumulated in 4D CT sequences of the lungs, this embodiment combines the advantages of global long-span prediction with local short-time cascading. Specifically, the registration process not only retains the traditional approach of directly predicting long-span deformations (such as directly predicting T00 to T50), but also employs a recursive sequence registration strategy.

[0042] like Figure 2 As shown, based on the assumption of the continuity of respiratory motion, a complex long-term large deformation task is decomposed into a series of continuous short-term small deformation estimation subtasks.

[0043] Specifically, for 4D CT image sequences (In this embodiment) to (Corresponding to T00 to T50), image pairs at adjacent time points are recursively processed in the form of a sliding window. .

[0044] At every moment Registration network We need to focus on capturing the previous moment. up to the current moment Deformation field between This refers to the instantaneous deformation field at time t. This recursive processing method reduces the difficulty of solving a single registration and effectively avoids the feature mismatch problem caused by excessive displacement, thereby accurately capturing subtle dynamic changes in the breathing process while maintaining the integrity of the local topology.

[0045] After obtaining a series of discrete deformation fields at intermediate times, in order to reconstruct the deformation field from the initial time T00 to any other time... The global motion trajectory is obtained, and the deformation field is solved. For any time... Cumulative deformation field It is not a simple linear superposition of instantaneous deformation fields at intermediate moments, but rather a result of deformation fields. Cumulative deformation field at the previous time (t-1) The result is obtained by performing a compound operation.

[0046] Its recursive calculation formula is as follows: (1); in, This is a compound operation; Let be the instantaneous deformation field at time t.

[0047] Based on this recursive relationship, from the initial moment... At any time The cumulative deformation field can be expanded into a cascaded composite of the instantaneous deformation fields at all intermediate moments: (2).

[0048] This completes image registration, integrating a series of tiny local instantaneous deformation fields into a global deformation trajectory spanning half a respiratory cycle, thereby enabling continuous tracking of the dynamic displacement of lung tissue.

[0049] The registration network will be introduced below.

[0050] To achieve high-precision modeling of complex spatiotemporal deformations in 4D CT image sequences, this embodiment constructs as follows: Figure 3 The multi-scale implicit-convolutional registration network (denoted as MSTI-CRN) shown here enhances temporal features. This registration network is composed of, for example, a multi-scale implicit-convolutional registration network (denoted as MSTI-CRN). Figure 4 The multi-scale temporal fusion network shown is constructed with shared network parameters during temporal registration. The multi-scale temporal fusion network employs a hybrid architecture combining explicit and implicit methods, but deeply integrates temporal information into each stage of feature extraction and deformation prediction.

[0051] Specifically: like Figure 4 As shown, the multi-scale temporal fusion network consists of a dual temporal domain encoding module, a spatiotemporal implicit neural representation module, and an explicit convolution module with embedded temporal attention mechanism.

[0052] To enable the network to perceive the periodic phase of respiratory motion and the span between adjacent frames, a dual time-domain coding module is first constructed. Based on the physical characteristics of respiratory motion, this module designs a high-frequency feature mapping mechanism for the time dimension.

[0053] First, the entire coordinate network of the image is downsampled and then input into the dual time-domain encoding module. Considering the cyclical nature of respiratory motion, periodic phase encoding is employed, utilizing a sinusoidal position encoding mechanism to normalize the time scalar. It is mapped to a high-dimensional feature vector.

[0054] Considering the respiratory half-cycle in this embodiment Set to 5, the encoding formula is defined as: (3); Where L is the maximum number of frequency bands set (e.g., L=2), and k is the frequency index.

[0055] This mapping makes time (End of inhalation) and Although they do not overlap in the feature space at the end of exhalation, they maintain a rigorous periodic mathematical correlation.

[0056] Secondly, in order to explicitly inform the network of the current time span of the registration task so as to adaptively adjust the deformation amplitude prediction, inter-frame distance coding is introduced to encode the frame interval. Linear normalization is: (4); Ultimately, the concatenation of the two generates a hybrid temporal encoding vector. It is then used as a global condition variable and synchronously injected into subsequent implicit and explicit networks, providing temporal context information for the entire spatiotemporal registration task.

[0057] At the coarse-resolution level of the network, this embodiment designs a spatiotemporal implicit neural representation module to capture global motion trends throughout the entire respiratory cycle. Unlike traditional INR, which only uses the spatial coordinates of lung images... As input, this module constructs a spatiotemporal conditional neural field, which uses a multilayer perceptron (MLP) to combine spatial coordinates with a hybrid temporal encoded vector. Perform feature concatenation.

[0058] The mapping relationship is expressed as follows: (5); Of particular note is that, to ensure the differentiability of the deformation field in the continuous domain and the high-fidelity reconstruction of the signal, this MLP employs a periodic activation function. This design gives the network-generated deformation field the excellent mathematical property of being differentiable of any order. Combining the inherent low-frequency priority spectral characteristics of the MLP network and the strong temporal prior containing phase and span information, this module can effectively filter out local noise and focus on fitting the time-varying global large-amplitude nonlinear deformation field, thus obtaining the implicit global deformation field. This design not only ensures the smoothness of the generated deformation field in the time dimension, but also provides a physically reasonable initial deformation field for subsequent fine registration, thereby reducing the topological folding rate in large deformation scenarios.

[0059] To compensate for the lack of implicit representation in high-frequency texture details at the fine-scale level of the network, this embodiment constructs an explicit convolutional module with an embedded temporal attention mechanism. This module is based on the U-Net architecture, and its core innovation lies in the temporal improvement of the traditional SE (Squeeze-and-Excitation Networks) module, proposing a temporally-aware SE module (TA-SE).

[0060] In this embodiment, the lung 4D CT image sequence to be registered uses a predetermined phase as the moving image and the remaining phases as the fixed images. After processing by a Convolutional Neural Network (CNN), the multi-scale feature maps obtained after convolution are... Perform global average pooling to obtain channel descriptors. Then it is combined with the mixed temporal encoding vector Fusion, through learning and generating time-aware channel weights via fully connected layers. : (6); in, For the Sigmoid function, It is the ReLU activation function; , These are the weights for the fully connected layer.

[0061] Among them, such as Figure 4 As shown, the processing procedure of the SE module includes: (1) Spatial information compression (Squeeze). The receiving dimension is... The 3D input feature map is used, and through pooling operations (usually global average pooling), the spatial pixels of height h and width w within each channel are compressed into a single value, thereby obtaining the channel descriptor. .

[0062] This step refines the originally bulky 3D feature map into a single image of length [length missing]. The one-dimensional feature vector is used to extract the global spatial context information of the current frame.

[0063] (2) Fusion of temporal information. This step does not directly use the spatial feature vector (i.e., channel descriptor) obtained in step (1). Instead, it introduces a hybrid time-coded vector that incorporates both periodic coding and distance coding. Through the concatenation operation, this With channel descriptor Merging along the channel dimension reduces the vector length from Extended to This tightly binds spatial features with the contextual information of the time series.

[0064] (3) Channel weight calculation (Excitation). The long vector that incorporates spatiotemporal information is fed into a small neural network consisting of a fully connected layer (FC) and an activation function. Specifically: it first passes through the first fully connected layer for dimensionality reduction and ReLU activation to fuse features, and then passes through the second fully connected layer for dimensionality increase and a Sigmoid activation function to remap the dimension back to the original number of channels. The output values ​​are then normalized to a range of 0 to 1. The final generated... The vector is the importance score of each channel calculated based on both spatiotemporal cues, which is also the channel weight. .

[0065] (4) Feature rescaling (Scale). The calculated channel weights are then scaled. , with the original input The feature map is multiplied channel by channel. In this way, time-aware channel weights can be used to automatically amplify feature channels that are more useful for the current task, while suppressing unimportant noise channels, and finally outputting a brand-new feature map with optimized weights.

[0066] Therefore, the feature maps obtained after convolution processing at different scales are transformed into a dynamic, time-varying system through the improved SE module. This system can dynamically reweight the response intensity of feature channels according to the current respiratory phase (such as inspiration or expiration), thereby adaptively capturing the instantaneous local deformation of lung microstructures (such as vascular terminals and bronchial walls). Through upsampling and splicing of skip connections, followed by two convolutions, a high-precision explicit residual deformation field is output. .

[0067] In summary, the multi-scale implicit-convolutional hybrid network for enhancing temporal features proposed in this embodiment transforms discrete temporal information into continuous high-dimensional features through dual temporal encoding, and achieves joint optimization of the spatiotemporal dimensions using a complementary strategy combining explicit and implicit methods. The final instantaneous deformation field is derived from the implicit global deformation field after trilinear upsampling. With explicit residual deformation field Composite superposition ( It is formed by ).

[0068] This end-to-end hybrid architecture not only maintains its advantages in large deformation topologies, but also significantly solves the problems of temporal jitter and trajectory discontinuity in 4D CT registration through deep embedding of temporal priors and dynamic attention mechanisms, achieving high-precision and robust modeling of lung respiratory motion.

[0069] To address the closed-loop temporal characteristics of 4D CT, this embodiment introduces a cyclic consistency loss as the core spatiotemporal topological constraint. This loss aims to eliminate the accumulated error in long-term recursive predictions. Based on the aforementioned composite strategy, the registration network completes the process from... arrive After the forward recursive registration, further prediction is needed from... Return to The reverse deformation. Ideally, the forward cumulative deformation field. With reverse regression deformation field The superposition of vectors should strictly approximate the zero vector.

[0070] Therefore, the cycle consistency loss function Defined as the root mean square error of the closed-loop residual: (7); in, The total number of pixels or voxels in the image; These are the spatial coordinates (pixel or voxel positions) in the image.

[0071] This constraint forces the registration network to learn that the forward and reverse deformations are inverse mappings of each other, which enhances the physical interpretability of the model for large deformation trajectories under unsupervised conditions.

[0072] In addition, to drive end-to-end training of the spatiotemporal multi-scale network, this embodiment constructs a composite objective function with multiple constraints. Given the significant differences in deformation amplitude at different time steps in 4D CT respiratory sequences, using fixed weights would lead to overfitting of the network on simple steps and underfitting on difficult steps with large deformations.

[0073] Therefore, this embodiment defines a total loss function. The weighted sum of the losses at all recursive time steps, superimposed with the global loop consistency constraint: (8); in, , Representing the first The image similarity loss function, smoothness loss function, and Jacobian determinant regularization term for each iteration; Let I be the cycle consistency loss function; I is the total number of iterations. The loss function hyperparameters are common to the smoothness loss function and the Jacobian determinant regularization term; represents the hyperparameters of the cycle consistency loss function; Representing the The dynamic weight coefficients of each iteration are not fixed hyperparameters, but are updated in real time during training through the performance-and-motion adaptive weighting (PMAW) strategy proposed in this embodiment, so as to dynamically balance the learning weights at different times.

[0074] Image similarity loss function Normalized cross-correlation (NCC) was used as a similarity measure to evaluate the similarity between the deformed moving image and the stationary image. (9); in, p For fixed image The point in the middle, A set of points; Indicated by point The set of points within a local window of a cube with a side length of n, centered at n (set in this embodiment) ); and Representing fixed images With deformed images Average intensity within a local window; For deformation field; for; To use points in a fixed image The pixel value of each point within a cube local window centered at n with side length n.

[0075] Using negative NCC as a similarity loss measure, the smaller the loss value, the more similar the two images are.

[0076] Smoothness loss function Abrupt local mutations are suppressed by penalizing the spatial gradient of the deformation field. For multi-scale network architectures, this constraint is imposed on the deformation field at all output scales. Specifically, it is defined as the spatial gradient of the deformation field. Sum of norms: (10); in, Representative scale quantity (corresponding) ), Indicates the first Deformation fields at various scales Spatial gradient at a given location.

[0077] To further ensure the topology preservation properties of the deformation field and prevent tissue folding or penetration, a Jacobian determinant regularization term is introduced. For deformation fields Its point Jacobian matrix at the location ,in The identity matrix describes the local volume variation. When the Jacobian determinant... When this happens, it means that a topological fold has occurred at that location.

[0078] Therefore, the Jacobian determinant regularization term aims to penalize all non-positive determinant values: (11).

[0079] To guide the model to automatically focus on challenging respiratory phases, this strategy first introduces an Exponential Weighted Moving Average (EWMA) algorithm to track the trend of historical similarity loss and calculate a dynamic decision threshold. .

[0080] Based on this threshold, the weights for the current time step are adjusted using performance feedback: (12); in, This represents the dynamic decision threshold for the i-th iteration; This represents the dynamic decision threshold for the (i-1)th iteration; Let represent the average similarity loss in the i-th iteration; The adjusted weights; The maximum weight is set; This is the minimum weight set.

[0081] When the registration error of a certain step exceeds the historical threshold, it is classified as a "difficult case" and its weight is significantly amplified. Conversely, the weight is reduced ( ), ).

[0082] This mechanism is similar to course learning, forcing the network to automatically allocate computing resources to the weakest link where the current registration effect is the worst, thereby avoiding local optima.

[0083] In addition to performance feedback, this strategy further incorporates the physical prior of inter-frame time span to address the closed-loop regression step (T50). T00) is a problem caused by large deformation due to its large span.

[0084] Based on normalized inter-frame distance (Right now Calculate the distance influence coefficient And inject it into the final weight update formula: (13); in, The maximum inter-frame distance; This represents the minimum inter-frame distance.

[0085] By introducing The gain term explicitly assigns a higher optimization priority to registration tasks with large spans. This composite weighting mechanism, which combines data-driven and physical perception (motion prior), effectively solves the optimization problem caused by uneven deformation amplitude in long-term registration, ensuring that the model can achieve high-precision topological closure throughout the entire breathing cycle.

[0086] To fully verify the performance of the registration network proposed in this embodiment in handling complex respiratory movements of the lungs, two internationally available benchmark datasets with complementary characteristics were selected: the DIR-Lab dataset and the POPI dataset. The DIR-Lab dataset contains 10 chest 4D-CT image sequences (Case 1-10), each covering 10 phases (T00-T90) of a complete respiratory cycle, with an image resolution of [missing information]. to As a benchmark for large deformation registration, this dataset not only provides 300 pairs of expert-annotated anatomical landmarks between the maximum inspiratory phase (T00) and the maximum expiratory phase (T50) to evaluate registration accuracy under extreme deformation, but more importantly, for each case, the dataset also provides an additional temporal tracking subset containing 75 feature points, which are precisely located in all 10 respiratory phases. This feature allows the DIR-Lab dataset to not only validate end-to-end large deformation capabilities but also effectively evaluate the algorithm's accuracy in capturing continuous motion trajectories throughout the entire respiratory process.

[0087] To further enhance the assessment of temporal consistency and trajectory smoothness, the POPI dataset was introduced. This dataset contains six high-quality 4D-CT lung image sequences acquired under respiratory gating, each consisting of 10 temporal phases of 3D-CT. For the first three patients, experts identified and tracked dense anatomical landmarks (approximately 40-100 per patient) across all 10 temporal phases. For the latter three patients, the landmarks were primarily distributed during the extreme phases of maximal inspiration and expiration.

[0088] In the data preprocessing stage, to meet the input specifications of the deep learning network and focus on the lung motion region, a standardized cleaning process was performed on the two datasets mentioned above. First, a threshold segmentation method was used to remove background air, treatment bed panels, and extrathoracic tissue from the CT images, and the smallest bounding box containing the complete lung field was cropped based on the generated lung mask to reduce invalid background calculations. Second, considering the grayscale characteristics of CT images, the Henlein unit (HU) values ​​of the images were truncated to a certain value. Within this range, encompassing the main density regions from lung parenchyma to soft tissue, it was subsequently normalized to a linear mapping. To eliminate the interference of abnormally highlighted skeletal tissue on similarity metrics. Finally, for the temporal training task, the processed 3D images are reassembled into 4D sequence groups in chronological order. The corresponding inter-frame temporal distance vector is calculated and used as input data for the dual time coding module.

[0089] All experiments were performed on a unified high-performance computing platform to ensure the fairness and reproducibility of the evaluation results. Regarding the network training strategy, considering the individual anatomical differences in lung 4D-CT data, this experiment adopted a single-sample instance optimization method, i.e., independent end-to-end training was performed on each 4D-CT sequence without relying on large-scale external datasets for pre-training. The training process used the Adam optimizer, which combines momentum and adaptive learning rate, with an initial learning rate set to [value missing]. Furthermore, a StepLR learning rate scheduling strategy was introduced to achieve more refined parameter search in the later stages of training. The total number of training iterations for each case was set to 1500, during which the global random seed was fixed at 3407 to eliminate randomness interference. Meanwhile, the initial values ​​of the spatial smoothness weight and the cycle consistency weight were set to 0.5 and 0.1, respectively, and dynamically adjusted with subsequent adaptive strategies to ensure that the model efficiently completes deep learning of spatiotemporal features and deformation field prediction.

[0090] In the evaluation methodology phase, the target registration error (TRE) and Jacobian determinant folding rate were considered. In addition, to overcome the limitations of traditional single closed-loop evaluation (which only validates T00), further consideration was given to... The T50 timeframe can easily mask the limitations of intermediate phase drift. Therefore, a multi-start closed-loop evaluation strategy is proposed for the ablation experiment. Conventional closed-loop indices often fixate on the maximum inspiratory phase (T00) as the starting point, which can lead to overfitting of the model only on specific paths, neglecting the topological stability at other points in the respiratory cycle. This method introduces a polling mechanism, polling each phase from T00 to T50. Each frame is sequentially set as an independent starting reference frame. For any selected starting phase... A complete closed-loop trajectory is constructed, evolving from the current phase to the maximum expiratory phase (T50) and then regressing back to the initial phase. This design simulates the physical process of whether lung tissue can accurately reposition itself after undergoing large deformations from any respiratory state, thereby enabling a comprehensive detection of the cumulative error and robustness of the registration algorithm in long-term prediction.

[0091] Based on the above strategy, this embodiment defines the multi-starting-point closed-loop consistency error (MS-LCE) as the average of the landmark point regression errors of all starting phases. It is assumed that the sequence contains... The kth time phase An initial phase with a closed-loop variable field. It is composed of forward cumulative deformation and backward regression deformation.

[0092] The formula for calculating MS-LCE is as follows: (14); in, This represents the total number of anatomical landmarks during that time period. The actual coordinates of the road sign points. These are the predicted coordinates after closed-loop transformation. This represents the Euclidean distance.

[0093] This metric does not rely on any supervision signals during the registration process; it is a purely physical consistency measure. A lower MS-LCE value means that the model is not only accurate at the endpoints but also maintains extremely high topological closure and trajectory smoothness across any cross-section of the entire respiratory trajectory. It is a standard for measuring whether a 4D registration algorithm has grasped the laws of respiratory motion.

[0094] (1) Quantitative evaluation and comparative analysis of the DIR-Lab dataset.

[0095] Table 1 presents the quantitative evaluation results of the proposed Multi-Scale Implicit-Convolutional Registration Network (MSTI-CRN) for enhancing temporal features in this embodiment on 10 cases of the DIR-Lab dataset, and provides a comprehensive comparison with current mainstream registration algorithms. The comparison methods include pairwise registration networks based on deep learning (such as VoxelMorph, an open-source deep learning framework designed for medical image registration, whose core lies in using convolutional neural networks to achieve fast and accurate unsupervised image alignment), and temporal registration networks (such as GroupRegNet (a deep learning model for processing time-series images, such as dynamic medical image alignment tasks, which achieves cross-time point structural change analysis or motion tracking by accurately aligning images at different time points in the sequence) and Lung-CRNet (a convolutional recurrent neural network for registering four-dimensional computed tomography images of the lungs)). The evaluation metrics mainly focus on the ratio of target registration error (TRE) to negative Jacobian determinant (Folds(%), representing the topological rationality of the deformation field, i.e. (proportion).

[0096] From the TRE metric, the method in this embodiment demonstrates significant advantages in handling large-scale lung motion. First, compared with basic deep learning pairwise registration methods, the average TRE of the method in this embodiment is significantly reduced. This is mainly attributed to the combination of multi-scale strategies and implicit neural representations, which enables the network to effectively capture the drastic nonlinear deformations in the DIR-Lab dataset (especially Cases 6-10).

[0097] Secondly, compared with temporal registration networks, the method in this embodiment still maintains leading or highly competitive performance. Although temporal registration methods also introduce temporal dimension information, the dual temporal encoding combined with the closed-loop recursive strategy proposed in this embodiment can more accurately characterize the continuous changes in respiratory phase. Experimental data show that the average TRE index of this embodiment is reduced by approximately 0.10 mm compared to the temporal baseline model, especially in the large-span registration task between maximal inspiration and maximal expiration, where it exhibits stronger robustness.

[0098] While ensuring high accuracy, the physical rationality of the deformation field is a key indicator for evaluating the clinical usability of 4D registration algorithms. From Table 1, the ratio of negative values ​​in the Jacobian determinant (…) As can be seen, the method in this embodiment performs excellently in terms of topology preservation, with an average folding rate of only 0.003%.

[0099] Experimental results demonstrate that the method in this embodiment achieves lower TRE while maintaining an extremely low folding rate. This indicates that the deformation field generated by this embodiment is not only numerically accurate but also more consistent with the biomechanical characteristics of lung tissue in terms of anatomical structure, thus exhibiting higher safety for clinical application.

[0100] Table 1. Comparison of TRE and folding rate of different methods on the Dir-Lab dataset; .

[0101] (2) Quantitative evaluation and comparative analysis of the POPI dataset.

[0102] Table 2 shows the performance of the method in this embodiment on the POPI dataset. Compared to the DIR-Lab dataset, the POPI dataset is cited less frequently in existing deep learning registration research, and there are fewer benchmark methods available for direct comparison. Nevertheless, to verify the effectiveness of the algorithm, representative algorithms such as Vandemeulebroucke (Vandemeulebroucke is involved in medical image registration work, mainly focusing on medical image registration, deep learning registration networks, etc.), GDL-FIRE (4D medical image registration framework), and GroupRegNet were selected for horizontal comparison.

[0103] From the TRE metric, the spatiotemporal multi-scale network proposed in this embodiment achieved excellent registration accuracy in all six datasets, with the average TRE reduced to 0.94 mm. Experimental results show that the method in this embodiment performs well at the moment of maximum deformation while maintaining low error. Compared with similar networks, the method in this embodiment exhibits strong robustness in both deformation field smoothness and alignment accuracy.

[0104] Regarding the physical plausibility of the deformation field, existing literature generally lacks quantitative reports on the Jacobian determinant folding rate on the POPI dataset, making it difficult to conduct fair cross-benchmark comparisons. Nevertheless, preliminary records from this experiment show that the deformation field generated by the method in this embodiment has an average folding rate of 0.0015%, ensuring trajectory accuracy without producing significant non-physical deformation. Given that topology preservation capability is crucial for evaluating the clinical value of 4D registration algorithms, and that a single static folding rate cannot fully reflect temporal dynamic characteristics, a deeper analysis of topological stability and spatiotemporal consistency will be discussed in detail during the ablation experiments.

[0105] Table 2. Comparison of TREs for different methods on the POPI dataset; .

[0106] (3) Ablation experiment.

[0107] To systematically decouple and quantify the effectiveness of the core components of the Multi-Scale Implicit-Convolutional Registration Network (MSTI-CRN) for enhancing temporal features proposed in this embodiment, a series of controlled variable experiments were designed, focusing first on the verification of training strategies and temporal mechanisms. To enhance the temporal awareness of the input, an MSTI-No Temporal Encoding variant was constructed. This model removes the dual temporal domain encoding module, aiming to explore the network's ability to capture the dynamic patterns of respiratory motion in the absence of explicit phase and frame spacing priors. Regarding constraint mechanisms, an MSTI-No CycleLoss variant was set up. By eliminating the cycle consistency loss term, the contribution of closed-loop topology constraints in suppressing the drift of accumulated errors in long-term recursive predictions was evaluated. Simultaneously, to demonstrate the superiority of the 4D-CT temporal registration strategy over traditional methods, an MSTI-Pairwise variant was designed. This variant transforms the training mode into direct pairwise registration, directly predicting large deformations from T00 to T50 without using intermediate frames as relays. This demonstrates the core role of this recursive strategy in reducing optimization difficulty and improving robustness to large deformations.

[0108] In addition, ablation analysis at the network architecture level verifies the irreplaceability of each branch in this hybrid architecture combining explicit and implicit representations. To this end, a variant of MSTI-NoMultiScale-CNN is constructed, which removes the coarse-scale spatiotemporal implicit neural representation module and relies solely on U-Net for feature extraction. This aims to reveal the key role of the low-frequency priority characteristic of MLP in constructing a smooth global anatomical skeleton and preventing large deformation topological folding. Correspondingly, the variant MSTI-NoMultiscale-INR removes the fine-scale temporal-aware U-Net module, retaining only the output of the implicit network as the final deformation field. This demonstrates the necessity of explicit convolutional features in capturing high-frequency texture details such as pulmonary vessels and bronchial terminal lesions, as well as correcting local nonlinear alignment errors.

[0109] To validate the contributions of each module in challenging anatomical scenarios, the ablation experiments focused on the DIR-Lab dataset. This dataset was chosen as the primary analysis object because DIR-Lab contains the extreme deformation of the lungs from maximal inspiration to maximal expiration (with maximum displacement exceeding 30 mm), and this intense nonlinear motion is optimal for testing temporal consistency and topological robustness.

[0110] Table 3 lists the TREs of different ablation variants on 10 cases in the DIR-Lab dataset, focusing on evaluating the alignment accuracy of extreme deformations between maximum inspiration (T00) and maximum expiration (T50). Overall, MSTI-CRN exhibits the best comprehensive performance, with an average TRE reduced to 1.16 ± 0.80 mm, outperforming other variants. Specifically, temporal awareness has the most significant impact on registration accuracy; after removing the dual temporal coding module, the average TRE is 1.68 mm, particularly in the large deformation case Case 8, where the error ranges from 1.24 mm to 4.05 mm. The explicit phase and inter-frame spacing priors are crucial for the network to handle complex respiratory motions and distinguish different deformation amplitudes. Secondly, the choice of registration strategy is equally important; the MSTI-Pairwise variant, with its pairwise registration, achieves an average TRE of 1.58 mm. This indicates that directly predicting large-span deformations from T00 to T50 is prone to getting trapped in local extrema. The recursive streaming strategy proposed in this embodiment reduces the optimization difficulty and improves robustness under large-amplitude motion by using intermediate frames as deformation relays.

[0111] Table 3. Registration results of different ablation models in the DIR-Lab dataset; .

[0112] To visually evaluate the actual performance of different strategies in anatomical structure alignment Figures 5-8 The image shows the intensity difference before and after registration at times T00 and T50 in Case 8. Four representative slices along the Z-axis (Z=25, 45, 60, 85) were selected. Figures 5-8 In the diagram, (a) represents a slice with Z=25, (b) represents a slice with Z=45, (c) represents a slice with Z=60, and (d) represents a slice with Z=85. The blue and red areas represent areas with larger registration errors, while the white or light gray areas represent areas with precise alignment.

[0113] For example Figure 5 The original unregistered image (00-50) shows a significant, large-area error region in dark color, as shown in the image. Figure 6 The difference map generated by the complete model (MSTI-CRN) shown exhibits a relatively pure gray-white tone, indicating a high degree of overlap in lung texture and diaphragmatic boundaries. In contrast, as... Figure 7 The MSTI-No Temporal Encoding variant shown is similar to... Figure 8The MSTI-Pairwise variant shown still exhibits noticeable red and blue spots near the diaphragm, corresponding to its higher TRE value. This intuitively reflects the lack of temporal priors or recursive guidance, leading to local mismatches in areas of large-amplitude movement. Furthermore, the visual differences between the MSTI-No Cycle Loss variant and the MSTI-No Dynamic Weights variant (variants without dynamic weights) are difficult to distinguish with the naked eye. However, in slices with dense blood vessels, the texture residuals of the complete model still appear smoother and weaker. Combining the quantitative TRE data with the qualitative visualization results demonstrates the necessity of introducing temporal encoding and recursive strategies to solve the registration problem in lungs with large deformations.

[0114] To verify the trajectory invertibility and topological stability of the model in long-term recursive prediction, the multi-starting-point loop closure consistency error (MS-LCE) was evaluated on the DIR-Lab dataset. Table 4 summarizes the average loop closure error of each ablation variant across 10 cases.

[0115] Table 4. Average loop closure error of each ablation variant across 10 cases; .

[0116] As shown in Table 4, MSTI-CRN significantly outperformed all variants, with an average MS-LCE of only 0.94 ± 0.63 mm across all cases. In contrast, the MSTI-No Temporal Encoding variant exhibited more severe instability, with an average error of 1.62 ± 0.97 mm. These results indicate that without explicit phase and inter-frame spacing guidance, the network is prone to temporal ambiguity when performing long-term recursive predictions, making it difficult to accurately pinpoint the relative position of the current deformation within the breathing cycle. Furthermore, the MSTI-No Cycle loss variant led to an increase in the average error to 1.25 ± 0.06 mm, an increase of approximately 35.9%, indicating that… This is a crucial mechanism for ensuring that forward and reverse deformations are inversely mapped to each other, guaranteeing the physical reversibility of the closed-loop trajectory. Regarding network architecture and optimization strategies, the integrity of the hybrid architecture is equally important for maintaining long-term robustness. The MSTI-No MultiScale-CNN variant, which removes the implicit neural representation module, has an average loop closure error of 1.37 ± 1.29 mm. This result is higher than that of the MSTI-No MultiScale-INR variant, which removes the explicit convolutional module. This indicates that relying solely on explicit U-Net is insufficient to construct an accurate global anatomical skeleton to cope with large-scale nonlinear deformations over long time. Implicit neural representations must be introduced to provide smooth and continuous global deformation field constraints. Meanwhile, the results of the MSTI-No Dynamic Weights variant are slightly inferior to the complete model, indicating that the adaptive weight strategy plays a fine-tuning role in balancing the learning contributions of different difficulty phases.

[0117] To visually verify the model's trajectory tracking accuracy in continuous time series, Case 1 from the DIR-Lab dataset was selected for multi-starting-point closed-loop trajectory visualization, such as... Figure 9 As shown, six closed-loop evolution processes with T00 to T50 as the starting time phases are illustrated. Figure 9 (a)-(f) in the figure represent the motion trajectories from time T00 to time T50, respectively. It can be seen that the trajectory lines exhibit a high degree of overlap across the six initial phases, confirming the high motion capture fidelity of the proposed network. Whether using the maximum inspiratory phase (T00) or intermediate transition phases (such as T20, T30) as the loop closure starting point, the predicted trajectory closely follows the actual anatomical path and regresses accurately, without significant path deviation. In summary, MSTI-CRN, through the organic integration of temporal coding, cyclic constraints, and an implicit-explicit hybrid architecture, achieves sub-millimeter-level loop closure consistency across all test cases and initial phases, possessing the spatiotemporal reliability required for clinical 4D sequence analysis.

[0118] To further verify the generalization ability of the method in this embodiment on different data sources, the ablation experiments described above were repeated on the POPI dataset, focusing on registration accuracy and image alignment quality. Overall, the experimental results on the POPI dataset show a performance trend consistent with DIR-Lab. As shown in Table 5, in terms of quantitative evaluation, MSTI-CRN still maintains the lowest average TRE, outperforming other variants that remove the core module, demonstrating the general effectiveness of the spatiotemporal recursive architecture in processing lung images from different sources. Qualitative results are as follows... Figures 10-13 As shown, the intensity difference diagrams before and after registration at times T00 and T50 in Case 8 are displayed. Four representative slices along the Z-axis (Z=25, 45, 60, 85) are selected. Figures 10-13In the diagram, (a) represents the slice with Z=25, (b) represents the slice with Z=45, (c) represents the slice with Z=60, and (d) represents the slice with Z=85. By comparing the intensity difference maps before and after registration, it is clearly observed that the difference map generated by the complete model is the smoothest and cleanest compared to the high-intensity error area (red and blue spots) remaining at the lung wall edge in the ablation variant. This result corroborates the previous conclusions, further demonstrating the contribution of strategies such as dual temporal encoding and cyclic consistency to improving the accuracy of respiratory motion registration.

[0119] Table 5. Registration results of different ablation variants in the POPI dataset; .

[0120] This embodiment addresses the problem of temporal jitter and voxel trajectory mismatch in deformation field caused by the correlation of the fragmented temporal dimension in 4D-CT sequence registration. It expands the research dimension from spatial to spatiotemporal, proposing a multi-scale implicit-convolutional registration network (MSTI-CRN) to enhance temporal features. This method no longer treats 4D registration as an independent pairwise optimization problem, but rather models it as a physically constrained temporal closed-loop process, achieving high-precision, spatiotemporally continuous modeling of lung respiratory motion.

[0121] The core work is mainly reflected in the following three dimensions: First, a dual temporal domain encoding mechanism is constructed for basic representation and temporal feature extraction. By introducing periodic phase encoding and inter-frame distance encoding, a one-dimensional time scalar is mapped to a high-dimensional feature space, giving the implicit network an explicit perception of breathing phase and time span. Combined with explicit CNN extraction of high-frequency textures for residual compensation, this effectively compensates for the shortcomings of traditional discrete-time sampling in capturing continuous dynamic changes.

[0122] Secondly, regarding network topology and optimization strategies, a multi-scale implicit-convolutional hybrid architecture for spatiotemporal joint optimization is proposed. Combined with a recursive streaming registration strategy and cyclic consistency physical constraints, a closed-loop optimization framework for forward cumulative deformation and backward regression deformation is established. Furthermore, an adaptive loss function optimization strategy based on dynamic weights is designed to automatically adjust the regularization and data term weights at different time points during training, effectively suppressing error drift in long-term multi-step prediction.

[0123] Finally, for experimental validation and performance evaluation, extensive comparative and ablation experiments were conducted on the DIR-Lab and POPI datasets. By introducing a multi-starting-point closed-loop consistency evaluation strategy (MS-LCE), the regression accuracy of the model at the start of any respiratory phase was comprehensively tested. Experimental results show that the MSTI-CRN model not only outperforms mainstream pairwise and temporal registration algorithms in target registration error, but also exhibits good trajectory smoothness and topological closure in long-term tracking, generating a dynamic deformation field that meets the requirements of biomechanical continuity.

[0124] Example 2 This embodiment provides a multi-scale implicit-convolutional image registration system for enhancing temporal features, including: The image processing module is configured to use a trained registration network to calculate the instantaneous deformation field between adjacent image pairs on the acquired 4D CT image sequence of the lungs. The registration network's processing includes: encoding the time scalar and frame spacing separately to obtain a hybrid time-coded vector; and concatenating the image spatial coordinates with the hybrid time-coded vector to obtain an implicit global deformation field. After convolution processing of 4D CT images of the lungs, multi-scale feature maps are obtained. Pooling operation is performed on the multi-scale feature maps to obtain channel descriptors. The channel descriptors are fused with the hybrid temporal coding vector, and temporally aware channel weights are generated through a fully connected layer. The multi-scale feature maps are then weighted channel by channel to obtain an explicit residual deformation field. The instantaneous deformation field is obtained by superimposing the implicit global deformation field and the explicit residual deformation field. The registration module is configured to obtain the cumulative deformation field from the initial time to the current time by cascading the instantaneous deformation fields of all intermediate times between the initial time and the current time, thereby completing the image registration.

[0125] It should be noted that the above modules correspond to the steps described in Embodiment 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0126] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in Embodiment 1. For brevity, further details are omitted here.

[0127] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0128] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.

[0129] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.

[0130] The method in Example 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.

[0131] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.

[0132] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.

[0133] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.

[0134] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.

[0135] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.

[0136] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A multi-scale implicit-convolutional image registration method for enhancing temporal features, characterized in that, include: For the acquired 4D CT image sequence of the lungs, a trained registration network is used to calculate the instantaneous deformation field between image pairs at adjacent time points; The registration network's processing includes: encoding the time scalar and frame spacing separately to obtain a hybrid time-coded vector; and concatenating the image spatial coordinates with the hybrid time-coded vector to obtain an implicit global deformation field. After convolution processing of 4D CT images of the lungs, multi-scale feature maps are obtained. Pooling operation is performed on the multi-scale feature maps to obtain channel descriptors. The channel descriptors are fused with the hybrid temporal coding vector, and temporally aware channel weights are generated through a fully connected layer. The multi-scale feature maps are then weighted channel by channel to obtain an explicit residual deformation field. The instantaneous deformation field is obtained by superimposing the implicit global deformation field and the explicit residual deformation field. The encoding of the time scalar is as follows: periodic phase encoding is used, and a sinusoidal position encoding mechanism is employed to normalize the time scalar. The mapping is to a high-dimensional feature vector, and the encoding formula is: ; For frame spacing The encoding is: ; in, For a respiratory half-cycle, L is the maximum set number of frequency bands, and k is the frequency index; The cumulative deformation field from the initial moment to the current moment is obtained by cascading the instantaneous deformation fields of all intermediate moments between the initial moment and the current moment, thereby completing image registration.

2. The multi-scale implicit-convolutional image registration method for enhancing temporal features as described in claim 1, characterized in that, After concatenating the image spatial coordinates with the hybrid temporal coding vector, the implicit global deformation field is obtained by mapping through a multilayer perceptron.

3. The multi-scale implicit-convolutional image registration method for enhancing temporal features as described in claim 1, characterized in that, The total loss function of the registration network during training for: ; ; in, , Representing the first The image similarity loss function, smoothness loss function, and Jacobian determinant regularization term for each iteration; Let this be the cycle consistency loss function; For the first The weight of the next iteration; I is the total number of iterations; This represents the total number of pixels in the image. Image spatial coordinates; For positive cumulative deformation field, For inverse regression deformation field; The loss function hyperparameters are common to the smoothness loss function and the Jacobian determinant regularization term; is the hyperparameter of the cycle consistency loss function.

4. The multi-scale implicit-convolutional image registration method for enhancing temporal features as described in claim 3, characterized in that, The weight update process is as follows: Calculate the dynamic decision threshold for the i-th iteration. ; Weights Adjustments will be made: ; Based on normalized inter-frame distance Calculate the distance influence coefficient ; The adjusted weights are updated as follows: ; in, This represents the dynamic decision threshold for the (i-1)th iteration; Let represent the average similarity loss in the i-th iteration; The adjusted weights; The updated weights; The maximum weight is set; The minimum weight is set; The maximum inter-frame distance; This is the minimum frame spacing.

5. A multi-scale implicit-convolutional image registration system for enhancing temporal features, characterized in that, include: The image processing module is configured to use a trained registration network to calculate the instantaneous deformation field between adjacent image pairs on the acquired 4D CT image sequence of the lungs. The registration network's processing includes: encoding the time scalar and frame spacing separately to obtain a hybrid time-coded vector; and concatenating the image spatial coordinates with the hybrid time-coded vector to obtain an implicit global deformation field. After convolution processing of 4D CT images of the lungs, multi-scale feature maps are obtained. Pooling operation is performed on the multi-scale feature maps to obtain channel descriptors. The channel descriptors are fused with the hybrid temporal coding vector, and temporally aware channel weights are generated through a fully connected layer. The multi-scale feature maps are then weighted channel by channel to obtain an explicit residual deformation field. The instantaneous deformation field is obtained by superimposing the implicit global deformation field and the explicit residual deformation field. The encoding of the time scalar is as follows: periodic phase encoding is used, and a sinusoidal position encoding mechanism is employed to normalize the time scalar. The mapping is to a high-dimensional feature vector, and the encoding formula is: ; For frame spacing The encoding is: ; in, For a respiratory half-cycle, L is the maximum set number of frequency bands, and k is the frequency index; The registration module is configured to obtain the cumulative deformation field from the initial time to the current time by cascading the instantaneous deformation fields of all intermediate times between the initial time and the current time, thereby completing the image registration.

6. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-4.

8. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-4.

Citation Information

Patent Citations

  • Multi-scale implicit-explicit hybrid medical image registration method and system

    CN120876558A

  • Self-adaptive medical image sequence registration method and system based on space-time cooperative coding

    CN121482428A