Medical image registration method and system based on image-to-image translation

By combining image-to-image translation and pose initialization modules, the problems of imaging modality differences and inaccurate initial poses are solved, and efficient and high-precision 2D/3D medical image registration is achieved, which is suitable for real-time surgical navigation.

CN120765709AActive Publication Date: 2025-10-10SHANDONG UNIV

Patent Information

Application Number
CN202511276960.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-10
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Existing 2D/3D medical image registration methods face challenges in imaging modality differences, inaccurate initial pose estimation, and optimization efficiency and accuracy, making it difficult to meet the requirements of real-time and high precision.

Method used

An image-to-image translation module is used to translate linear X-ray images into simulated DRR images. The pose initialization module is combined with self-supervised learning to obtain the initial pose, and the pose is optimized through a multi-scale optimization strategy. The two modules are integrated into a unified framework to reduce the random search time of traditional methods.

Benefits of technology

Real-time registration is achieved with sub-millimeter accuracy, which improves the robustness and accuracy of registration, especially in low-contrast areas and complex anatomical structures, with the registration time controlled within 3.2 seconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765709A_ABST
    Figure CN120765709A_ABST
Patent Text Reader

Abstract

The invention provides a medical image registration method and system based on image-to-image translation, and belongs to the technical field of medical image registration, and the method comprises the steps: obtaining a linear X-ray image, processing the linear X-ray image, and inputting the processed linear X-ray image to an image-to-image translation module to obtain a simulation DRR image; inputting the obtained linear X-ray image into a pose initialization module to obtain an initialized pose; setting the obtained initialized pose as a global variable, inputting the global variable into an optimizer, down-sampling a simulation DRR image and a standard DRR image to a certain scale, calculating the image similarity between the images, inputting the image similarity into the optimizer, continuously updating the pose output by the optimizer, and carrying out optimization of different scales to obtain a final pose. And calculating a target registration error based on the final pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image registration, and in particular relates to a medical image registration method and system based on image-to-image translation. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] 2D / 3D medical image registration is a key technology in image-guided surgery (IGS). Its core goal is to align preoperative high-resolution 3D computed tomography (CT) images with intraoperative 2D X-ray images in real time, thereby providing surgeons with precise surgical navigation.

[0004] However, existing registration methods still face many challenges due to differences in imaging modalities, initial pose estimation errors, and the strict clinical requirements for real-time performance and high accuracy.

[0005] For example, differences in imaging modalities can make registration difficult. Traditional registration methods, such as those based on normalized cross correlation (NCC) or mutual information (MI), rely on intensity similarity between X-ray images and digitally reconstructed radiographs (DRR images) generated by CT. However, X-ray images are affected by scattering, noise, and low contrast, while DRR images are simulated images generated based on idealized physical models. The two differ significantly in grayscale distribution, noise patterns, and structural details. In practical training, X-ray images are typically preprocessed. First, a logarithmic transformation (logarithm of the maximum pixel value minus the logarithm of each pixel value) is applied to transform the attenuation of the X-ray image from nonlinear to linear. 50 pixels are cropped from all sides to eliminate the effects of the collimator and ensure compatibility with the DRR image projector. In the following steps, these preprocessed X-ray images are collectively referred to as linear X-ray images. Furthermore, the original DRR images are collectively referred to as standard DRR images, the DRR images obtained through the image-to-image translation network are collectively referred to as simulated DRR images, and the images obtained through data augmentation of the standard DRR images are collectively referred to as simulated X-ray images.

[0006] Unsupervised image-to-image translation methods, such as CycleGAN, do not rely on paired data and may introduce structural distortion or artifacts during translation, destroying key anatomical features such as bone edges, leading to subsequent registration failure.

[0007] Supervised image-to-image translation methods, such as pix2pix, rely on paired data. However, high-quality paired datasets of medical X-ray and DRR images are scarce, and manual annotation is extremely expensive. Differences in imaging modalities make intensity-based registration algorithms perform poorly in low-contrast areas such as soft tissue or complex anatomical structures like the pelvis and spine, increasing registration errors.

[0008] There's also the issue of inaccurate initial pose estimation, which can hinder optimization convergence. For example, marker-based methods, such as PnP, rely on manually annotated 2D-3D marker matching. However, intraoperative X-ray images may contain occlusions, metal artifacts, or missing markers, leading to matching failures. Marker positioning errors are directly propagated to the pose estimate, affecting registration accuracy.

[0009] Methods based on random sampling points, such as SCRNet and RayEmb, establish 2D-3D correspondences through scene coordinate regression or learning ray embedding space, avoiding dependency on landmarks. However, these methods are computationally complex, require a large number of sampling points, and are sensitive to image quality. Pose regression methods, such as PoseNet and DiffPose, use convolutional networks to directly regress 6D poses from images. However, due to differences in imaging modalities, models trained on DRR images, for example, struggle to adapt to X-ray images, resulting in large initial pose errors. Initial pose deviations can cause gradient optimization to fall into local optima, significantly increasing the registration failure rate in complex anatomical structures or with large pose offsets. While XVR uses data augmentation during pose regression to improve the generalization of initial pose estimates for X-ray images, the segmentation and projection of CT images during pose optimization increases the modality difference between DRR and X-ray images, leaving room for further improvement in registration accuracy.

[0010] Furthermore, optimization efficiency is low, making it difficult to meet real-time requirements. Traditional optimization methods, such as CMA-ES, employ random search strategies. While this can expand the capture range, it converges slowly, with a single registration taking up to several minutes, failing to meet real-time intraoperative requirements. Gradient optimization methods, such as DiffDRR, accelerate gradient computation through differentiable rendering, but are still limited by local optima and are sensitive to initial pose. Specifically, the Trilinear projection method used by DiffDRR is a DRR image projection method based on trilinear interpolation. This method simulates the cumulative attenuation effect of X-rays traversing the CT volume by uniformly sampling points in 3D space along a ray and interpolating the values ​​of adjacent voxels. This method discretizes the ray into multiple sampling points, performs differentiable trilinear interpolation using the PyTorch function grid_sample, and finally integrates (sums or maximizes) the values ​​along the ray to generate the DRR image. Compared to exact ray tracing, such as the Siddon method, this method is more computationally efficient and supports gradient propagation, making it suitable for joint optimization with deep learning models. However, its accuracy is slightly lower than that of physically based ray tracing methods.

[0011] There's also the issue of low registration accuracy. Single-scale registration, relying solely on a single resolution or fixed receptive field, struggles to simultaneously capture the consistency of global spatial structure (such as large-scale displacements or rotations) and the precise alignment of subtle local deformations (such as tissue edges or texture details). This limitation results in insufficient global consistency or local accuracy in the registration results, impacting overall registration accuracy.

[0012] Therefore, the current 2D / 3D medical image registration has problems such as inaccurate registration due to differences in imaging modalities, inaccurate initial image pose estimation due to labeling, and low optimization efficiency and registration accuracy during pose optimization. Summary of the Invention

[0013] To overcome the above-mentioned deficiencies of the prior art, the present invention provides a medical image registration method and system based on image-to-image translation, which integrates image-to-image translation, pose initialization and optimization into a unified framework, reduces the time-consuming random search problem of traditional methods, and realizes real-time and high-precision registration.

[0014] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: In a first aspect, a medical image registration method based on image-to-image translation is disclosed, comprising: Acquire a linear X-ray image, process the linear X-ray image, and input it into an image-to-image translation module to obtain a simulated DRR image; Input the linear X-ray image into the pose initialization module to obtain the initial pose; The obtained initialization pose is set as a global variable and input into the optimizer. After the translated simulated DRR image and the projected standard DRR image are downsampled to a certain scale, the image similarity between the images is calculated and input into the optimizer. The pose output by the optimizer is continuously updated. After optimization at different scales, the final pose is obtained, and the target registration error is calculated based on the final pose.

[0015] As a further technical solution, the image-to-image translation module adopts a generative adversarial network, including a generator and a discriminator; During the training phase, the input of the generator is the masked linear X-ray image, and the input of the discriminator is the standard DRR image obtained by segmenting and projecting the CT volume in the real pose; The generator translates the input linear X-ray image into a simulated DRR image, and the discriminator judges the authenticity of the simulated DRR image translated by the generator based on a standard DRR image.

[0016] As a further technical solution, the image-to-image translation module adopts multi-scale discriminator joint optimization to ensure that the translated simulated DRR image retains key anatomical structure features through adversarial loss, feature matching loss and VGG perceptual loss.

[0017] As a further technical solution, during the training phase, the posture initialization module samples a predefined posture range to obtain the true posture, and then projects the true posture to obtain a standard DRR image. The standard DRR image is further subjected to data enhancement operations to obtain a large number of simulated X-ray images. The simulated X-ray image is then input into the posture regressor to obtain the predicted posture, and then the predicted posture is projected to obtain the standard DRR image.

[0018] As a further technical solution, the pose initialization module jointly optimizes the multi-scale normalized cross-correlation loss and the manifold-based geometric constraint loss through self-supervised learning at the end of the training phase so that the pose regressor outputs the predicted initialization pose.

[0019] In a second aspect, a medical image registration system based on image-to-image translation is disclosed, comprising: The image-to-image translation module is configured to: receive an input linear X-ray image, and process the linear X-ray image to obtain a simulated DRR image; The pose initialization module is configured to: process the acquired linear X-ray image to obtain an initial pose; The pose optimization module is configured to: set the obtained initial pose as a global variable and input into an optimizer, calculate the image similarity between the simulated DRR image and the standard DRR image after down-sampling to a certain scale, and input into the optimizer, constantly update the pose output by the optimizer, obtain the final pose after optimization at different scales, and calculate the target registration error based on the final pose.

[0020] The above one or more technical solutions have the following beneficial effects: In view of the problem that imaging modal difference leads to registration difficulty: the image-to-image translation module is designed in the technical solution of the embodiment, the neural network used by the module is an improved high-resolution image-to-image translation network, the core of which is composed of a generator and a discriminator, which is specially used to translate linear X-ray images into simulated DRR images. The global generator improves the detail fidelity, and multiple sub-discriminators judge the image authenticity at different resolutions, forcing the generator to retain the spatial consistency of the anatomical structure, which helps to achieve accurate registration later.

[0021] The technical solution of the embodiment compensates for the imaging modal difference between the linear X-ray image and the standard DRR image through image-to-image translation, thereby improving the robustness of the registration algorithm based on intensity similarity, and achieving more accurate registration in low-contrast areas (soft tissue) and complex anatomical structures (pelvis, spine).

[0022] In view of the problem that inaccurate initial pose estimation affects optimization convergence: the pose initialization module is designed in the technical solution of the embodiment, the pose regressor used by the module is obtained by jointly optimizing the image multi-scale normalized cross-correlation loss and the manifold-based geometric constraint loss through self-supervised learning to obtain the initial pose, without relying on artificial marker points, reducing the matching failure caused by metal artifacts or occlusion.

[0023] In view of the problem that the optimization efficiency and registration accuracy are not high during the pose optimization process: the pose optimization module is designed in the technical solution of the embodiment, the multi-scale optimization strategy used by the module realizes fast global pose estimation at a coarse scale, and then accurately optimizes high-frequency details such as bone microstructure at a fine scale, thereby significantly improving the registration efficiency while ensuring higher registration accuracy.

[0024] The technical solution of the embodiment integrates the image-to-image translation, pose initialization and optimization modules into a unified framework, reduces the time-consuming of random search of traditional methods such as CMA-ES, and realizes real-time and high-precision registration.

[0025] The advantages of the additional aspects of the application will be partially given in the following description, partially become obvious from the following description, or be known by the practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0026] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0027] Figure 1 CT visualization structure diagram and label under 3D Slicer; Figure 2 This is a network structure diagram of PixPose according to an embodiment of the present invention; Figure 3 This is an example of a simulated X-ray image after data augmentation for the PixPose network structure pose initialization module. Figure 4 This is an example diagram of each module of the PixPose network structure; Figure 5 Schematic diagram of TRE visualization of different methods on the DeepFluoro dataset; Figure 6 A diagram showing the visualization of the registration time of different methods on the DeepFluoro dataset; Figure 7 Schematic diagram for visualizing the true and predicted landmarks of different methods on the DeepFluoro dataset. DETAILED DESCRIPTION

[0028] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0029] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.

[0030] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.

[0031] Example 1 The key technical improvement of this embodiment is to build a 2D / 3D medical image registration framework PixPose. Specifically, a medical image registration method based on image-to-image translation is disclosed, including: Step 1: Obtain a linear X-ray image, process the linear X-ray image, and input it into the image-to-image translation module to obtain a simulated DRR image; Step 2: Input the acquired linear X-ray image into the pose initialization module to obtain the initial pose; Step 3: Set the obtained initialization pose as a global variable and input it into the optimizer. After downsampling the simulated DRR image and the standard DRR image to a certain scale, calculate the image similarity between the images and input it into the optimizer. Continuously update the pose output by the optimizer. After optimization at different scales, obtain the final pose and calculate the target registration error based on the final pose.

[0032] In one implementation example, regarding the imaging modality difference problem in step 1, an image-to-image translation module based on the pix2pixHD network is designed. The network used in this module is an improved high-resolution image-to-image translation network, whose core consists of a generator and a discriminator, and is specifically used to translate linear X-ray images into simulated DRR images. The generator adopts a multi-scale global-local structure. The global generator extracts overall features through a U-Net structure of downsampling-residual block-upsampling, while the local enhancer hierarchically refines the low-resolution features to improve the fidelity of details. The discriminator adopts a multi-scale PatchGAN architecture, which uses multiple sub-discriminators to judge the authenticity of images at different resolutions, forcing the generator to preserve the spatial consistency of the anatomical structure.

[0033] In the training phase, see the black solid line in module I. Figure 2 As shown in Figure 2, the input of this module is a linear X-ray image and a standard DRR image of the same pose. In order to adapt to the input of the pix2pixHD network, the image size is downsampled from 1436 x 1436 to 512 x 512.

[0034] In order to make full use of the characteristics of the DeepFluoro dataset, this module further masks the linear X-ray image and segments and projects the CT volume, reducing the influence of the anatomical structure around the pelvis on the image-to-image translation. Since the labels of the linear X-ray image mask are 1 to 6, the labels retained by the segmentation (Segment, S) are also set to 1 to 6, as shown in Figure 1 shown.

[0035] The generator (G) translates the input linear X-ray image into a simulated DRR image G fake To deceive the discriminator (Discriminator, D), the discriminator judges G based on the standard DRR image generated by the segmented CT volume (Volume, V) and the real pose transformation (Transformation, T) projection fake Is it true or false.

[0036] In terms of specific implementation details, this module adopts multi-scale discriminator joint optimization, and ensures the final generation of simulated DRR images with both high resolution and structural accuracy through the adversarial loss corresponding to the following formula (1), the feature matching loss corresponding to the formula (2), and the VGG perceptual loss corresponding to the formula (3), providing high-quality input for subsequent alignment. In the testing phase, see the red solid line in module I. Figure 2 As shown in the figure, the simulated DRR image obtained by directly inputting the masked and translated linear X-ray image can be used. This design reduces the imaging modality difference between the linear X-ray image and the standard DRR image during the pose optimization process, significantly improving the registration accuracy.

[0037] About adversarial loss: ; in, is a linear X-ray image, is a standard DRR image, is a simulated DRR image, represents the last layer of different discriminators, represents the probability that the discriminator judges the standard DRR image to be true, represents the probability that the discriminator judges the simulated DRR image to be true, represents the number of discriminators, represents the weights of different discriminators, Represents the average loss of a set of images. This loss function ensures that the generated simulated DRR images are visually indistinguishable from standard DRR images through adversarial training of the generator and the discriminator.

[0038] About feature matching loss: ; in, is a linear X-ray image, is a standard DRR image, is a simulated DRR image, Represents the first layer, Indicates the first time the discriminator judges the standard DRR image as true The feature map of the layer, Indicates the first time that the discriminator judges the simulated DRR image as true The feature map of the layer, represents the weights of different layers of the discriminator, represents the total number of layers of each discriminator, represents the number of discriminators, represents the weights of different discriminators, Represents the average loss of a set of images. This loss function constrains the generator to preserve the structural features of the image and avoids deviation in deep features between the simulated DRR image and the standard DRR image by matching the feature maps (such as edges and textures) of the intermediate layers of the discriminator.

[0039] About VGG perceptual loss: ; in, is a linear X-ray image, is a standard DRR image, is a simulated DRR image, Represents the first layer, Indicates the simulated X-ray image in the VGG network The feature map of the layer, Indicates that the standard DRR image is in the VGG network The feature map of the layer, Represents the different weights of each layer of the VGG network, Represents the total number of layers of the VGG network, The loss function uses the pre-trained VGG network to extract high-level semantic features and further aligns the global semantic content of the simulated DRR images with the standard DRR images, such as organ shape and spatial layout.

[0040] In one implementation example, a pose initialization module based on a pose regressor is designed to address the problem of inaccurate initial pose estimation in step 2. The ResNet-18 used in the pose regressor is a classic deep convolutional neural network architecture, which is mainly used to solve the gradient vanishing problem in deep network training. The network contains 18 weighted layers. Its core innovation is the introduction of a residual connection (skip connection) mechanism, which allows the gradient to be directly back-propagated to the shallow layer, thereby effectively training deeper networks. ResNet-18 serves as the feature extraction backbone network. The features extracted from its last layer are decoded into rotation (3 degrees of freedom) and translation (3 degrees of freedom) parameters by two linear layers to achieve end-to-end pose regression.

[0041] During the training phase, Figure 2 As shown in the black solid line of module II, this module first samples the three translation components and three rotation components in the predefined pose range to obtain the true pose , and then the DRR image is obtained by projection , and then obtain a large number of simulated X-ray images through data augmentation (DA) ,like Figure 3As shown, the pose regressor (P) is then input to obtain the predicted pose , and then the DRR image is obtained by projection Finally, the multi-scale normalized cross-correlation loss corresponding to formula (4) and the geometric constraint loss based on SE(3) manifold corresponding to formula (5) are jointly optimized through self-supervised learning.

[0042] During the testing phase, Figure 2 As shown in the red solid line in Module II, the linear X-ray image can be directly input to obtain the initial pose. This design avoids the reliance on traditional marker methods, generates synthetic training data through differentiable rendering, and performs random data augmentation operations on the synthetic training data to obtain simulated X-ray images, improving the generalization and robustness of the initial pose estimation.

[0043] About multi-scale normalized cross-correlation loss: ; in, and Respectively represent the simulated DRR image block and the standard DRR image block after normalization at a specific scale. The pixel value at the pixel coordinate, Indicates the size of the image block, Represents weights of different scales, Represents the average loss of a set of images. This loss function measures the local similarity between simulated DRR images and standard DRR images at different scales and optimizes the pose regressor through normalized cross-correlation.

[0044] Geometric constraint loss based on SE(3) manifold: ; in, and The true poses are The rotational and translational components of and They are respectively predicted pose The rotational and translational components of represents the focal length of the projection, This loss function directly constrains the distance between the predicted pose and the true pose on the SE(3) manifold, ensuring the physical rationality of the rotation and translation components.

[0045] In one implementation example, a pose optimization module based on a multi-scale strategy was designed to address the issues of low optimization efficiency and registration accuracy during pose optimization in step three. The core of this strategy is to balance the efficiency and accuracy of the registration process by gradually reducing the image resolution: rapid and rough alignment is performed at the high-resolution level (small scale) to capture large-scale transformations; while fine adjustments are performed at the low-resolution level (large scale) to optimize local details. This strategy effectively expands the optimization convergence domain, avoids local optimal solutions, and significantly improves computational efficiency. In the specific implementation, each scale level dynamically adjusts the learning rate and adopts an early stopping mechanism to ensure optimal registration results with limited computing resources.

[0046] This module is used directly in the testing phase, such as Figure 2 As shown in the red solid line of module III, the linear X-ray image is input into the pose regressor trained in step 2 to obtain the initial pose , then Set as a global variable and input it into the Adam optimizer. The linear X-ray image after the mask in step 1 is translated through the trained generator to obtain the simulated DRR image, and step 3 is based on After the DRR images generated by the segmented CT volume projection are downsampled to a certain scale, the pose is continuously updated according to the image similarity (Sim) between the corresponding images according to formula (7). , after optimization at different scales, the final pose is obtained , thereby calculating the target registration error (TRE) corresponding to formula (8).

[0047] In terms of specific implementation details, since the preoperative CT and intraoperative X-ray images are acquired independently, the patient's femur may move slightly during this process. Therefore, the femoral segments labeled 5 and 6 are excluded when using the segmented projection.

[0048] Introduce gradient normalized cross-correlation similarity calculation in the pose optimization module: ; in, and They represent the normalized simulated DRR image and the standard DRR image processed by the gradient operator. The pixel value at the pixel coordinate, and Indicates the size of the image, Represents the average similarity of a group of images. This loss function uses the gradient operator to strengthen image edge alignment and improve the accuracy of pose optimization.

[0049] About the similarity calculation between images: ; in, and Represent the simulated DRR images and standard DRR images of different scales respectively, Indicates the size of the image block, Represents weights of different scales, This similarity calculation combines multi-scale NCC and gradient NCC to construct a pyramid optimization target, achieving coarse-to-fine pose adjustment.

[0050] About target registration error calculation: ; in, Represents the 3D coordinates of the marker point, and represent the true pose and the predicted pose respectively, Indicates the number of markers. This evaluation metric calculates the Euclidean distance between the markers in the transformation between the true pose and the predicted pose.

[0051] In the technical solution of this embodiment, the entire process integrates three sub-modules: image-to-image translation (Module I), pose initialization (Module II), and optimization (Module III). Ultimately, the registration time is controlled to 3.2 seconds while maintaining submillimeter accuracy (5% quantile error <1mm).

[0052] The technical solution of this example embodiment uses the pix2pixHD network for image-to-image translation, combined with a multi-scale discriminator, adversarial loss, feature matching loss, and VGG perceptual loss to ensure that the generated DRR image-like images retain key anatomical structures such as bone edges, avoiding the structural distortion problem of unsupervised image-to-image translation methods such as CycleGAN.

[0053] The technical solution of this embodiment introduces differentiable rendering technology (such as DiffDRR) to generate standard DRR images. Through data augmentation, the generalization ability of the pose regressor for estimating the pose of linear X-ray images is improved, avoiding the initial pose error caused by imaging modality differences in traditional methods.

[0054] To address the problems of low optimization efficiency and registration accuracy during pose optimization, the technical solution of this embodiment proposes a multi-scale optimization strategy, which can achieve rapid global registration at the coarse scale and accurately optimize high-frequency details such as bone microstructure at the fine scale, thereby improving registration efficiency and accuracy.

[0055] In one implementation example, hardware acceleration (such as GPU parallel computing) and lightweight network design are used to ensure efficient deployment in clinical environments (such as surgical navigation systems) and avoid process interruptions caused by manual intervention.

[0056] In one implementation example, a method for constructing a high-quality dataset of paired linear X-ray images and standard DRR images is proposed. The existing DeepFluoro dataset consists of six pelvic CT scans and their corresponding segmentations. Each CT scan contains several original X-ray images, corresponding mask images, and ground-truth poses. Based on this dataset, standard DRR images are generated using the ground-truth poses and CT volume projections corresponding to each linear X-ray image. Due to the scarcity of linear X-ray images, all linear X-ray images and standard DRR images for each CT scan form a separate test set, while all linear X-ray images and standard DRR images from all remaining CT scans form a training set.

[0057] The comparison of the PixPose framework proposed in this embodiment's sub-technical solution with the existing technology is as follows: Figure 4 The visual comparison of the dataset samples in the preprocessing, image-to-image translation and registration stages is shown. The images registered using the PixPose method are almost completely aligned with the linear X-ray images. Figure 5 Table 1 and Table 2 show the 2D / 3D registration performance comparison of various methods on the DeepFluoro dataset, covering a variety of accuracy indicators (including the 5th, 50th, and 95th percentile errors, success rates at different thresholds) and registration time. PixPose shows obvious advantages in many key indicators, especially in tasks with high precision requirements. For example, on all sub-datasets, PixPose's 5th percentile error and failure rate at a threshold of 1mm are significantly lower than those of other methods, showing excellent robustness and accuracy. In addition, this method also achieves or is close to the best results in indicators such as the 50th and 95th percentile errors and failure rates at thresholds of 10mm and 5mm, while maintaining a low running time, such as Figure 6 As shown in Figure 2, the time taken is typically about 3.2 seconds, demonstrating high efficiency and real-time performance in practical deployments. These results highlight the significant advantages of image translation technology in resolving the imaging modality differences between linear X-ray images and standard DRR images. Figure 7 The visual comparison of the predicted landmark points and the true landmark points after registration by each method is shown. It can be clearly seen that the predicted landmark points obtained by PixPose are closest to the true landmark points.

[0058] Table 1

[0059] Example 2 The purpose of this embodiment is to provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.

[0060] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0061] A computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the above method.

[0062] Example 4 The purpose of this embodiment is to provide a medical image registration system based on image-to-image translation, including: The image-to-image translation module is configured to: receive an input linear X-ray image, and process the linear X-ray image to obtain a simulated DRR image; The pose initialization module is configured to: process the acquired linear X-ray image to obtain an initial pose; The pose optimization module is configured as follows: setting the obtained initialization pose as a global variable and inputting it into the optimizer; downsampling the simulated DRR image and the standard DRR image to a certain scale; calculating the image similarity between the images; and inputting it into the optimizer; continuously updating the pose output by the optimizer; after optimization at different scales, obtaining the final pose; and calculating the target registration error based on the final pose.

[0063] Example 5 The purpose of this embodiment is to provide a computer program product containing instructions, which, when running on a computer, enables the computer to execute the methods and functions involved in any of the above embodiments. The steps involved in the apparatus of the above embodiment correspond to those of the method embodiment 1. For detailed implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and causing the processor to perform any method of the present invention.

[0064] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0065] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.

Claims

1. A medical image registration method based on image-to-image translation, characterized by: include: Acquire a linear X-ray image, process the linear X-ray image, and input it into an image-to-image translation module to obtain a simulated DRR image; Input the acquired linear X-ray image into the pose initialization module to obtain the initial pose; The obtained initialization pose is set as a global variable and input into the optimizer. After downsampling the simulated DRR image and the standard DRR image to a certain scale, the image similarity between the images is calculated and input into the optimizer. The pose output by the optimizer is continuously updated. After optimization at different scales, the final pose is obtained, and the target registration error is calculated based on the final pose.

2. The medical image registration method based on image-to-image translation according to claim 1, wherein: The image-to-image translation module uses a generative adversarial network, including a generator and a discriminator; During the training phase, the input of the generative adversarial network is the standard DRR image obtained by segmenting and projecting the CT volume under the corresponding real pose; The generator translates the input X-ray image into a simulated DRR image, and the discriminator judges the authenticity of the simulated DRR image translated by the generator based on a standard DRR image.

3. The medical image registration method based on image-to-image translation according to claim 1, wherein: The image-to-image translation module adopts multi-scale discriminator joint optimization to ensure that the generated simulated DRR images retain key anatomical structure features through adversarial loss, feature matching loss and VGG perceptual loss.

4. The medical image registration method based on image-to-image translation according to claim 1, wherein: During the training phase, the pose initialization module samples the predefined pose range to obtain the real pose, then projects the real pose to obtain a standard DRR image, performs data augmentation on the standard DRR image to obtain a simulated X-ray image, then inputs the simulated X-ray image into the pose regressor to obtain the predicted pose, and then projects the predicted pose to obtain the standard DRR image.

5. The medical image registration method based on image-to-image translation according to claim 1, wherein: At the end of the training phase, the pose initialization module jointly optimizes the multi-scale normalized cross-correlation loss and the manifold-based geometric constraint loss through self-supervised learning so that the pose regressor outputs the predicted initialization pose.

6. A medical image registration system based on image-to-image translation, characterized by: include: The image-to-image translation module is configured to: receive an input linear X-ray image, and process the linear X-ray image to obtain a simulated DRR image; The pose initialization module is configured to: process the acquired linear X-ray image to obtain an initial pose; The pose optimization module is configured as follows: setting the obtained initialization pose as a global variable and inputting it into the optimizer; downsampling the simulated DRR image and the standard DRR image to a certain scale; calculating the image similarity between the images and inputting it into the optimizer; continuously updating the pose output by the optimizer; after optimization at different scales, obtaining the final pose; and calculating the target registration error based on the final pose.

7. The medical image registration system based on image-to-image translation according to claim 6, characterized in that: The image-to-image translation module uses a generative adversarial network, including a generator and a discriminator; During the training phase, the input of the generative adversarial network is the standard DRR image obtained by segmenting and projecting the CT volume under the corresponding real pose; The generator translates the input X-ray image into a simulated DRR image, and the discriminator judges the authenticity of the simulated DRR image translated by the generator based on a standard DRR image.

8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method described in any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are performed.

Citation Information

Patent Citations

  • Multi-modal medical image translation method based on target perception generative adversarial network

    CN114373532A

  • X-ray image and CT image registration method

    CN114511597A

  • Intraoperative X-ray and CT image registration method and device

    CN118037793A

  • 2D-3D medical image registration method and device, computer equipment and storage medium

    CN118608583A

  • Generating Synthetic X-ray Images and Object Annotations from CT Scans for Augmenting X-ray Abnormality Assessment Systems

    US20220292742A1

Cited By

  • Three-dimensional ultrasound reconstruction method and device, electronic equipment and storage medium

    CN122435169A