A registration method and system based on CT feature projection and multi-view X-ray
Patent Information
- Application Number
- CN202610710732.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-05-22
AI Technical Summary
基于特征点匹配的方法需要人工提取图像中的特征点(如骨骼边缘、病灶轮廓),操作繁琐、主观性强,且在特征点模糊或缺失时配准精度大幅下降;基于图像灰度相似度的方法通常通过生成DRR(数字重建放射影像)图像,将CT图像与X光图像的配准转化为DRR图像与X光图像的相似度匹配,但传统DRR投影多为不可微分操作,无法融入端到端神经网络训练,导致配准效率低、收敛速度慢
[0017]The beneficial effects of this invention are as follows: The registration method based on CT feature projection and multi-view X-ray is applicable to the registration of CT and multi-view X-ray images of different target objects (such as human bones and organs). Compared with the prediction model trained using only a single CT image, it effectively improves the generalization ability of the prediction model, can adapt to different clinical application scenarios, and has broad promotion value and practical significance. The spatial constraint effect of multi-view X-ray images is used to make up for the shortcomings of single-view registration, further ensuring the accuracy of iterative registration. Differentiable projection is used to integrate the DRR generation process into the end-to-end neural network training, and the multiple constraints of geodesic loss and cross-geodesic loss are combined to significantly improve the accuracy of initial pose prediction. The initial registration pose is quickly predicted by a deep learning model, and the backpropagation characteristics of differentiable projection are used to accelerate iterative optimization, greatly shortening the registration time and achieving fully automatic and efficient registration. The style alignment of X-ray images and DRR images is completed by iteratively traversing the maximum recovery value method, eliminating modal difference interference, and the multi-view constraints effectively resist the influence of image occlusion and noise, ensuring the stability and robustness of registration in complex scenes.
Smart Images

Figure CN122265358B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to a registration method and system based on CT feature projection and multi-view X-rays. Background Technology
[0002] In the field of medical imaging, CT (computed tomography) images provide three-dimensional structural information of target objects (such as human bones and organs), offering rich detail and high spatial resolution. X-ray images, on the other hand, offer advantages such as fast imaging speed, low radiation dose, and portable equipment, making them widely used in real-time intraoperative guidance and rapid diagnosis. Image registration is a crucial technology for mapping two different modalities of medical images to the same spatial coordinate system. Its registration accuracy directly affects the accuracy of clinical diagnosis and the safety of surgical navigation.
[0003] Currently, CT and X-ray image registration methods are mainly divided into two categories: feature point matching-based and image grayscale similarity-based methods. Feature point matching-based methods require manual extraction of feature points (such as bone edges and lesion contours) from the image, which is cumbersome, subjective, and suffers from a significant drop in registration accuracy when feature points are blurred or missing. Image grayscale similarity-based methods typically generate DRR (Digital Reconstructed Radiography) images, transforming the registration of CT and X-ray images into similarity matching between DRR and X-ray images. However, traditional DRR projections are mostly non-differentiable operations, making it impossible to incorporate end-to-end neural network training, resulting in low registration efficiency and slow convergence. Furthermore, existing registration methods often use single-view X-ray images, making them susceptible to image occlusion and noise interference, hindering accurate registration in complex scenes.
[0004] Therefore, it is necessary to provide a new registration method based on CT feature projection and multi-view X-ray. Summary of the Invention
[0005] Based on the aforementioned problems in the existing technology, the purpose of this invention is to provide a registration method based on CT feature projection and multi-view X-ray. By using differentiable projection and multi-view constraints, it achieves fully automatic and high-precision registration of three-dimensional CT and multi-view X-ray images, thereby improving registration accuracy and robustness.
[0006] To achieve the above objectives, the technical solution adopted by this invention is: a registration method based on CT feature projection and multi-view X-rays, comprising:
[0007] S1, construct a CT training dataset, obtain a set of 3D CT images from the dataset and preprocess them to obtain the first CT image; S2, Construct an image registration model, which includes a first feature encoder, a second feature encoder, a differentiable projection, feature fusion, and pose prediction; S3, different random perturbations are added to the standard pose and differentiable projections are performed to obtain the first DRR image and the second DRR image; S4, input the first CT image into the first feature encoder to obtain the first feature, process the first feature under the standard pose through differentiable projection to obtain the second feature; input the preprocessed first DRR image into the second feature encoder to obtain the third feature; input the second feature and the third feature into feature fusion and pose prediction to obtain the predicted perturbation pose, combine with the standard pose to obtain the first predicted pose, and then obtain the first predicted DRR image through differentiable projection. Repeat this step to obtain the second predicted pose and the second predicted DRR image. S5, calculate the similarity loss, geodesic loss and cross-geodesic loss respectively, and train the image registration model by combining all the losses; S6: Obtain the 3D CT image and multi-view X-ray image of the target object and input them into the trained image registration model to obtain the initial registration pose. Iteratively optimize based on the multi-view constraint loss to obtain the accurate registration pose.
[0008] Furthermore, the three-dimensional preprocessing of a set of three-dimensional CT images in the dataset includes resampling and cropping of the CT images, truncation and normalization of CT voxels.
[0009] Furthermore, the first feature encoder consists of a three-dimensional residual convolutional network for extracting three-dimensional features of CT images; the second feature encoder consists of a two-dimensional residual convolutional network for extracting two-dimensional features of DRR images; the differentiable projection is implemented based on ray tracing algorithms and deep learning frameworks to complete DRR projection calculations; the feature fusion consists of channel-dimensional concatenation and convolutional networks to effectively fuse the projected two-dimensional CT features with the two-dimensional DRR features, enhancing the feature representation capability; and the pose prediction consists of a multilayer perceptron to output perturbation parameters of the projected pose, thereby predicting the projected pose.
[0010] Furthermore, the step of adding different random perturbations to the standard pose and performing differentiable projection to obtain the first DRR image and the second DRR image includes: adding different random perturbations to the standard pose to generate the first projection pose and the second projection pose, calculating the first transformation matrix of the first projection pose and the second projection pose, and projecting the first CT image according to the first projection pose and the second projection pose respectively through differentiable projection to obtain the first DRR image and the second DRR image. The calculation of the first transformation matrix between the first projected pose and the second projected pose includes using a rigid body transformation algorithm to calculate the first transformation matrix between the first projected pose and the second projected pose. The first transformation matrix includes a rotation matrix and a translation matrix. The addition of random perturbation to the standard pose to generate the first projected pose includes setting the standard pose as the orthogonal pose of the target skeleton, adding random perturbation to the standard pose, and generating the first projected pose.
[0011] Furthermore, step S4 includes: The first feature encoder of the image registration model inputs the first CT image and uses three-dimensional convolution, normalization and activation functions to extract the three-dimensional features of the first CT image to obtain the first feature; The first feature is subjected to a differentiable projection according to the standard pose, and the three-dimensional feature is projected into a two-dimensional feature. Then, the normalization process is performed to obtain the second feature. The first DRR image is preprocessed in two dimensions. The preprocessed first DRR image is then input into the second feature encoder of the model. Two-dimensional convolution, normalization and activation functions are used to extract the two-dimensional features of the first DRR image to obtain the third feature. Based on the predicted perturbation pose and the standard pose, the first predicted pose is calculated, and then the first CT image is projected onto the first predicted pose using differentiable projection to obtain the first predicted DRR image.
[0012] Furthermore, step S4 also includes: The first CT image is re-input into the first feature encoder to obtain the first feature, which is then projected onto the standard pose using a differentiable projection to obtain the second feature; the second DRR image is preprocessed in two dimensions and then input into the second feature encoder to extract the two-dimensional features of the second DRR image to obtain the fourth feature. The second and fourth features are input into the feature fusion module. The two features are concatenated along the channel dimension to obtain the fused feature. The fused feature is then input into the pose prediction module to output the second predicted perturbation pose. Based on the second predicted perturbation pose and the standard pose, the second predicted pose is calculated, and then the first CT image is projected onto the second predicted pose using differentiable projection to obtain the second predicted DRR image.
[0013] Furthermore, step S5 includes: Calculate the geodesic loss between the first projected attitude and the first predicted attitude, and between the second projected attitude and the second predicted attitude, to constrain the accuracy of attitude prediction; The cross geodesic loss includes the geodesic loss calculated from the first projected attitude and the attitude obtained from the second predicted attitude and the first transformation matrix, and the geodesic loss calculated from the second projected attitude and the attitude obtained from the first predicted attitude and the first transformation matrix.
[0014] A registration system based on CT feature projection and multi-view X-rays, applied to the aforementioned registration method based on CT feature projection and multi-view X-rays, the system comprising: The data preprocessing module is used to construct a CT training dataset, obtain a set of 3D CT images from the dataset, and preprocess them to obtain the first CT image. The modeling module is used to construct an image registration model, which includes a first feature encoder, a second feature encoder, a differentiable projection, feature fusion, and pose prediction. The training sample generation module is used to add different random perturbations to the standard pose and perform differentiable projection to obtain the first DRR image and the second DRR image. The attitude prediction module is used to input the first CT image into the first feature encoder to obtain the first feature, and process the first feature under the standard attitude through differentiable projection to obtain the second feature; input the preprocessed first DRR image into the second feature encoder to obtain the third feature; input the second feature and the third feature into feature fusion and attitude prediction to obtain the predicted perturbation attitude, combine it with the standard attitude to obtain the first predicted attitude, and then obtain the first predicted DRR image through differentiable projection. This step is repeated to obtain the second predicted attitude and the second predicted DRR image. The loss calculation training module is used to train the image registration model by combining the similarity loss, geodesic loss, and cross-geodesic loss separately. The registration iterative optimization module is used to obtain the initial registration pose by inputting the 3D CT image and multi-view X-ray image of the target object into the trained image registration model, and then iteratively optimizes it based on the multi-view constraint loss to obtain the accurate registration pose.
[0015] Embodiments of the present invention also provide a network-side server, comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described registration method based on CT feature projection and multi-view X-ray.
[0016] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described registration method based on CT feature projection and multi-view X-ray.
[0017] The beneficial effects of this invention are as follows: The registration method based on CT feature projection and multi-view X-ray is applicable to the registration of CT and multi-view X-ray images of different target objects (such as human bones and organs). Compared with the prediction model trained using only a single CT image, it effectively improves the generalization ability of the prediction model, can adapt to different clinical application scenarios, and has broad promotion value and practical significance. The spatial constraint effect of multi-view X-ray images is used to make up for the shortcomings of single-view registration, further ensuring the accuracy of iterative registration. Differentiable projection is used to integrate the DRR generation process into the end-to-end neural network training, and the multiple constraints of geodesic loss and cross-geodesic loss are combined to significantly improve the accuracy of initial pose prediction. The initial registration pose is quickly predicted by a deep learning model, and the backpropagation characteristics of differentiable projection are used to accelerate iterative optimization, greatly shortening the registration time and achieving fully automatic and efficient registration. The style alignment of X-ray images and DRR images is completed by iteratively traversing the maximum recovery value method, eliminating modal difference interference, and the multi-view constraints effectively resist the influence of image occlusion and noise, ensuring the stability and robustness of registration in complex scenes. Attached Figure Description
[0018] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0019] In the picture: Figure 1 A flowchart of a registration method based on CT feature projection and multi-view X-ray provided for the first embodiment of the present invention; Figure 2 A schematic diagram of the structure of the image registration model provided in the first embodiment of the present invention; Figure 3 A flowchart illustrating the iterative optimization registration method based on multi-view constraints provided for the first embodiment of the present invention; Figure 4 A schematic diagram of a registration system based on CT feature projection and multi-view X-ray provided for the second embodiment of the present invention; Figure 5 This is a schematic diagram of the network-side server provided according to the third embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] First Implementation Method The first embodiment of the present invention provides a registration method based on CT feature projection and multi-view X-ray, comprising: constructing a CT training dataset and preprocessing it to obtain a first CT image; constructing an image registration model; adding different random perturbations to a standard pose to generate a first projection pose and a second projection pose, calculating a first transformation matrix of the first projection pose and the second projection pose, projecting the first CT image onto the first projection pose and the second projection pose respectively using differentiable projection to obtain a first DRR image and a second DRR image; extracting CT features and DRR features respectively using a dual feature encoder, performing feature fusion and pose prediction on the CT features after differentiable projection and DRR features to obtain a predicted pose and a predicted DRR image; training the model using a combination of similarity loss, geodesic loss, and cross-geodesic loss; acquiring a three-dimensional CT image and a multi-view X-ray image of the target object, preprocessing them, and inputting them into the trained image registration model to obtain an initial registration pose; and iteratively optimizing the registration parameters based on the second transformation matrix between the initial registration pose and the X-ray image using differentiable projection and multi-view constraint loss to obtain an accurate registration pose. This invention presents a registration method based on CT feature projection and multi-view X-rays. This method is applicable to the registration of CT and multi-view X-ray images of different target objects (such as human bones and organs). Compared with prediction models trained using only a single CT image, it effectively improves the generalization ability of the prediction model, can adapt to different clinical application scenarios, and has broad promotional value and practical significance. It utilizes the spatial constraints of multi-view X-ray images to compensate for the shortcomings of single-view registration, further ensuring the accuracy of iterative registration. It adopts differentiable projection to integrate the DRR generation process into end-to-end neural network training, and combines multiple constraints of geodesic loss and cross-geodesic loss to significantly improve the accuracy of initial pose prediction. It uses a deep learning model to quickly predict the initial registration pose, and combines the backpropagation characteristics of differentiable projection to accelerate iterative optimization, greatly shortening the registration time and achieving fully automatic and efficient registration. It completes style alignment between X-ray images and DRR images by iteratively traversing the maximum recovery value method, eliminating modal difference interference, and combines multi-view constraints to effectively resist image occlusion and noise effects, ensuring the stability and robustness of registration in complex scenes.
[0022] The following details the implementation of the registration method based on CT feature projection and multi-view X-ray in this embodiment. The following details are provided for ease of understanding and are not essential for implementing this solution. The specific process of this embodiment is as follows: Figure 1 As shown.
[0023] Step S1: Construct a CT training dataset, obtain a set of 3D CT images from the dataset, and preprocess them to obtain the first CT image.
[0024] Specifically, a CT training dataset is constructed, containing 3D CT images of different objects. A set of 3D CT images from the dataset undergoes 3D preprocessing, including resampling and cropping of the CT images, and truncation and normalization of CT voxels. Resampling adjusts the CT images to a uniform resolution, ensuring consistency in feature extraction; cropping removes invalid regions from the images, reducing redundant information interference; truncation and normalization of CT voxels map the numerical range to a preset interval (e.g., [0,1]), accelerating model convergence and improving the stability of feature extraction.
[0025] As an example, the CT image is first resampled using linear interpolation to adjust the voxel spacing to 1mm×1mm×1mm and unify the image resolution. Then, the image is cropped to a fixed size of 256×256×256 according to the target region. Next, the Hu value of the CT is truncated according to [500, 1000]. Finally, the min-max normalization method is used to map the voxel values to the [0,1] interval to obtain the first CT image.
[0026] Step S2: Construct an image registration model, which includes a first feature encoder, a second feature encoder, differentiable projection, feature fusion, and pose prediction.
[0027] Specifically, such as Figure 2 As shown, the image registration model is a trainable neural network used for predicting projection pose.
[0028] The first feature encoder consists of a 3D residual convolutional network or any other network model for 3D feature extraction, used to extract 3D features from CT images. As an example, the first feature encoder uses a 3D ResNet network, extracting 3D features from CT images through multi-layer 3D convolution, batch normalization, and the ReLU activation function.
[0029] The second feature encoder consists of a 2D residual convolutional network or any other network model used for 2D feature extraction, and is used to extract 2D features from DRR images. As an example, the second feature encoder uses a 2D ResNet network, which extracts 2D features from DRR images through multiple layers of 2D convolution, batch normalization, and the ReLU activation function.
[0030] Differentiable projection, implemented using ray tracing algorithms and deep learning frameworks, supports backpropagation and can be applied to the training process of end-to-end neural networks to perform DRR projection calculations. Compared to traditional non-differentiable projection, this differentiable projection integrates the projection process into the backpropagation chain of the neural network, enabling the model to simultaneously optimize projection parameters and network parameters through gradient descent, significantly improving registration efficiency and accuracy.
[0031] Feature fusion, consisting of channel-dimensional concatenation and convolutional networks or other arbitrary feature fusion modules, is used to effectively fuse projected 2D CT features with 2D DRR features, enhancing the representational power of the features. As an example, the feature fusion module uses channel-dimensional concatenation to stitch together 2D CT features and 2D DRR features, and then uses a multi-layer 2D convolutional network to perform dimensionality reduction and downsampling processing on the fused features to obtain a fused feature map.
[0032] Pose prediction consists of a multilayer perceptron or other arbitrary numerical regression prediction module, used to output perturbation parameters of the projected pose, thereby predicting the projected pose. As an example, pose prediction uses a 3-layer multilayer perceptron with 1024, 512, and 6 neurons in the hidden layers, respectively, and outputs perturbation pose parameters containing 3 rotation angles and 3 translation amounts.
[0033] Step S3: Add different random perturbations to the standard pose to generate a first projection pose and a second projection pose. Calculate the first transformation matrix of the first projection pose and the second projection pose. Project the first CT image onto the first projection pose and the second projection pose using differentiable projection to obtain the first DRR image and the second DRR image.
[0034] Specifically, adding random perturbations to the standard pose to generate the first projected pose includes: setting the standard pose as the orthogonal pose of the target skeleton, adding random perturbations to the standard pose, and generating the first projected pose; wherein the rotation angle perturbation range of the random perturbation is [-50°, 50°], and the translation perturbation range is [-100mm, 100mm].
[0035] Adding random perturbations to the standard pose to generate a second projected pose includes: setting the standard pose as the orthogonal pose of the target skeleton, adding random perturbations different from those mentioned above to the standard pose, and generating a second projected pose; wherein, the rotation angle perturbation range of the random perturbation is [-50°, 50°], and the translation perturbation range is [-100mm, 100mm].
[0036] Calculating the first transformation matrix between the first projected attitude and the second projected attitude involves using a rigid body transformation algorithm to calculate the first transformation matrix between the first projected attitude and the second projected attitude. The first transformation matrix includes a rotation matrix and a translation matrix.
[0037] The method involves projecting a first CT image onto a first CT image using a differentiable projection method based on a first projection pose to generate a first DRR image with a resolution of 512×512. The method supports backpropagation and can integrate the projection process into neural network training to achieve synchronous optimization of projection parameters and network parameters.
[0038] The method involves projecting a first CT image onto a second CT image using differentiable projection based on a second projection pose to obtain a second DRR image. This includes: using differentiable projection based on a ray tracing algorithm and the PyTorch deep learning framework to project the first CT image onto the second projection pose, generating a second DRR image with a resolution of 512×512. Differentiable projection supports backpropagation and can integrate the projection process into neural network training, achieving synchronous optimization of projection parameters and network parameters.
[0039] Step S4: The first CT image is input into the first feature encoder to obtain the first feature, and processed under the standard pose by differentiable projection to obtain the second feature; the preprocessed first DRR image is input into the second feature encoder to obtain the third feature; the second feature and the third feature are input into feature fusion and pose prediction to obtain the predicted perturbation pose, combined with the standard pose to obtain the first predicted pose, and then the first predicted DRR image is obtained by differentiable projection. This step is repeated to obtain the second predicted pose and the second predicted DRR image.
[0040] Specifically, step S4 includes the following steps: Step S41: The first feature encoder of the input image registration model of the first CT image adopts a 3D ResNet network. Through two layers of three-dimensional convolution, batch normalization and ReLU activation function, the three-dimensional features of the first CT image are extracted. The feature dimension is 256×256×256×256, and the first feature is obtained.
[0041] Step S42: Perform differentiable projection on the first feature according to the standard pose to project the three-dimensional feature into a two-dimensional feature with a feature dimension of 256×64×64. Then, use the min-max normalization method to normalize the two-dimensional feature to obtain the second feature.
[0042] Step S43: Perform two-dimensional preprocessing on the first DRR image. The preprocessed first DRR image is input into the second feature encoder of the model. The second feature encoder uses a 2D ResNet network. Through four layers of two-dimensional convolution, batch normalization and ReLU activation function, two-dimensional features of the first DRR image are extracted. The feature dimension is 256×64×64, and the third feature is obtained.
[0043] Specifically, two-dimensional preprocessing of DRR images includes resampling, cropping, and pixel normalization. Resampling adjusts the DRR image to a uniform resolution, ensuring consistency in feature extraction. Cropping removes invalid regions from the image, reducing redundant information interference. Pixel normalization maps numerical ranges to a preset interval (e.g., [0,1]), accelerating model convergence and improving the stability of feature extraction. As an example, pixel normalization uses the min-max normalization method, mapping pixel values to the [0,1] interval.
[0044] Step S44: Input the second and third features into the feature fusion module. Use channel-dimensional concatenation to concatenate the two features along the channel dimension to obtain a fused feature with a feature dimension of 512×64×64. Then, use a two-layer two-dimensional convolutional network to perform dimensionality reduction and downsampling processing on the fused feature to obtain a fused feature map with a dimension of 128×16×16. Input the fused feature into the pose prediction module using a three-layer multilayer perceptron with 1024, 512, and 6 hidden layer neurons, respectively. Output the predicted perturbation pose, which includes three rotation angles and three translation amounts.
[0045] Step S45: Based on the predicted perturbation pose and the standard pose, a first predicted pose is calculated. Then, the first CT image is projected onto the first CT image using differentiable projection to obtain a first predicted DRR image. Wherein, the first predicted pose = standard pose + predicted perturbation pose.
[0046] Step S46: The first CT image is re-input into the first feature encoder to obtain the first feature, and then projected into the standard pose by differentiable projection to obtain the second feature; the second DRR image is input into the second feature encoder after two-dimensional preprocessing, and a 2D ResNet network is used to extract the two-dimensional features of the second DRR image through 4 layers of two-dimensional convolution, batch normalization and ReLU activation function. The feature dimension is 256×64×64, and the fourth feature is obtained.
[0047] Step S47: Input the second feature and the fourth feature into the feature fusion module. Use channel-dimensional concatenation to concatenate the two features according to the channel dimension to obtain a fused feature with a feature dimension of 512×64×64. Then, use a two-layer two-dimensional convolutional network to perform dimensionality reduction and downsampling processing on the fused feature to obtain a fused feature map with a dimension of 128×16×16. Input the fused feature into the pose prediction module. Use a three-layer multilayer perceptron with hidden layer neurons of 1024, 512 and 6 respectively to output the second predicted perturbation pose.
[0048] Step S48: Based on the second predicted perturbation pose and the standard pose, the second predicted pose is calculated, and then the first CT image is projected onto the second predicted pose using differentiable projection to obtain the second predicted DRR image. Wherein, the second predicted pose = standard pose + second predicted perturbation pose.
[0049] Step S5: Calculate the similarity loss between the corresponding DRR image and the predicted DRR image, the geodesic loss between the two projected poses and the two predicted poses, and the cross-geodesic loss. Combine all the losses to train the image registration model.
[0050] Specifically, step S5 includes the following steps: Step S51: Calculate the similarity loss between the first DRR image and the first predicted DRR image, and between the second DRR image and the second predicted DRR image, using Normalized Cross-Correlation (NCC). The formula for calculating the similarity loss is:
[0051] in, For similarity loss, For normalized cross-correlation, For the first DRR image, For the first predicted DRR image, For the second DRR image, This is the second predicted DRR image.
[0052] Step S52: Calculate the geodesic loss between the first projected attitude and the first predicted attitude, and between the second projected attitude and the second predicted attitude, to constrain the accuracy of attitude prediction. The formula for calculating the geodesic loss is:
[0053] in, For geodesic loss, For geodesic distance function, This is the first projection pose. The first predicted posture, This is the second projection pose. This is the second predicted posture.
[0054] Geodesic distance function The calculation formula is:
[0055] in, For geodesic distance function, Let be the focal length of the projection device, Tr() be the trace of the matrix, and T be the transformation matrix for the projection orientation. Let R be the transformation matrix for the predicted attitude, R be the rotation matrix in the transformation matrix for the projected attitude, and t be the translation amount in the transformation matrix for the projected attitude. The rotation matrix is the transformation matrix used to predict the pose. This represents the translation amount in the transformation matrix for predicting the attitude.
[0056] Step S53, the cross geodesic loss includes the geodesic loss calculated from the first projected attitude and the attitude obtained from the second predicted attitude and the first transformation matrix, and the geodesic loss calculated from the second projected attitude and the attitude obtained from the first predicted attitude and the first transformation matrix. The formula for calculating the cross geodesic loss is:
[0057]
[0058]
[0059] in, For cross geodesic loss, For geodesic distance function, To change attitude in the reverse crossover, This is the first projection pose. For positive crossover attitude transformation, This is the second projection pose. This is the first predicted pose output by the model. This is the first transformation matrix. This is the second predicted pose output by the model.
[0060] Step S54: Calculate the joint loss function, use the Adam optimizer, set the learning rate to 1e-4, the number of iterations to 100 rounds, and the batch size to 2, and train the image registration model until the loss function converges to obtain the trained model.
[0061] Specifically, the formula for calculating the joint loss function is as follows:
[0062] in, For the total loss, For similarity loss, For geodesic loss, This is the loss due to cross geodesics.
[0063] Step S6: Acquire the 3D CT image and multi-view X-ray image of the target object, preprocess them, and input them into the trained image registration model to obtain the initial registration pose; based on the second transformation matrix between the initial registration pose and the X-ray image, use differentiable projection and multi-view constraint loss for iterative optimization to update the registration parameters and obtain the accurate registration pose.
[0064] Specifically, step S6 includes the following steps: Step S61: Obtain the 3D CT image and multi-view X-ray image of the target object. Perform 3D preprocessing on the CT image. Perform style alignment with the DRR image on the multi-view X-ray image by iteratively traversing the maximum recovery value method. Select one of the X-ray images for 2D preprocessing. Input the preprocessed CT image and X-ray image into the trained image registration model to predict the initial registration pose.
[0065] Specifically, such as Figure 3 As shown, acquiring the 3D CT image and multi-view X-ray image of the target object includes acquiring the 3D CT image of the target object and at least two X-ray images from different perspectives. For example, acquiring the 3D CT image and three multi-view X-ray images (anteroposterior, lateral, and oblique views) of the target object (e.g., a patient with a spinal fracture). The CT image is processed according to the 3D preprocessing steps in step S1, including resampling, cropping, CT voxel truncation, and normalization.
[0066] Style alignment of multi-view X-ray images using an iterative maximum recovery value (MRF) method involves: traversing the grayscale range of the X-ray image, adjusting grayscale contrast and brightness to ensure the grayscale distribution of the X-ray image matches that of the DRR image, thus eliminating style differences. Style alignment can reduce interference from style differences between real X-ray images and DRR images, improving registration accuracy.
[0067] Select an orthogonal X-ray image and process it according to the two-dimensional preprocessing steps in step S43, including resampling, cropping, and pixel normalization. Input the preprocessed CT image and orthogonal X-ray image into the trained image registration model to predict the initial registration pose.
[0068] Step S62: Obtain the second transformation matrix between multi-view X-ray images. Based on the initial registration pose, iteratively optimize using differentiable projection and multi-view constraint loss to update the registration parameters and obtain the accurate registration pose.
[0069] Specifically, a second transformation matrix is obtained between multi-view X-ray images. This second transformation matrix includes multiple transformation matrices derived from the calibration parameters used during intraoperative X-ray image acquisition, enabling spatial transformation between any two viewpoints. As an example, a second transformation matrix is obtained between three multi-view X-ray images. This second transformation matrix is derived from the calibration parameters used during X-ray machine image acquisition and contains three transformation matrices, corresponding to spatial transformations between anteroposterior and lateral views, anteroposterior and oblique views, and lateral and oblique views, respectively.
[0070] The multi-view predicted pose is obtained based on the initial registration pose and the second transformation matrix. Differentiable projection algorithm is used to obtain the DRR images corresponding to each viewpoint. The similarity loss between the multi-view X-ray images and the corresponding DRR images is calculated as the multi-view constraint loss. A gradient descent iterative optimization method based on a deep learning framework is employed, aiming to minimize the multi-view constraint loss, and the registration parameters (rotation angle and translation) are iteratively updated. As an example, a gradient descent iterative optimization method based on the PyTorch framework is used with a learning rate of 1e-5. Iteration stops when the error converges to an acceptable range, resulting in an accurate registration pose, thus completing the registration of CT and multi-view X-ray images.
[0071] Multi-view constraint loss utilizes the initial pose and a second transformation matrix to obtain the predicted pose from multiple views, then uses a differentiable projection algorithm to obtain the corresponding DRR image, and calculates the similarity loss between the multi-view X-ray image and the corresponding DRR image. The second transformation matrix is obtained during intraoperative X-ray image acquisition and contains multiple transformation matrices, enabling spatial transformation between any two views. Multi-view constraint can fully utilize X-ray image information from different views, compensate for the limitations of single-view registration, and improve the robustness and accuracy of registration, making it particularly suitable for registration scenarios involving complex anatomical structures.
[0072] This invention presents a registration method based on CT feature projection and multi-view X-rays. This registration method is applicable to the registration of CT and multi-view X-ray images of different target objects (such as human bones and organs). Compared to prediction models trained using only a single CT image, it effectively improves the generalization ability of the prediction model, adapts to different clinical application scenarios, and has broad promotional value and practical significance. Simultaneously, it utilizes the constraint effect of multi-view X-ray images to compensate for the shortcomings of single-view registration, further improving the iterative registration accuracy and meeting the needs of scenarios with high registration accuracy requirements, such as clinical surgical navigation. It employs differentiable projection, integrating the projection process into end-to-end neural network training. The system combines geodesic loss and cross-geodesic loss constraints to effectively improve the accuracy of initial pose prediction. It rapidly predicts the initial registration pose using a deep learning model, and integrates differentiable projection to support backpropagation, accelerating the iterative registration process. Compared to traditional registration methods, this significantly shortens registration time, eliminates the need for manual intervention, and achieves fully automated registration, improving the efficiency of clinical applications. It uses an iterative maximum recovery value method to achieve style alignment between X-ray and DRR images, reducing interference from style differences between the two types of images. Multi-view constraints effectively resist the effects of image occlusion and noise, ensuring stable and accurate registration even in complex scenarios.
[0073] Second implementation method: like Figure 4As shown, the second embodiment of the present invention provides a registration system based on CT feature projection and multi-view X-ray, the system including: a data preprocessing module 201, a modeling module 202, a training sample generation module 203, a pose prediction module 204, a loss calculation training module 205, and a registration iterative optimization module 206.
[0074] Specifically, the data preprocessing module 201 is used to construct a CT training dataset, obtain a set of 3D CT images from the dataset, and preprocess them to obtain a first CT image; the modeling module 202 is used to construct an image registration model, which includes a first feature encoder, a second feature encoder, differentiable projection, feature fusion, and pose prediction; the training sample generation module 203 is used to add different random perturbations to the standard pose to generate a first projection pose and a second projection pose, calculate a first transformation matrix of the first projection pose and the second projection pose, and project the first CT image according to the first projection pose and the second projection pose respectively through differentiable projection to obtain a first DRR image and a second DRR image; the pose prediction module 204 is used to input the first CT image into the first feature encoder to obtain a first feature, process it under the standard pose through differentiable projection to obtain a second feature; and input the preprocessed first DRR image into the second feature encoder. The system obtains the third feature; the second and third features are input features and the pose prediction is used to obtain the predicted perturbation pose, which is combined with the standard pose to obtain the first predicted pose, and then differentiable projection is used to obtain the first predicted DRR image. This step is repeated to obtain the second predicted pose and the second predicted DRR image. The loss calculation training module 205 is used to calculate the similarity loss between the corresponding DRR image and the predicted DRR image, the geodesic loss between the two projected poses and the two predicted poses, and the cross-geodesic loss. All losses are combined to train the image registration model. The registration iteration optimization module 206 is used to acquire the three-dimensional CT image and multi-view X-ray image of the target object, and after preprocessing, input them into the trained image registration model to obtain the initial registration pose. Based on the second transformation matrix between the initial registration pose and the X-ray image, iterative optimization is performed using differentiable projection and multi-view constraint loss to update the registration parameters and obtain the accurate registration pose.
[0075] It is not difficult to see that this embodiment is a system implementation corresponding to the first embodiment, and this embodiment can be implemented in conjunction with the first embodiment. The relevant technical details mentioned in the first embodiment are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the first embodiment.
[0076] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.
[0077] The third embodiment of the present invention relates to a network-side server, such as... Figure 5 As shown, it includes at least one processor 302; and a memory 301 communicatively connected to at least one processor 302; wherein the memory 301 stores instructions executable by at least one processor 302, the instructions being executed by at least one processor 302 to enable at least one processor 302 to perform the above-described data processing method.
[0078] The memory 301 and processor 302 are connected via a bus, which may include any number of interconnecting buses and bridges. The bus connects various circuits of one or more processors 302 and memory 301 together. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 302 is transmitted over a wireless medium via an antenna, which further receives data and transmits it to processor 302.
[0079] Processor 302 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory 301 can be used to store data used by processor 302 during operation.
[0080] The fourth embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the registration method based on CT feature projection and multi-view X-rays in the first embodiment.
[0081] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0082] The above descriptions are merely embodiments of the present invention. Commonly known structures and characteristics are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.
[0083] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A registration method based on CT feature projection and multi-view X-ray, characterized in that, include: S1, construct a CT training dataset, obtain a set of 3D CT images from the dataset and preprocess them to obtain the first CT image; S2, Construct an image registration model, which includes a first feature encoder, a second feature encoder, a differentiable projection, feature fusion, and pose prediction; S3, different random perturbations are added to the standard pose and differentiable projections are performed to obtain the first DRR image and the second DRR image; S4, input the first CT image into the first feature encoder to obtain the first feature, process the first feature through differentiable projection under the standard pose to obtain the second feature; input the preprocessed first DRR image into the second feature encoder to obtain the third feature; input the second feature and the third feature into feature fusion and pose prediction to obtain the predicted perturbation pose, combine with the standard pose to obtain the first predicted pose, and then obtain the first predicted DRR image through differentiable projection. Repeat this step to obtain the second predicted pose and the second predicted DRR image, including: The first feature encoder of the first CT image input image registration model uses three-dimensional convolution, normalization and activation functions to extract the three-dimensional features of the first CT image to obtain the first feature; The first feature is subjected to a differentiable projection according to the standard pose, and the three-dimensional feature is projected into a two-dimensional feature. Then, the normalization process is performed to obtain the second feature. The first DRR image is preprocessed in two dimensions. The preprocessed first DRR image is then input into the second feature encoder of the model. Two-dimensional convolution, normalization and activation functions are used to extract the two-dimensional features of the first DRR image to obtain the third feature. Based on the predicted perturbation pose and the standard pose, the first predicted pose is calculated, and then the first CT image is projected onto the first predicted pose using differentiable projection to obtain the first predicted DRR image. The first CT image is re-input into the first feature encoder to obtain the first feature, which is then projected onto the standard pose using a differentiable projection to obtain the second feature; the second DRR image is preprocessed in two dimensions and then input into the second feature encoder to extract the two-dimensional features of the second DRR image to obtain the fourth feature. The second and fourth features are input into the feature fusion module. The two features are concatenated along the channel dimension to obtain the fused feature. The fused feature is then input into the pose prediction module to output the second predicted perturbation pose. The second predicted pose is calculated based on the second predicted perturbation pose and the standard pose. Then, the first CT image is projected onto the second predicted pose using differentiable projection to obtain the second predicted DRR image. S5, calculate the similarity loss, geodesic loss and cross-geodesic loss respectively, and train the image registration model by combining all the losses; S6: Obtain the 3D CT image and multi-view X-ray image of the target object and input them into the trained image registration model to obtain the initial registration pose. Iteratively optimize based on the multi-view constraint loss to obtain the accurate registration pose.
2. The registration method based on CT feature projection and multi-view X-ray according to claim 1, characterized in that, The three-dimensional preprocessing of a set of three-dimensional CT images in the dataset includes resampling and cropping of the CT images, truncation and normalization of CT voxels.
3. The registration method based on CT feature projection and multi-view X-ray according to claim 1, characterized in that, The first feature encoder is composed of a three-dimensional residual convolutional network, used to extract three-dimensional features of CT images; The second feature encoder consists of a two-dimensional residual convolutional network, used to extract two-dimensional features of the DRR image; the differentiable projection is implemented based on the ray tracing algorithm and deep learning framework, used to complete the DRR projection calculation. Feature fusion consists of channel-dimensional concatenation and convolutional networks, used to effectively fuse projected 2D CT features with 2D DRR features, thereby enhancing the representational power of the features. Attitude prediction consists of a multilayer perceptron, which outputs perturbation parameters of the projected attitude to achieve the prediction of the projected attitude.
4. The registration method based on CT feature projection and multi-view X-ray according to claim 1, characterized in that, The step of adding different random perturbations to the standard pose and performing differentiable projection to obtain the first DRR image and the second DRR image includes: adding different random perturbations to the standard pose to generate the first projection pose and the second projection pose, calculating the first transformation matrix of the first projection pose and the second projection pose, and projecting the first CT image according to the first projection pose and the second projection pose respectively through differentiable projection to obtain the first DRR image and the second DRR image. The calculation of the first transformation matrix between the first projected pose and the second projected pose includes using a rigid body transformation algorithm to calculate the first transformation matrix between the first projected pose and the second projected pose. The first transformation matrix includes a rotation matrix and a translation matrix. The addition of random perturbation to the standard pose to generate the first projected pose includes setting the standard pose as the orthogonal pose of the target skeleton, adding random perturbation to the standard pose, and generating the first projected pose.
5. The registration method based on CT feature projection and multi-view X-ray according to claim 1, characterized in that, Step S5 includes: Calculate the geodesic loss between the first projected attitude and the first predicted attitude, and between the second projected attitude and the second predicted attitude, to constrain the accuracy of attitude prediction; The cross geodesic loss includes the geodesic loss calculated from the first projected attitude and the attitude obtained from the second predicted attitude and the first transformation matrix, and the geodesic loss calculated from the second projected attitude and the attitude obtained from the first predicted attitude and the first transformation matrix.
6. A registration system based on CT feature projection and multi-view X-ray, characterized in that, The system, applied to the registration method based on CT feature projection and multi-view X-ray as described in any one of claims 1-5, comprises: The data preprocessing module is used to construct a CT training dataset, obtain a set of 3D CT images from the dataset, and preprocess them to obtain the first CT image. The modeling module is used to construct an image registration model, which includes a first feature encoder, a second feature encoder, a differentiable projection, feature fusion, and pose prediction. The training sample generation module is used to add different random perturbations to the standard pose and perform differentiable projection to obtain the first DRR image and the second DRR image. The pose prediction module is used to input a first CT image into a first feature encoder to obtain a first feature, and process the first feature under a standard pose through differentiable projection to obtain a second feature; input the preprocessed first DRR image into a second feature encoder to obtain a third feature; input the second feature and the third feature into feature fusion and pose prediction to obtain a predicted perturbation pose, combine it with the standard pose to obtain a first predicted pose, and then obtain a first predicted DRR image through differentiable projection. This step is repeated to obtain a second predicted pose and a second predicted DRR image. Specifically, it includes: inputting the first CT image into the first feature encoder of the image registration model, using three-dimensional convolution, normalization and activation functions to extract the three-dimensional features of the first CT image to obtain the first feature; The first feature is subjected to a differentiable projection according to the standard pose, and the three-dimensional feature is projected into a two-dimensional feature. Then, the normalization process is performed to obtain the second feature. The first DRR image is preprocessed in two dimensions. The preprocessed first DRR image is then input into the second feature encoder of the model. Two-dimensional convolution, normalization and activation functions are used to extract the two-dimensional features of the first DRR image to obtain the third feature. Based on the predicted perturbation pose and the standard pose, the first predicted pose is calculated, and then the first CT image is projected onto the first predicted pose using differentiable projection to obtain the first predicted DRR image. The first CT image is re-input into the first feature encoder to obtain the first feature, which is then projected onto the standard pose using a differentiable projection to obtain the second feature; the second DRR image is preprocessed in two dimensions and then input into the second feature encoder to extract the two-dimensional features of the second DRR image to obtain the fourth feature. The second and fourth features are input into the feature fusion module. The two features are concatenated along the channel dimension to obtain the fused feature. The fused feature is then input into the pose prediction module to output the second predicted perturbation pose. The second predicted pose is calculated based on the second predicted perturbation pose and the standard pose. Then, the first CT image is projected onto the second predicted pose using differentiable projection to obtain the second predicted DRR image. The loss calculation training module is used to calculate the similarity loss, geodesic loss and cross-geodesic loss respectively, and to train the image registration model by combining all the losses. The registration iterative optimization module is used to obtain the initial registration pose by inputting the 3D CT image and multi-view X-ray image of the target object into the trained image registration model, and then iteratively optimizes it based on the multi-view constraint loss to obtain the accurate registration pose.
7. A network-side server, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the registration method based on CT feature projection and multi-view X-ray as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the registration method based on CT feature projection and multi-view X-ray as described in any one of claims 1 to 5.