A preoperative three-dimensional CT registration method based on a single X-ray image

By using a single-view CT reconstruction network in CT registration, converting X-ray images into three-dimensional CT, and combining ICP algorithm for three-dimensional spatial registration, the problem that two-dimensional matching in the existing technology cannot accurately describe the three-dimensional alignment relationship is solved, and high-precision CT registration is achieved.

CN119693424BActive Publication Date: 2025-05-27DALIAN UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510194940.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The existing CT registration methods are matched in two-dimensional space, and cannot fully accurately describe the alignment relationship in three-dimensional space, resulting in low registration accuracy and susceptible to artifacts and noise.

Method used

A single-view CT based on a single X-ray image is used to reconstruct the network, convert the two-dimensional X-ray image into three-dimensional CT, and CT registration is performed in three-dimensional space using the ICP algorithm.

Benefits of technology

It improves the accuracy of CT registration, realizes submillimeter-level registration, simplifies the operation process, and is especially suitable for spinal orthopedic surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693424B_ABST
    Figure CN119693424B_ABST
Patent Text Reader

Abstract

The present invention belongs to the fields of medical imaging and artificial intelligence, and discloses a preoperative three-dimensional CT registration method based on a single X-ray image. The present invention proposes to first use an X-ray image to reconstruct the CT of the patient's current pose, and then use the reconstructed CT and the preoperative CT for CT registration in three-dimensional space, solving the problem of dimensional mismatch in previous registration methods. The present invention proposes a single-view CT reconstruction network to convert a two-dimensional X-ray image into a three-dimensional CT, and then uses the ICP algorithm to achieve the registration of the CT generated by the network and the preoperative CT. The present invention realizes CT registration in three-dimensional space only using X-ray images, effectively improving the accuracy of CT registration and achieving sub-millimeter-level CT registration. Especially in spinal orthopedic surgeries, it can accurately determine the position of the patient's anatomical structure, helping doctors perform surgeries precisely and efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of medical imaging and artificial intelligence, and particularly relates to a preoperative three-dimensional CT registration method based on a single X-ray image. Background Art

[0002] CT (Computed Tomography) registration is a technique for aligning CTs acquired from the same patient at different time points or using different devices. In orthopedic surgery, doctors will perform precise preoperative planning based on the patient's preoperative CT, such as determining the optimal implantation path and angle of screws, and simulating the resection of diseased intervertebral discs. Since the patient's pose during surgery is different from that during the acquisition of the preoperative CT, CT registration is required. By making the preoperative CT consistent with the patient's pose during surgery through CT registration, the positional relationship of anatomical structures can be accurately determined, providing accurate three-dimensional positioning information for the intraoperative navigation system and helping doctors to monitor the relative position of the instrument and the target tissue in real time. Through precise navigation after registration, doctors can effectively reduce the errors caused by visual occlusion or tissue displacement during the operation, avoid important structures such as bone marrow, blood vessels, and nerves, and perform the operation accurately and efficiently.

[0003] Currently, the widely used registration method in orthopedic surgery is to take two X-ray images of the patient, a frontal view and a lateral view, during the operation, and calculate the rotation and offset of the patient's body position, that is, the registration transformation matrix, by comparing the bone features in the simulated projections (DRRs) of these two orientations of the preoperative CT. Then, the registration transformation matrix is applied to the preoperative CT to complete the CT registration, making the preoperative CT consistent with the patient's current pose. However, CT is in three-dimensional space, while X-ray images and DRR images are in two-dimensional space. Due to the complex anatomical structure of the human body and tissue overlap during the projection process, two-dimensional image matching cannot fully and accurately describe the alignment relationship in three-dimensional space, and the small errors generated by two-dimensional space matching will be amplified in three-dimensional space. At the same time, X-ray images are easily affected by artifacts and noise, which will affect the extraction and matching of bone features. The registration accuracy of this method is usually at the millimeter level.

[0004] In order to improve the accuracy of CT registration, researchers have been exploring new registration methods. For example, Patent CN114511597B improves on the above method. When comparing the frontal view and lateral view X-ray images and DRR images, both types of images are downsampled to multiple resolution sizes, and the differences between the images are calculated at different scales. It also proposes a mutual information measure method to evaluate the similarity of registration. Patent CN118037793B proposes to utilize the ability of neural network feature extraction. During registration, multiple X-ray images of the patient are taken, and the anatomical structure features in the two-dimensional X-ray images and three-dimensional CT are extracted and corresponding in the feature space, thereby completing the CT registration.

[0005] Different from the existing CT registration methods that compare the differences between X-ray images and DRR images in a two-dimensional space or directly extract the features of two-dimensional X-ray images and match them with three-dimensional CT features, the present invention proposes a preoperative three-dimensional CT registration method based on a single X-ray image. First, a three-dimensional CT of the patient's current pose is restored using a single X-ray image, and then the restored three-dimensional CT is registered with the preoperative CT in three-dimensional space, solving the problem of dimension mismatch in the direct registration of two-dimensional X-ray images and three-dimensional CT in the previous methods, and effectively improving the accuracy of CT registration. Specifically, the present invention proposes a CT reconstruction method based on single-view projection, which converts a two-dimensional X-ray image into a three-dimensional CT by training a single-view CT reconstruction network, and then uses the ICP (Iterative Closest Point) algorithm to register the CT reconstructed by the network with the preoperative CT in three-dimensional space. The present invention overcomes the problem of dimension mismatch in the previous registration methods and can improve the registration accuracy to the sub-millimeter level. Summary of the Invention

[0006] The present invention provides a preoperative three-dimensional CT registration method based on a single X-ray image, which can achieve CT registration in three-dimensional space using a single X-ray image, effectively improving the registration accuracy. The CT registration method proposed by the present invention includes two parts: a single-view CT reconstruction network and the ICP algorithm, and the overall architecture is as Figure 1 shown. The specific steps are as follows:

[0007] Step 1: Construct an "X-ray - CT" dataset

[0008] Obtain the preoperative CT of the patient and a single X-ray image taken in the intraoperative frontal view. Perform displacement and rotation of the preoperative CT in various poses, and generate corresponding DRR images by simulating the projection of the preoperative CT according to the imaging parameters of the taken X-ray image to construct a dataset for training and validating the single-view CT reconstruction network.

[0009] Step 2: Build a single-view CT reconstruction network

[0010] The single-view CT reconstruction network converts a two-dimensional X-ray image into a three-dimensional CT. The single-view CT reconstruction network includes an encoder module, a feature enhancement module, a decoder module, and a feature conversion module, and the network architecture is as Figure 2 shown.

[0011] (1) Encoder module: The encoder module includes a channel expansion module and four downsampling modules with the same structure, which perform feature extraction and progressive downsampling on the two-dimensional X-ray image, converting it into a compact representation with a higher number of channels. The channel expansion module is mainly composed of a convolutional layer and a residual block. The input of the convolutional layer is the two-dimensional X-ray image with an input channel number of 1 and an output channel number of 210. The convolutional layer is used to increase the channel number of the two-dimensional X-ray image and enhance the feature representation ability of the single-view CT reconstruction network. The residual block is divided into residual block A and residual block B, which are connected in the form of "residual block A - residual block B - residual block B - residual block A". The structure of residual block A is as Figure 3 shown. Residual block A doubles the channel number. The structure of residual block B is as Figure 4 shown. Residual block B keeps the channel number unchanged and enhances the feature extraction ability of the single-view CT reconstruction network. The downsampling module is mainly composed of a max pooling layer and a residual block, which are connected in the form of "max pooling layer - residual block B - residual block B - residual block B". The 2×2 max pooling layer can halve the size of the feature map. Residual block B can extract deep features. During the downsampling process, the number of feature channels remains unchanged, and the size of the feature map gradually becomes smaller. The single-view CT reconstruction network extracts deeper feature information.

[0012] (2) Feature enhancement module: The feature enhancement module is used to explore the dependency relationships between feature maps and enhance the single-view CT reconstruction network's understanding of global information. After multiple layers of feature extraction by the encoder module, the underlying semantic features in the input feature map are obtained, which contain the structural information required for accurate reconstruction. A self-attention-based feature enhancement module is added to the bottom layer of the single-view CT reconstruction network to fully capture the long-range dependency relationships in the feature map, enhance the single-view CT reconstruction network's understanding of global information, and achieve accurate CT reconstruction. The input of the feature enhancement module is the feature map output by the encoder module. First, position encoding is performed on the feature map, and then it is sent to the Transformer module for processing. The Transformer module is composed of a normalization layer, a multi-head attention layer, and a feed-forward network layer, which are connected in the form of "normalization layer - multi-head attention layer - normalization layer - feed-forward network layer". The multi-head attention layer captures the dependency relationships between each feature map. After passing through the normalization layer to stabilize the training process, the feed-forward network layer finally enhances the feature representation ability. Finally, the feature enhancement module outputs the enhanced feature map.

[0013] (3) Decoder module: The decoder module is used to gradually restore the spatial resolution of the enhanced feature map and integrate the encoded features through the feature transformation module, and gradually convert the enhanced feature map into a complete CT. The decoder module includes an upsampling module and a one-dimensional convolutional layer. The upsampling module corresponds to the downsampling module in the encoder module and gradually restores the size of the enhanced feature map. Then, the number of channels is reduced from 840 to 128 through a 1×1 convolutional layer, and finally a complete three-dimensional CT is output.

[0014] (4) Feature transformation module: The feature transformation module is used to solve the problem of inconsistent information dimensions between the projection domain and the image domain. The input of the feature transformation module is the projection domain feature maps with different resolutions output by the encoder module. After the projection domain feature maps are input, they first pass through a feature screening module, which consists of a global average pooling layer, a one-dimensional convolutional layer, and a Sigmoid activation function. The global average pooling layer can reduce the computational amount, the one-dimensional convolutional layer adaptively adjusts the weights of different semantics in the projection domain feature maps, and the Sigmoid activation function performs channel weighting on the input projection domain features to obtain the screened projection domain features with rich information. Then, the projection domain features are converted into CT features in the image domain using a max pooling layer, a 2D residual layer, and a transposed convolutional layer. Finally, the feature transformation module outputs the image domain feature map after semantic transformation to the decoder module.

[0015] Step 3: Train the single-view CT reconstruction network

[0016] In the dataset established in Step 1, the paired DRR images and CTs are used as the input and supervision of the single-view CT reconstruction network respectively, where the DRR images are named x and the corresponding CTs are named q. The loss function used consists of the MSE loss and the MAE loss, and are the weight parameters of the loss, as shown in Equation (1.1):

[0017] (1.1)

[0018] After training for a certain number of epochs, converges, and the training of the single-view CT reconstruction network is completed. When the training is completed, an X-ray image taken in the intraoperative direct view position obtained in Step 1 is input into the single-view CT reconstruction network, and its corresponding three-dimensional CT is output.

[0019] Step 4: Apply the ICP algorithm for CT registration in three-dimensional space

[0020] Input an X-ray image taken in the intraoperative frontal view obtained in Step 1 into the single-view CT reconstruction network trained in Step 3, and a three-dimensional CT will be output. First, convert the preoperative CT and the CT generated by the single-view CT reconstruction network into point cloud representations, namely the source point cloud and the target point cloud. The objective function of the ICP algorithm is shown in Equation (1.2):

[0021] (1.2)

[0022] where, and are points in the source point cloud and the target point cloud respectively, is the total number of points in the source point cloud, is the rotation matrix, is the translation vector. Each iteration of the ICP algorithm minimizes this objective function. First, for each point in the source point cloud, find the nearest point in the target point cloud as the corresponding point pair. Then, based on the current corresponding point pair, according to the principle of minimizing the objective function, calculate the optimal rotation matrix and translation vector such that the source point cloud aligns with the target point cloud after transformation. Apply the calculated and to the source point cloud. If the change in transformation is less than the preset threshold or the maximum number of iterations is reached, the algorithm terminates. Otherwise, repeat the steps of finding the corresponding point pair, calculating the rotation matrix and translation vector and applying them to the source point cloud. The finally obtained registration transformation matrix can accurately describe the translational and rotational relationships in space between the preoperative CT and the CT generated by the network. Apply the registration transformation matrix to the preoperative CT to complete CT registration.

[0023] Effects and benefits of the present invention:

[0024] (1) The present invention proposes a three-dimensional CT registration method based on a single X-ray image, which realizes CT registration in three-dimensional space only using X-ray images, effectively improves the accuracy of CT registration, achieves sub-millimeter-level CT registration, especially in spinal orthopedic surgeries, can accurately determine the positions of the patient's anatomical structures, and helps doctors perform surgeries precisely and efficiently.

[0025] (2) The present invention only requires one X-ray to achieve preoperative CT registration. Compared with the previous CT registration method that requires at least two X-ray films taken in the frontal view and the lateral view, it simplifies the operation process and is of great significance for some scenarios in orthopedic surgeries where it is difficult to take lateral X-ray films. Brief Description of the Drawings

[0026] Figure 1 is the overall architecture diagram of the single X-ray and preoperative CT registration method of the present invention.

[0027] Figure 2 It is the architecture diagram of the single-view CT reconstruction network of the present invention.

[0028] Figure 3 It is the architecture diagram of residual block A in the single-view CT reconstruction network of the present invention.

[0029] Figure 4 It is the architecture diagram of residual block B in the single-view CT reconstruction network of the present invention. Detailed implementation manner

[0030] The following further illustrates the detailed implementation manner of the present invention in combination with the accompanying drawings and technical solutions.

[0031] First, on the premise of ensuring the privacy and rights of patients, preoperative CT of patients and a frontal X-ray image taken during surgery are obtained from a cooperative hospital. A dataset is constructed according to Step 1 for training and validating the single-view CT reconstruction network.

[0032] Then, according to Step 2 and in combination with Figure 2 Build a single-view CT reconstruction network: First, construct an encoder module. Input a two-dimensional X-ray image, expand the number of channels to 210 through the convolutional layer of the channel expansion module, and extract rich features through the cascading of multiple residual blocks. Subsequently, the feature map enters four downsampling modules with the same structure in sequence, gradually reducing the size of the feature map from 512×512 to 32×32. Secondly, build a feature enhancement module behind the encoder module, capture long-range dependence relationships through the self-attention mechanism, and improve the global information modeling ability in combination with positional encoding and multi-head attention. Then, construct a decoder module, which is symmetrically designed based on the downsampling structure of the encoder module, replaces the downsampling layer with an upsampling layer, gradually restores the resolution of the feature map to 512×512, and finally reduces the number of channels to 128 through 1×1 convolution to output a high-quality CT. Finally, construct a feature transformation module, which connects the encoder module and the decoder module to solve the semantic difference problem between the projection domain and the image domain.

[0033] Secondly, train the single-view CT reconstruction network according to Step 3. Specific details: Select 300 pairs of DRR images and CT data as the training set and 10 pairs as the validation set. Train according to the process of Step 3, and it takes 7 hours to complete the training under RTX4090 GPU.

[0034] After completing the training and validation of the network, input a frontal X-ray image taken during surgery into the network according to Step 3 to obtain a three-dimensional CT. Then, convert the CT generated by the network and the preoperative CT into point clouds according to Step 4, apply the ICP algorithm to calculate the registration transformation matrix, and apply the registration transformation matrix to the preoperative CT to complete CT registration.

Claims

1. A preoperative 3D CT registration method based on a single X-ray image, characterized in that: This preoperative 3D CT registration method based on a single X-ray image includes two parts: a single-view CT reconstruction network and an ICP algorithm; Here are the steps: Step 1: Build the "X-ray-CT" dataset Obtain the patient's preoperative CT and an X-ray image taken in the embrasure position during the operation; perform displacement and rotation of the preoperative CT in various postures, simulate projection of the preoperative CT according to the imaging parameters of the X-ray image to generate the corresponding DRR image to construct a data set for training and verifying the single-view CT reconstruction network; Step 2: Build a single-view CT reconstruction network The single-view CT reconstruction network converts a two-dimensional X-ray image into a three-dimensional CT. The single-view CT reconstruction network includes an encoder module, a feature enhancement module, a decoder module, and a feature conversion module. (1) Encoder module: The encoder module includes a channel expansion module and four downsampling modules with the same structure. It extracts features and gradually downsamples the two-dimensional X-ray image to convert it into a compact representation with a higher number of channels. The channel expansion module is mainly composed of a convolutional layer and a residual block. The input of the convolutional layer is a two-dimensional X-ray image with 1 input channel and 210 output channels. The convolutional layer is used to increase the number of channels of the two-dimensional X-ray image and enhance the feature representation capability of the single-view CT reconstruction network. The residual block is divided into residual block A and residual block B. According to the "residual block A-residual block B", the residual block A is the input of the two-dimensional X-ray image with 1 input channel and 210 output channels. The ... The residual block A doubles the number of channels, and the residual block B keeps the number of channels unchanged to enhance the feature extraction capability of the single-view CT reconstruction network. The downsampling module is mainly composed of a maximum pooling layer and a residual block, which are connected in the form of "maximum pooling layer-residual block B-residual block B-residual block B". The 2×2 maximum pooling layer halves the size of the feature map, and the residual block B can extract deep features. In the process of downsampling, the number of feature channels remains unchanged, and the size of the feature map gradually decreases. The single-view CT reconstruction network extracts deeper feature information. (2) Feature enhancement module: The feature enhancement module is used to mine the dependencies between feature maps and enhance the single-view CT reconstruction network's understanding of global information. After multi-layer feature extraction by the encoder module, the underlying semantic features of the input feature map are obtained, which contain the structural information required for accurate reconstruction. A self-attention-based feature enhancement module is added to the bottom layer of the single-view CT reconstruction network to fully capture the long-distance dependencies in the feature map, enhance the single-view CT reconstruction network's understanding of global information, and achieve accurate CT reconstruction. The input of the feature enhancement module is the feature map of the underlying semantic features of the encoder module. The feature map is first positionally encoded and then sent to the Transformer module for processing. The Transformer module consists of a normalization layer, a multi-head attention layer, and a feedforward network layer, which are connected in the form of "normalization layer-multi-head attention layer-normalization layer-feedforward network layer". The multi-head attention layer captures the dependencies between the feature maps. After the normalization layer stabilizes the training process, the feature representation capability is finally enhanced through the feedforward network layer. The final feature enhancement module outputs the enhanced feature map; (3) Decoder module: The decoder module is used to gradually restore the spatial resolution of the enhanced feature map and integrate the encoded features through the feature conversion module to gradually convert the enhanced feature map into a complete CT. The decoder module includes an upsampling module and a one-dimensional convolutional layer. The upsampling module corresponds to the downsampling module in the encoder module and gradually restores the size of the enhanced feature map. Then, the number of channels is reduced from 840 to 128 through a 1×1 convolutional layer, and finally a complete three-dimensional CT is output. (4) Feature conversion module: The feature conversion module is used to solve the problem of inconsistent information dimensions between the projection domain and the image domain. The input of the feature conversion module is the projection domain feature map of different resolutions output by the encoder module. After the projection domain feature map is input, it first passes through a feature screening module. The feature screening module consists of a global average pooling layer, a one-dimensional convolution layer, and a Sigmoid activation function. The global average pooling layer is used to reduce the amount of calculation. The one-dimensional convolution layer adaptively adjusts the weights of different semantics in the projection domain feature map. The Sigmoid activation function performs channel weighting on the input projection domain feature map to obtain a filtered information-rich projection domain feature map. Then, the maximum pooling layer, the 2D residual layer, and the deconvolution layer are used to convert the filtered information-rich projection domain feature map into a CT feature map in the image domain. Finally, the feature conversion module outputs the CT feature map of the image domain after semantic conversion to the decoder module; Step 3: Train the single-view CT reconstruction network In the dataset established in step 1, paired DRR images and CT are used as input and supervision of the single-view CT reconstruction network, respectively. The DRR image is named x, and the corresponding CT is named q. The loss function used is composed of MSE loss and MAE loss. λ1 and λ2 are the weight parameters of the loss, as shown in formula (0.2): LOSS=λ1MSE(x,q)+λ2MAE(x,q)(0.2) After training to a certain number of generations, LOSS converges and the single-view CT reconstruction network training is completed; when the training is completed, an X-ray image taken in the embrasure position during the operation obtained in step 1 is input into the single-view CT reconstruction network, and its corresponding three-dimensional CT is output; Step 4: Apply the ICP algorithm for CT registration in 3D space Input an intraoperative X-ray image obtained in step 1 and taken in the embrasure position into the single-view CT reconstruction network trained in step 3, and output a three-dimensional CT. First, the preoperative CT and the CT generated by the single-view CT reconstruction network are converted into point cloud representations, namely, source point cloud and target point cloud. The objective function of the ICP algorithm is shown in formula (1.2): Among them, p i and q i are points in the source point cloud and the target point cloud respectively, n is the total number of points in the source point cloud, R is the rotation matrix, and t is the translation vector; the ICP algorithm minimizes the objective function in each iteration; first, for each point in the source point cloud, find the nearest point in the target point cloud as the corresponding point pair; then, based on the current corresponding point pair, according to the principle of minimizing the objective function, calculate the optimal rotation matrix R and translation vector t, so that the source point cloud is aligned with the target point cloud after transformation; apply the calculated R and t to the source point cloud, if the transformation change is less than the preset threshold or the maximum number of iterations is reached, the algorithm terminates; otherwise, repeat the steps of finding corresponding point pairs, calculating the rotation matrix R and translation vector t, and applying them to the source point cloud; the final registration transformation matrix can accurately describe the spatial translation and rotation relationship between the preoperative CT and the network-generated CT, and the registration transformation matrix is ​​applied to the preoperative CT to complete the CT registration.

Citation Information

Patent Citations

  • A registration method and device for intraoperative X-ray and CT images

    CN118037793B

  • Image registration method and device, terminal equipment and storage medium

    CN113628260A

  • Medical low-dose CT image restoration method based on frequency domain Transform

    CN117541500A