Single perspective image and CT image registration method based on multi-mode joint optimization
Through a multimodal joint optimization method, combined with depth camera point cloud and preoperative three-dimensional images, the problem of insufficient accuracy and robustness in the registration of single fluoroscopic images with CT images was solved, and precise registration of single fluoroscopic images was achieved, saving surgical time.
Patent Information
- Application Number
- CN202510877057.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-26
AI Technical Summary
Existing technologies have problems with accuracy and robustness in the registration of single fluoroscopic images with CT images. Multiple fluoroscopic images are usually required to improve the registration effect, which prolongs the operation time.
A method based on multimodal joint optimization is adopted, combining depth camera point cloud data and preoperative three-dimensional images. By constructing a posture transformation matrix and a gradient descent optimization algorithm, accurate registration of a single fluoroscopic image and a CT image is achieved, avoiding the need to take multiple fluoroscopic images.
The accuracy and robustness of registration are improved, the operation time is reduced, the patient's radiation exposure is avoided, and the registration efficiency is improved.
Smart Images

Figure CN120765702A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image registration methods, and in particular relates to a single fluoroscopic image and CT image registration method based on multimodal joint optimization. Background Art
[0002] The surgical navigation system based on preoperative CT images and intraoperative X-ray fluoroscopy images uses the registration of X-ray images and CT images to map the preoperative surgical plan to the intraoperative stage, thereby achieving automatic surgical navigation and path guidance, which can greatly improve surgical accuracy and save surgical time.
[0003] The registration of X-ray fluoroscopy images and CT images usually adopts an iterative optimization method based on digitally reconstructed radiographs (DRR). This technology generates DRR from CT three-dimensional volume data through a ray casting algorithm, establishes an objective function based on grayscale similarity, and uses an optimization algorithm to iteratively adjust the spatial transformation parameters (including translation, rotation, and scaling degrees of freedom) to maximize the objective function value. In order to solve the problem of insufficient parameter constraints caused by the limited spatial information provided by single-view X-ray images, orthogonal dual-view (anteroposterior / lateral) imaging mode is often used in clinical practice. By establishing multi-plane constraint equations, the accuracy of three-dimensional spatial registration is effectively improved.
[0004] However, traditional methods have problems with accuracy and robustness in the registration of single fluoroscopic images and CT images. It is usually necessary to take multiple fluoroscopic images to improve the registration effect, which not only increases the registration time but also prolongs the operation time. Therefore, a single fluoroscopic image and CT image registration method based on multimodal joint optimization is designed to solve the above problems. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a method for the registration of a single fluoroscopic image and a CT image based on multimodal joint optimization to solve the problems of insufficient accuracy and robustness when using a single fluoroscopic image and a CT image for registration, which avoids the need to take multiple fluoroscopic images, improves the registration efficiency, and saves surgical time.
[0006] The technical solutions of the present invention are as follows:
[0007] A single fluoroscopic image and CT image registration method based on multimodal joint optimization includes the following steps:
[0008] S1. Scan the target area of the patient before surgery to obtain a preoperative three-dimensional image, and pre-process the image to obtain a first processed image;
[0009] S2. Intraoperatively scan the target area of the patient to obtain an intraoperative two-dimensional perspective image and depth camera point cloud data, and preprocess the intraoperative two-dimensional perspective image to obtain a second processed image;
[0010] S3. Obtaining a posture transformation matrix between the preoperative 3D image and the intraoperative 2D fluoroscopic image based on a positional relationship of the patient relative to the 2D fluoroscopic imaging device, transforming the first processed image, and then generating a digitally reconstructed radiographic image using a ray casting algorithm;
[0011] S4, calculating a normalized cross-correlation coefficient between the digitally reconstructed radiographic image and the second processed image, and converting the normalized cross-correlation coefficient into a first function that can be minimized;
[0012] S5. Performing threshold segmentation on the preoperative three-dimensional image to obtain a skin region image, thereby acquiring single-layer skin surface point cloud data;
[0013] S6. transforming the depth camera point cloud using the pose transformation matrix, and constructing a second function that can minimize the depth camera point cloud data and the single-layer skin surface point cloud data;
[0014] S7. Construct an objective function based on the first function and the second function, and use a gradient descent optimization algorithm to minimize the objective function, successively update the posture transformation matrix until the set conditions are met, obtain the minimized objective function value, and finally obtain the target posture transformation matrix of the three-dimensional image and the two-dimensional perspective image.
[0015] Furthermore, in step S5, the preoperative three-dimensional image is threshold segmented to obtain a skin region image, and then single-layer skin surface point cloud data is obtained, including the steps of:
[0016] S51, segmenting the preoperative 3D image based on a set threshold to obtain the skin region image, performing a morphological closing operation on the segmented image to obtain a closed skin region image, and performing denoising on the closed skin region image using a mean filtering method;
[0017] S52. Using an edge detection operator to perform a convolution operation on the denoised skin area image to obtain an initial edge image of the skin area, and then performing non-maximum suppression on the image to extract single-layer skin surface point cloud data.
[0018] Furthermore, the first processed image in step S1 is an image obtained by magnifying the original voxel values of the bone region in the pre-operative three-dimensional image by a set multiple.
[0019] Furthermore, the pre-processing in step S2 includes the following steps:
[0020] S21, performing logarithmic mapping on the intraoperative two-dimensional fluoroscopic image after smoothing to obtain a logarithmic mapping result, and subtracting the logarithmic mapping result from the logarithmic value of the initial energy to generate a negative image;
[0021] S22, linearly normalize the negative image to a set range, and adjust to a set size through bilinear interpolation to obtain a second processed image.
[0022] Further, after obtaining the depth camera point cloud through the step S2 and the single-layer skin surface point cloud through the step S5, the two are respectively subjected to down-sampling processing to obtain a down-sampled depth camera point cloud and a down-sampled skin surface point cloud.
[0023] Further, the step S4 specifically comprises the following steps:
[0024] S41, calculate the normalized cross-correlation coefficient of the digital reconstructed radiographic image and the second processed image, and the calculation formula is:
[0025]
[0026] wherein, and are the mean values of the pixel values of the second processed image and the digital reconstructed radiographic image respectively, <, > is the inner product, |·| is the two-norm of the vector, and the numerical range of NCC is [0, 1];
[0027] S42, construct a first function that can be minimized according to the normalized cross-correlation coefficient in the step S41, and the specific construction is as follows:
[0028] L1=1-NCC.
[0029] Further, the second function that can be minimized in the step S6 is specifically a bidirectional nearest neighbor distance measurement function between the depth camera point cloud data and the single-layer skin surface point cloud data, and the calculation formula is as follows:
[0030]
[0031] wherein P={p i}, Q={q i}, for each point in P, find the nearest point in Q, calculate the distance square and take the average, and repeat the same operation for each point in Q, and the bidirectional nearest neighbor distance is the sum of the two.
[0032] Further, in the step S7, the target function constructed based on the first function and the second function is specifically as follows:
[0033] Loss=α L1+β L2
[0034] Loss=α(1-NCC(I Xray ,I DRR ))+βD charmfer (P,Q)
[0035] Wherein, alpha and beta are weight coefficients of the balance function, alpha is 1, and beta is 0.001.
[0036] Further, the threshold value set in the step S51 is selected according to the HU value of the skin, and the value range is [-50, 100].
[0037] Further, the value range of the set multiple is [2.5, 4.5].
[0038] Further, the set range in the step S22 is [0, 1], and the set size is 128*128 or 64*64.
[0039] Compared with the prior art, the present application has the following advantages:
[0040] 1. In the C-arm X-ray machine equipped with a depth camera, the present application introduces X-ray fluoroscopy images and depth camera point cloud data, and solves the problem of easily falling into a local optimal solution caused by registration using only preoperative three-dimensional images and intraoperative two-dimensional fluoroscopy image data by adding depth camera point cloud in the registration algorithm and extracting single-layer skin surface point cloud in the preoperative three-dimensional image, thereby improving the success rate of registration.
[0041] 2. The present application constructs a first function and a second function related to the pose transformation matrix of the preoperative three-dimensional image and the intraoperative two-dimensional fluoroscopy image, calculates a target function composed of the first function and the second function by using a gradient optimization descent algorithm, finally makes the values of the first function, the second function and the target function reach a minimum value, and the pose transformation matrix is no longer updated, so that the optimal solution of the target pose transformation matrix is finally obtained, thereby ensuring the accuracy of the registration of the preoperative three-dimensional image and the intraoperative two-dimensional fluoroscopy image.
[0042] 3. The present application only needs to register a single fluoroscopy image obtained during surgery with a CT image obtained before surgery, solves the problem of long surgery time caused by the need to take multiple fluoroscopy images in the traditional method, reduces the exposure time of the patient to radiation during surgery, optimizes the registration efficiency while ensuring the accuracy and robustness of the registration.
[0043] In summary, the present application solves the problem of insufficient accuracy and robustness of registration using a single fluoroscopy image and a CT image, avoids the need to take multiple fluoroscopy images, improves the registration efficiency, and saves surgery time. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 is a flowchart of the single fluoroscopy image and CT image registration method based on multi-modal joint optimization of the present application;
[0045] Figure 21 is an example of a human body model image diagram of the present invention, wherein (a) is a CT image diagram, (b) is a point cloud diagram of the human body model skin surface extracted using a skin estimation algorithm, and (c) is a point cloud diagram of a depth camera installed on a C-arm machine;
[0046] Figure 3 These are images of the lumbar spine and sacral region of the present invention, wherein (a) is an X-ray fluoroscopic image, (b) is an X-ray fluoroscopic image after 2D pre-processing, and (c) is a DRR image generated using the transformation relationship T after registration. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0048] like Figures 1 to 3 As shown, the present invention proposes a single fluoroscopic image and CT image registration method based on multimodal joint optimization, comprising the steps of:
[0049] S1. Scan the target area of the patient before surgery to obtain a preoperative three-dimensional image, and pre-process the image to obtain a first processed image;
[0050] Specifically, the original voxel value of the bone region in the pre-operative three-dimensional image (ie, the HU value of the bone region) is magnified by a set multiple to obtain a first processed image.
[0051] In the present invention, voxel points with HU values greater than 800 in the preoperative three-dimensional image are multiplied by a set multiple.
[0052] In a specific embodiment of the present invention, the value range of the multiple is set to [2.5, 4.5], preferably 3.5.
[0053] In the present invention, 3D pre-processing, i.e., remapping of voxel values, is performed to improve the grayscale value of the bone region in the digitally reconstructed radiographic image subsequently generated based on the first processed image, thereby strengthening the bone structure features and enhancing the discrimination of image registration.
[0054] In the present invention, the preoperative three-dimensional image is a CT image obtained by scanning the target area of the patient with a CT device before the operation, such as Figure 2 As shown in (a).
[0055] In a specific embodiment of the present invention, the target area of the patient is the lumbar spine and sacrum area of the patient;
[0056] S2. Scan the target area of the patient during surgery to obtain an intraoperative two-dimensional perspective image and depth camera point cloud data, and pre-process the intraoperative two-dimensional perspective image to obtain a second processed image;
[0057] The present invention uses a C-arm X-ray machine equipped with a depth camera (hereinafter referred to as the C-arm machine) to obtain an intraoperative two-dimensional fluoroscopic image of the patient's target area and depth camera point cloud data, and preprocesses the intraoperative two-dimensional fluoroscopic image to obtain a second processed image; the method specifically includes the following steps:
[0058] S21, smoothing the intraoperative two-dimensional fluoroscopic image to suppress noise interference and reduce image detail differences, thereby improving the robustness of subsequent similarity calculations; then performing logarithmic mapping to obtain a logarithmic mapping result, and subtracting the logarithmic mapping result from the logarithmic value of the initial energy to generate a negative image,
[0059] In the present invention, a Gaussian filter with a kernel size of 5 and a standard deviation of 1.0 is selected to smooth the intraoperative two-dimensional images.
[0060] In the present invention, logarithmic mapping is performed to obtain a logarithmic mapping result, and the logarithmic mapping result is subtracted from the logarithmic value of the initial energy to generate a negative image. Specifically, the corresponding logarithmic value is calculated for each pixel value in the intraoperative two-dimensional fluoroscopic image to obtain a logarithmic mapping result to enhance the contrast of low grayscale areas (bone areas); the logarithmic value of the initial energy of the X-ray is 65536 minus the logarithmic value corresponding to each pixel value in the intraoperative two-dimensional fluoroscopic image obtained above to generate a negative image, so that the grayscale distribution of the image is consistent with the subsequent digitally reconstructed radiological image.
[0061] S22, linearly normalizing the negative film image obtained in step S21 to a set range, and adjusting it to a set size through bilinear interpolation to obtain a second processed image;
[0062] In a specific embodiment of the present invention, the setting range of the linear normalization of the negative image is [0, 1], and the setting size is 128*128 or 64*64, in order to ensure consistency with the subsequent digitally reconstructed radiographic image (DRR) generated based on the preoperative three-dimensional image, unify the image resolution, reduce the amount of calculation and adapt to the subsequent similarity matching step.
[0063] S3. Obtaining a posture transformation matrix between the preoperative three-dimensional image and the intraoperative two-dimensional perspective image based on the positional relationship of the patient relative to the two-dimensional perspective imaging device to transform the first processed image, and then generating a digitally reconstructed radiographic image through a ray casting algorithm.
[0064] The two-dimensional perspective imaging device in the present invention is a C-arm X-ray machine (hereinafter referred to as C-arm machine) equipped with a depth camera.
[0065] Specifically, the position relationship of the patient relative to the C-arm machine includes:
[0066] (1) The patient lies prone with the C-arm on the patient's right side;
[0067] (2) The patient lies prone with the C-arm on the patient's left side;
[0068] (3) The patient lies supine with the C-arm on the patient's right side;
[0069] (4) The patient lies supine with the C-arm on the patient's left side;
[0070] In this embodiment, the initial value of the posture transformation matrix between the preoperative 3D image and the intraoperative 2D fluoroscopic image in step S3 is determined based on the relative position relationship of the patient's C-arm.
[0071] In a specific embodiment of the present invention, taking the patient lying prone and the C-arm on the right side of the patient during fluoroscopy, the initial value of the pose transformation matrix of the intraoperative 2D fluoroscopy image relative to the preoperative 3D image is:
[0072]
[0073] Where (cx, cy, cz) are the coordinates of the center point of the CT image, and the initial value of the above-mentioned posture transformation matrix can be converted into the form of T0 = (-1.209, 1.209, -1.209, cx, cy, cz), where the first three elements are rotation vectors and the last three elements are translation vectors.
[0074] Similarly, when the patient is lying prone and the C-arm is on the patient's left side, the initial value of the posture transformation matrix of the intraoperative 2D fluoroscopic image relative to the preoperative 3D image is T0 = (1.209, -1.209, -1.209, cx, cy, cz);
[0075] When the patient is lying supine and the C-arm is on the right side of the patient, the initial value of the posture transformation matrix of the intraoperative 2D fluoroscopic image relative to the preoperative 3D image is T0 = (1.209, 1.209, 1.209, cx, cy, cz);
[0076] When the patient lies supine with the C-arm on the patient's left side, the initial value of the posture transformation matrix of the intraoperative two-dimensional fluoroscopic image relative to the preoperative three-dimensional image is T0 = (-1.209, -1.209, 1.209, cx, cy, cz).
[0077] In the present invention, since the digitally reconstructed radiographic image to be generated is a two-dimensional image, it is necessary to convert the preoperative three-dimensional image from the CT image coordinate system to the C-arm coordinate system of the intraoperative two-dimensional fluoroscopic image. Therefore, based on the current position of the patient relative to the C-arm, the initial value of the transformation relationship between the C-arm coordinate system and the CT image coordinate system is obtained, that is, the preoperative three-dimensional image is transformed into the C-arm coordinate system. The preoperative three-dimensional image is preprocessed to obtain a first processed image, and the first processed image is then used to generate a digitally reconstructed radiographic image through a ray projection algorithm.
[0078] In the present invention, the digitally reconstructed radiographic image generated above is linearly normalized to the range of [0, 1] to ensure consistency with the second processed image obtained in step S2.
[0079] S4, calculating the normalized cross-correlation coefficient between the digitally reconstructed radiographic image and the second processed image, and converting the normalized cross-correlation coefficient into a first function that can be minimized, specifically comprising:
[0080] In the present invention, the similarity between the digitally reconstructed radiographic image and the image affected by the second processing is calculated. The similarity can be calculated using mutual information (MI) or normalized cross correlation (NCC).
[0081] In this invention, the normalized cross correlation (NCC) is used as the similarity measurement indicator. The normalized cross correlation coefficient calculation formula is as follows:
[0082]
[0083] in, and are the mean pixel values of the second processed image and the digitally reconstructed radiographic image, respectively. <, > are the inner products. |·| is the bi-norm of the vector. The value range of NCC is [0, 1].
[0084] The first function that can be minimized is constructed based on the normalized cross-correlation coefficient obtained in the above steps, as follows:
[0085] L1 = 1-NCC;
[0086] In the present invention, in order to adapt to the gradient descent optimization framework, the normalized correlation coefficient needs to be converted into a minimizable objective function, where the NCC value range is [0, 1], and is converted into a minimizable objective through 1-NCC, so that the result of maximizing the similarity is equivalent to the minimized first function value.
[0087] In the present invention, the closer the value of NCC is to 1, the more similar the digitally reconstructed radiographic image and the second processed image are, and the smaller the value of the first function L1 = 1-NCC is, the closer it is to 0.
[0088] In the present application, the first function is a function of the second processed image and the digital reconstructed radiograph, which is generated by the first processed image after being transformed by the initial value T0 of the pose transformation matrix, so the first function is a function of the pose transformation matrix T.
[0089] S5, threshold segmentation is performed on the preoperative three-dimensional image to obtain a skin region image, and then single-layer skin surface point cloud data is obtained;
[0090] Specifically, it includes the following steps:
[0091] S51, based on the set threshold, the preoperative three-dimensional image is segmented to obtain a skin region image, morphological closing operation is performed on the skin region image to obtain a closed skin region image, and mean filtering method is used to denoise the closed skin region image;
[0092] In the present application, the skin region is segmented according to the HU value
-50, 100
[0093] S52, an edge detection operator is used to perform convolution operation on the denoised skin region image to obtain an initial edge image of the skin region, and then non-maximum suppression is performed to extract single-layer skin surface point cloud data.
[0094] In the present application, 3D-edge detection operator is used to perform convolution operation on the denoised skin region image to extract edge information and obtain an initial edge image of the skin region.
[0095] In the present application, non-maximum suppression is used to thin the initial edge image of the skin region to single-layer voxel thickness, and then single-layer skin surface point cloud is obtained.
[0096] In the present application, after obtaining the single-layer skin surface point cloud through step S5, the depth camera point cloud obtained in step S2 and the single-layer skin surface point cloud obtained in step S5 can be respectively subjected to down-sampling processing to obtain down-sampled depth camera point cloud and down-sampled skin surface point cloud, which helps to retain key features of the point cloud while reducing data volume and subsequent calculation amount.
[0097] In the present embodiment, the down-sampling methods that can be used include but are not limited to uniform sampling, random sampling and voxel grid-based sampling.
[0098] S6, the depth camera point cloud is transformed through the pose transformation matrix, and a second function that can be minimized by the depth camera point cloud data and the single-layer skin surface point cloud data is constructed;
[0099] In this embodiment, the second function that can be minimized in step S6 is specifically a bidirectional nearest neighbor distance measurement function constructed between the depth camera point cloud data and the single layer skin surface point cloud data;
[0100] Specifically, in order to construct a differentiable bidirectional nearest neighbor distance (hamfer) metric function to realize the optimization process based on gradient descent, it is necessary to matrix-represent the depth camera point cloud data and the single-layer skin surface point cloud data; wherein, the single-layer skin surface point cloud data is defined as the matrix The hollow R represents a set of three-dimensional space points, and m represents the number of three-dimensional points contained in the single-layer skin surface point cloud data; the depth camera point cloud data is defined as a matrix n represents the 3D point data of the depth camera point cloud data. Based on this tensor representation, a differentiable bidirectional nearest neighbor distance metric function is established, ultimately achieving accurate registration of the depth camera point cloud and the single-layer skin point cloud.
[0101] The specific calculation formula of the bidirectional nearest neighbor distance metric function is as follows:
[0102]
[0103] Where P = {p i}, Q={q i For each point in P, find the nearest point in Q, calculate the square of the distance and take the average. For each point in Q, find the nearest point in P, calculate the square of the distance and take the average. At this time, the two are added to obtain the minimum value of the second function, that is, the bidirectional nearest neighbor distance between the single-layer skin point cloud data and the depth camera point cloud.
[0104] In the present invention, to more conveniently calculate the bidirectional nearest neighbor distance between the depth camera point cloud and the single-layer skin point cloud, the depth camera point cloud needs to be converted from the C-arm coordinate system to the CT image coordinate system. Therefore, the depth camera point cloud is transformed to the CT image coordinate system using the initial value of the pose transformation matrix between the C-arm coordinate system and the CT image coordinate system obtained in step S2. Therefore, the second function is a function of the pose transformation matrix T.
[0105] S7. Construct an objective function based on the first function and the second function, and use a gradient descent optimization algorithm to minimize the objective function, successively update the posture transformation matrix until the set conditions are met, obtain the minimized objective function value, and finally obtain the target posture transformation matrix of the three-dimensional image and the two-dimensional perspective image.
[0106] In the present invention, the objective function is constructed based on the first function and the second function, as follows:
[0107] Loss=α L1+β L2
[0108] Loss = α(1-NCC(I Xray , I DRR ))+βD Charmfer (P, Q)
[0109] Among them, α and β are weight coefficients set to balance the first function and the second function. Since the range of the NCC value is [0, 1] and the output range of the bidirectional nearest neighbor distance metric function is large, β is usually smaller than α. In this embodiment, α is set to 1 and β is set to 0.001.
[0110] In the present invention, since the first function and the second function are both functions of the posture transformation matrix T, the objective function is also a function of the posture transformation matrix T. The gradient descent optimization algorithm is used to optimize the objective function, which is actually a continuous optimization of the posture transformation matrix T. The posture transformation matrix is updated once each optimization is completed until the set conditions are met and the minimized objective function value is obtained. Finally, the target posture transformation matrix of the preoperative three-dimensional image and the intraoperative two-dimensional perspective image is obtained, and the alignment of the two is completed.
[0111] In the present invention, the Adam optimizer is used to perform gradient descent optimization on the objective function. The process of gradient descent optimization is described as follows:
[0112] a. Configure the optimizer parameters: learning rate η = 0.01, the step size of the rotation component of the pose transformation matrix T is 0.005, the step size of the translation component is set to 10, the step size is reduced every 50 iterations, and the decay coefficient is 0.5. Momentum β1 = 0.9, β2 = 0.999;
[0113] b. Set the initial value T0 of the posture transformation matrix T according to the position relationship of the patient relative to the C-arm machine;
[0114] c. Set the convergence threshold or maximum number of iterations;
[0115] d. Calculate the objective function Loss and find the gradient (i.e., derivative) of the pose transformation matrix T, then update the pose transformation matrix T, use the updated pose transformation matrix T to transform the first processed image and generate a corresponding digitally reconstructed radiographic image, and then repeat step S4 to calculate the normalized correlation coefficient between the updated digitally reconstructed radiographic image and the second processed image, so that the first function is updated; use the above-mentioned updated pose transformation matrix T to correspondingly transform the depth camera point cloud, repeat step S6 to calculate the bidirectional nearest neighbor distance metric function between the single-layer skin surface point cloud data and the depth camera point cloud data, the second function is updated, and then the objective function is updated.
[0116] It is determined whether the set conditions are met, and until the set conditions are met, the minimized objective function value is obtained, and finally the target posture transformation matrix of the preoperative three-dimensional image and the intraoperative two-dimensional perspective image is obtained.
[0117] The specific setting conditions are: comparing whether the norm of the gradient of the objective function Loss with respect to the pose transformation matrix T is less than the convergence threshold, or whether the maximum number of iterations has been reached. If the setting conditions are not met, run step d again.
[0118] In the present invention, the convergence threshold is set to 0.01, and the maximum number of iterations is set to 100 or 200.
[0119] In the present invention, as the posture transformation matrix T is updated successively, it gets closer and closer to the optimal solution, the generated digitally reconstructed radiographic image becomes more and more similar to the second processed image (that is, the value of the first function L1 is closer to 0), and the depth camera point cloud becomes closer and closer to the single-layer skin surface point cloud (that is, the value of the second function L2 is closer to 0), that is, the value of the objective function Loss is also closer to 0, until the optimal solution is reached, and a posture transformation matrix that is no longer updated is obtained, that is, the target posture transformation matrix of the preoperative three-dimensional image and the intraoperative two-dimensional perspective image is obtained. At this time, the precise alignment of the two is completed through the target transformation matrix.
[0120] In the present invention, completing the registration through a single fluoroscopic image can save surgical time and avoid taking multiple fluoroscopic images.
[0121] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A single fluoroscopic image and CT image registration method based on multimodal joint optimization, characterized by: Including steps: S1. Scan the target area of the patient before surgery to obtain a preoperative three-dimensional image, and pre-process the image to obtain a first processed image; S2. Scanning the target area of the patient during surgery to obtain an intraoperative two-dimensional perspective image and depth camera point cloud data, and preprocessing the intraoperative two-dimensional perspective image to obtain a second processed image; S3. Obtaining a posture transformation matrix between the preoperative 3D image and the intraoperative 2D fluoroscopic image based on a positional relationship of the patient relative to the 2D fluoroscopic imaging device, transforming the first processed image, and then generating a digitally reconstructed radiographic image using a ray casting algorithm; S4, calculating a normalized cross-correlation coefficient between the digitally reconstructed radiographic image and the second processed image, and converting the normalized cross-correlation coefficient into a first function that can be minimized; S5. Performing threshold segmentation on the preoperative three-dimensional image to obtain a skin region image, thereby acquiring single-layer skin surface point cloud data; S6. transforming the depth camera point cloud using the pose transformation matrix, and constructing a second function that can minimize the depth camera point cloud data and the single-layer skin surface point cloud data; S7. Construct an objective function based on the first function and the second function, and use a gradient descent optimization algorithm to minimize the objective function, successively update the posture transformation matrix until the set conditions are met, obtain the minimized objective function value, and finally obtain the target posture transformation matrix of the preoperative three-dimensional image and the intraoperative two-dimensional perspective image.
2. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 1, characterized in that: In step S5, threshold segmentation is performed on the preoperative three-dimensional image to obtain a skin region image, and then single-layer skin surface point cloud data is obtained, including the steps of: S51, segmenting the preoperative 3D image based on a set threshold to obtain the skin region image, performing a morphological closing operation on the segmented image to obtain a closed skin region image, and performing denoising on the closed skin region image using a mean filtering method; S52. Using an edge detection operator to perform a convolution operation on the denoised skin area image to obtain an initial edge image of the skin area, and then performing non-maximum suppression on the image to extract single-layer skin surface point cloud data.
3. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 1, characterized in that: The first processed image in step S1 is an image obtained by magnifying the original voxel values of the bone region in the pre-operative three-dimensional image by a set multiple.
4. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 1, characterized in that: The pre-processing in step S2 comprises the following steps: S21, performing logarithmic mapping on the intraoperative two-dimensional fluoroscopic image after smoothing to obtain a logarithmic mapping result, and subtracting the logarithmic mapping result from the logarithmic value of the initial energy to generate a negative image; S22: Linearly normalize the negative image to a set range, and adjust it to a set size through bilinear interpolation to obtain a second processed image.
5. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 1, characterized in that: After obtaining the depth camera point cloud through step S2 and the single-layer skin surface point cloud through step S5, downsampling processing is performed on the two to obtain a downsampled depth camera point cloud and a downsampled skin surface point cloud.
6. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 1, characterized in that: The step S4 specifically includes the following steps: S41, calculating the normalized correlation coefficient between the digitally reconstructed radiographic image and the second processed image, using the following formula: in, and are the mean pixel values of the second processed image and the digitally reconstructed radiographic image, respectively; <, > are inner products; |·| is the bi-norm of the vector; and the value range of NCC is [0, 1]; S42: Construct a first minimizable function based on the normalized cross-correlation coefficient in step S41, as follows: L1=1-NCC.
7. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 1, characterized in that: The second function that can be minimized in step S6 is specifically a bidirectional nearest neighbor distance measurement function between the depth camera point cloud data and the single-layer skin surface point cloud data, and the calculation formula is as follows: Where P = {p i }, Q={q i ), for each point in P, find the nearest point in Q, calculate the square of the distance and take the average, repeat the same operation for each point in Q, and the bidirectional nearest neighbor distance is the sum of these two.
8. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 1, characterized in that: In step S7, the objective function constructed based on the first function and the second function is as follows: Loss=αL1+βL2 Loss=α(1-NCC(I Xray ,I DRR ))+βD Charmfer (P,Q) Among them, α and β are the weight coefficients of the balance function, α is 1, and β is 0.
001.
9. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 2, wherein: The threshold value set in step S51 is selected according to the HU value of the skin, and the value range is [-50, 100].
10. The method for registering a single fluoroscopic image with a CT image based on multimodal joint optimization according to claim 3, characterized in that: The value range of the set multiple is [2.5, 4.5].
Citation Information
Patent Citations
Method and system for precise registration and coincidence of preoperative CT or nuclear magnetic image and corresponding focus in operation, storage medium, and equipment
CN112907642A
Multi-modal medical image three-dimensional point cloud registration optimization method
CN114119549A
Preoperative three-dimensional CT registration method based on single X-ray image
CN119693424A
Systems and methods for real-time multiple modality image alignment
US20220285009A1