An unsupervised registration method and system for pelvic reduction surgery navigation

Through the unsupervised registration method, the two-dimensional X-ray images are converted into three-dimensional CT images, and the image segmentation is performed using the MedSAM segmentation model, which solves the problem of insufficient accuracy and generalization ability of two-dimensional and three-dimensional image registration in pelvic reduction surgical navigation, achieving higher accuracy and wider posture prediction, and improving the efficiency of surgical navigation.

CN119515942BActive Publication Date: 2025-05-23BEIJING DADING FRONTIER MEDICAL TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510096083.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-23
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

In the prior art In the pelvic reduction surgical navigation, the registration of two-dimensional X-ray images and three-dimensional CT images has problems with insufficient accuracy and generalization ability, especially in the Z-axis direction, the prediction accuracy is poor, and the spatial information of preoperative CT cannot be fully utilized.

Method used

Unsupervised registration method is adopted to convert the two-dimensional X-ray image into three-dimensional CT image by reconstruction model, and image segmentation is performed using the MedSAM segmentation model, key points are determined and rigid transformation matrix is ​​calculated to achieve unsupervised registration in 2D/3D.

Benefits of technology

It improves registration accuracy and generalizes the ability. It can predict the position and posture of all patients without training the network for each patient. The prediction range of position and posture is wider, which enhances the speed and convenience of surgical navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119515942B_ABST
    Figure CN119515942B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised registration method and system for pelvic reduction surgical navigation, the method comprising: constructing a three-dimensional reconstruction network model, and completing training based on a generative adversarial loss function; inputting intraoperative two-dimensional X-ray images into the three-dimensional reconstruction network model to obtain a three-dimensional intraoperative CT image; segmenting the preoperative CT image and the intraoperative CT image using the MedSAM segmentation model to obtain preoperative and intraoperative segmented images of the pelvic region; determining the key points of the preoperative segmented image and the intraoperative segmented image by a convolutional neural network, and minimizing the distance of the key points according to a closed-form function to derive a rigid transformation matrix; and registering the intraoperative CT image according to the rigid transformation matrix and the preoperative CT image. Through the technical solution of the present invention, the registration accuracy and generalization ability are improved, and the prediction range of the posture and attitude is wider, providing good technical support for the quickness and convenience of surgical navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of surgical navigation systems, and in particular to an unsupervised registration method for pelvic reduction surgical navigation and an unsupervised registration system for pelvic reduction surgical navigation. Background Art

[0002] The pelvis is one of the most important structures in the human skeleton, connecting the spine, hip joints and lower limbs, and is crucial to the body's stability and motor function. Pelvic fractures are common injuries in trauma orthopedics, and their injury mechanisms are complex. The pelvis is rich in important soft tissues such as blood vessels and nerves. Once a serious pelvic fracture occurs, it will be life-threatening if the correct treatment plan is not selected in time.

[0003] Repairing bones requires a series of steps such as fracture reduction and fixation to restore the normal anatomical structure and function of the pelvis. Before fracture reduction, the doctor will calculate the parameters of fracture reduction in the preoperative planning step, including the translation size and angle required for the fracture reduction process. Pelvic reduction surgery is relatively complicated. In order to ensure the success rate of the operation, it is necessary to obtain accurate pelvic position and pelvic posture. Appropriate medical images can provide doctors with a lot of effective information during the operation. Surgical navigation can combine preoperative and intraoperative information to assist doctors in surgery. Fracture reduction surgery registration can assist doctors in judging the progress of surgical reduction. It combines preoperative and intraoperative images to obtain the spatial position and parameters of the current fracture reduction, and compares them with those obtained in preoperative planning, so as to prepare for the next surgical plan.

[0004] In clinical practice, high-resolution three-dimensional CT images are usually obtained before surgery to determine the patient's preoperative position and posture (posture). However, due to the high radiation and long imaging time when taking CT images, in order to obtain the patient's posture during surgery, two-dimensional X-ray images with shorter imaging time and less radiation are usually selected during surgery. By comparing the patient's preoperative and intraoperative medical images, the patient's posture information can be reflected in real time.

[0005] Surgical navigation is an automated tool that assists doctors in surgery. It provides real-time navigation of the patient's pelvis and surgical instruments, and assists doctors in completing surgery by combining the patient's fracture reduction situation before and during surgery. It requires registration of preoperative CT with intraoperative X-rays. Surgical navigation can unify preoperative 3D data and intraoperative 2D data into the same coordinate system, indirectly providing doctors with real-time surgical posture information of the patient. Due to the different modalities of data obtained before and during surgery, 3D CT is usually used to obtain DRR images using Digitally Reconstructed Radiograph technology to simulate the formation process of X-rays, so as to perform image registration of DRR images and X-rays.

[0006] The key technology of surgical navigation is to complete the rigid registration from 2D to 3D by combining preoperative and intraoperative images. Registration usually refers to the process of aligning multiple data or images to achieve consistency or comparability. This concept is widely used in many fields, including computer vision, medical imaging, remote sensing image processing, etc. In the field of medical imaging, image registration is used to align medical images acquired at different times, different modalities, or different imaging devices so that doctors can better observe and compare. The object will not deform during the rigid registration process. The accuracy of registration directly affects the success rate of the operation, and the patient's posture before and during the operation needs to be combined in real time to assist the doctor in completing the operation. With the continuous development of artificial intelligence, 2D / 3D registration methods based on deep learning have gradually replaced traditional algorithms. Many researchers have done a lot of research on the application of machine learning to medical image registration and proposed many practical methods.

[0007] Most of the learning-based methods in the prior art understand the prediction of spatial posture as a regression problem, that is, for an image, the 6-DOF posture parameters are output. Although this method has real-time performance during the reasoning process, it requires training the network for each patient and does not have strong generalization. In addition, the existing learning-based methods all have a known registration posture transformation range and use it to train the regression network, understanding the registration as a white box problem, and cannot accurately learn the posture that is not in the spatial range. From the perspective of solving the registration problem, most of the existing methods understand 2D-3D registration as 2D-2D registration, so there is a lack of information at the depth level, that is, the prediction accuracy in the Z-axis direction is poor, and the spatial information of preoperative CT cannot be fully utilized. Most pelvic CT images have interference information such as air and soft tissue, which interferes with the extraction of pelvic information. Summary of the invention

[0008] In response to the above problems, the present invention provides an unsupervised registration method and system for pelvic reduction surgical navigation. The method and system can fully utilize the depth information of the image through the reconstruction model to convert the two-dimensional X-ray image into a three-dimensional CT image, and can convert the 2D / 3D problem into a 3D / 3D registration problem, which serves as a bridge for the subsequent registration process. The MedSAM segmentation model is used for image segmentation, which is used as data expansion and segmentation network to improve the registration accuracy. The unsupervised registration framework effectively combines the key information of preoperative CT and has generalization ability. It can predict the position and posture of all patients without training the network for each patient, and the prediction range of the position and posture is wider, which provides good technical support for the speed and convenience of surgical navigation.

[0009] To achieve the above object, the present invention provides an unsupervised registration method for pelvic reduction surgery navigation, comprising:

[0010] The 3D reconstruction network model is constructed using the encoding module, cross-dimensional connection module and decoding module, and the training is completed based on the generative adversarial loss function;

[0011] Inputting the intraoperative two-dimensional X-ray image into the trained three-dimensional reconstruction network model to obtain a phase-mapped three-dimensional intraoperative CT image;

[0012] The preoperative CT image and the intraoperative CT image are segmented using the MedSAM segmentation model to obtain a preoperative segmentation image and an intraoperative segmentation image of the pelvic region;

[0013] Determine the key points of the preoperative segmentation image and the intraoperative segmentation image by using a convolutional neural network, and perform minimization calculation on the distance of the key points according to a closed-form function to derive a rigid transformation matrix;

[0014] The intraoperative CT image is registered according to the rigid transformation matrix and the preoperative CT image.

[0015] In the above technical solution, preferably, the encoding module, the cross-dimensional connection module and the decoding module are used to construct a 3D reconstruction network model, and the training is completed based on the generative adversarial loss function. The specific process includes:

[0016] The encoder-decoder structure is used to build a 3D reconstruction network model, and the encoder and decoder are connected through a cross-dimensional connection module;

[0017] The encoder is composed of a preset number of dense blocks DenseBlock, the decoder is composed of a preset number of upsampling modules and a three-dimensional feature extraction module Basic3D, the cross-dimensional connection module includes two cross-dimensional connection modules Connection A and Connection B, the Connection A is used to map the two-dimensional features of the encoder to a three-dimensional feature space, and the Connection B is used to fuse the two-dimensional features and the three-dimensional features in the encoding and decoding stages;

[0018] The three-dimensional reconstruction network model is trained using two-dimensional X-ray images and corresponding three-dimensional CT images, and based on the generation of an adversarial loss function, the training is completed when the three-dimensional reconstruction network model can reach a preset loss value.

[0019] In the above technical solution, preferably, the preoperative CT image and the intraoperative CT image are segmented using the MedSAM segmentation model respectively to obtain a preoperative segmentation image and an intraoperative segmentation image of the pelvic region, and the specific process includes:

[0020] The MedSAM segmentation model is constructed by using an image encoder, a prompt encoder and a mask decoder, wherein the image encoder is used to extract the global features of the CT image and generate an image embedding, the prompt encoder is used to generate a corresponding feature vector according to the prompt box information input by the user, and the mask decoder is used to fuse the output features of the image encoder and the prompt encoder, and generate a segmentation result of the target area through a decoding operation;

[0021] The preoperative CT image and the intraoperative CT image are input into the MedSAM segmentation model to obtain a preoperative segmentation image mask and an intraoperative segmentation image mask for the pelvic region.

[0022] In the above technical solution, preferably, the key points of the preoperative segmentation image and the intraoperative segmentation image are determined by a convolutional neural network, and the distance of the key points is minimized according to a closed-form function to derive a rigid transformation matrix, and the specific process includes:

[0023] The preoperative segmentation image and the intraoperative segmentation image are input into the convolutional neural network Conv-COM, and after multi-layer convolution, normalization, ReLU activation and downsampling operations, the features are output to the centroid layer;

[0024] Calculating the centroids of a preset number of activation maps and converting the activation maps into coordinates of key points, wherein each key point corresponds to a specific location in the anatomical structure;

[0025] According to the key point set of the preoperative segmentation image and the key point set of the intraoperative segmentation image, a closed function that minimizes the key point distance is derived using derivatives to derive an optimal solution of the rigid transformation matrix.

[0026] In the above technical solution, preferably, the intraoperative CT image is registered according to the rigid transformation matrix and the preoperative CT image, and the specific process includes:

[0027] Based on the preoperative CT image, a registration transformation is performed using a rigid transformation matrix of a key point set to obtain a registration result image;

[0028] The registration effect of the registration result image is judged according to the DICE loss between the registration result image and the intraoperative segmentation image.

[0029] The present invention further proposes an unsupervised registration system for pelvic reduction surgery navigation, which uses the unsupervised registration method for pelvic reduction surgery navigation disclosed in any one of the above technical solutions, including:

[0030] A reconstruction model building module is used to build a 3D reconstruction network model using an encoding module, a cross-dimensional connection module, and a decoding module, and complete training based on a generative adversarial loss function;

[0031] A three-dimensional image mapping module, used for inputting the intraoperative two-dimensional X-ray image into the trained three-dimensional reconstruction network model to obtain a phase-mapped three-dimensional intraoperative CT image;

[0032] A pelvic image segmentation module, used to segment the preoperative CT image and the intraoperative CT image using the MedSAM segmentation model to obtain a preoperative segmentation image and an intraoperative segmentation image of the pelvic region;

[0033] A registration matrix derivation module, used for determining key points of the preoperative segmentation image and the intraoperative segmentation image through a convolutional neural network, and performing minimization calculation on the distance of the key points according to a closed-form function to derive a rigid transformation matrix;

[0034] The pelvic image registration module is used to register the intraoperative CT image according to the rigid transformation matrix and the preoperative CT image.

[0035] In the above technical solution, preferably, the reconstruction model building module is specifically used to:

[0036] The encoder-decoder structure is used to build a 3D reconstruction network model, and the encoder and decoder are connected through a cross-dimensional connection module;

[0037] The encoder is composed of a preset number of dense blocks DenseBlock, the decoder is composed of a preset number of upsampling modules and a three-dimensional feature extraction module Basic3D, the cross-dimensional connection module includes two cross-dimensional connection modules Connection A and Connection B, the Connection A is used to map the two-dimensional features of the encoder to a three-dimensional feature space, and the Connection B is used to fuse the two-dimensional features and the three-dimensional features in the encoding and decoding stages;

[0038] The three-dimensional reconstruction network model is trained using two-dimensional X-ray images and corresponding three-dimensional CT images, and based on the generation of an adversarial loss function, the training is completed when the three-dimensional reconstruction network model can reach a preset loss value.

[0039] In the above technical solution, preferably, the three-dimensional image mapping module is specifically used for:

[0040] The MedSAM segmentation model is constructed by using an image encoder, a prompt encoder and a mask decoder, wherein the image encoder is used to extract the global features of the CT image and generate an image embedding, the prompt encoder is used to generate a corresponding feature vector according to the prompt box information input by the user, and the mask decoder is used to fuse the output features of the image encoder and the prompt encoder, and generate a segmentation result of the target area through a decoding operation;

[0041] The preoperative CT image and the intraoperative CT image are input into the MedSAM segmentation model to obtain a preoperative segmentation image mask and an intraoperative segmentation image mask for the pelvic region.

[0042] In the above technical solution, preferably, the registration matrix derivation module is specifically used for:

[0043] The preoperative segmentation image and the intraoperative segmentation image are input into the convolutional neural network Conv-COM, and after multi-layer convolution, normalization, ReLU activation and downsampling operations, the features are output to the centroid layer;

[0044] Calculating the centroids of a preset number of activation maps and converting the activation maps into coordinates of key points, wherein each key point corresponds to a specific location in the anatomical structure;

[0045] According to the key point set of the preoperative segmentation image and the key point set of the intraoperative segmentation image, a closed function that minimizes the key point distance is derived using derivatives to derive an optimal solution of the rigid transformation matrix.

[0046] In the above technical solution, preferably, the pelvic image registration module is specifically used for:

[0047] Based on the preoperative CT image, a registration transformation is performed using a rigid transformation matrix of a key point set to obtain a registration result image;

[0048] The registration effect of the registration result image is judged according to the DICE loss between the registration result image and the intraoperative segmentation image.

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows: by reconstructing the model, the depth information of the image is fully utilized to convert the two-dimensional X-ray image into a three-dimensional CT image, and the 2D / 3D problem can be converted into a 3D / 3D registration problem, which plays a bridging role in the subsequent registration process. The MedSAM segmentation model is used for image segmentation, which is used as data expansion and segmentation network to improve the registration accuracy. The unsupervised registration framework effectively combines the key information of preoperative CT and has generalization ability. The position and posture of all patients can be predicted without training the network for each patient, and the prediction range of the position and posture is wider, which provides good technical support for the quickness and convenience of surgical navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 A schematic flow chart of an unsupervised registration method for pelvic reduction surgery navigation disclosed in an embodiment of the present invention;

[0051] Figure 2 A schematic diagram of the principle structure of a three-dimensional reconstruction network model disclosed in an embodiment of the present invention;

[0052] Figure 3 A schematic diagram of the principle structure of a cross-dimensional connection module disclosed in an embodiment of the present invention;

[0053] Figure 4 A schematic diagram of the principle structure of the MedSAM segmentation model disclosed in one embodiment of the present invention;

[0054] Figure 5 A schematic diagram of a process of unsupervised key point registration disclosed in an embodiment of the present invention;

[0055] Figure 6 A schematic diagram of the implementation result of reconstruction and re-projection disclosed in an embodiment of the present invention;

[0056] Figure 7 A schematic diagram of a segmentation result disclosed in an embodiment of the present invention;

[0057] Figure 8 A schematic diagram of key point registration results disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0059] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0060] like Figure 1 As shown, an unsupervised registration method for pelvic reduction surgery navigation provided by the present invention includes:

[0061] The 3D reconstruction network model is constructed using the encoding module, cross-dimensional connection module and decoding module, and the training is completed based on the generative adversarial loss function;

[0062] Input the intraoperative two-dimensional X-ray image into the trained three-dimensional reconstruction network model to obtain a phase-mapped three-dimensional intraoperative CT image;

[0063] The preoperative CT images and intraoperative CT images were segmented using the MedSAM segmentation model to obtain preoperative segmentation images and intraoperative segmentation images of the pelvic area.

[0064] The key points of the preoperative segmentation image and the intraoperative segmentation image are determined by a convolutional neural network, and the distance between the key points is minimized according to a closed-form function to derive the rigid transformation matrix.

[0065] The intraoperative CT images are registered based on the rigid transformation matrix and the preoperative CT images.

[0066] In this embodiment, the two-dimensional X-ray image is converted into a three-dimensional CT image by reconstructing the model to fully utilize the depth information of the image, and the 2D / 3D problem can be converted into a 3D / 3D registration problem, which serves as a bridge for the subsequent registration process. The MedSAM segmentation model is used for image segmentation, which is used as data expansion and segmentation network to improve the registration accuracy. The unsupervised registration framework effectively combines the key information of preoperative CT and has generalization ability. It can predict the position and posture of all patients without training the network for each patient, and the prediction range of the position and posture is wider, which provides good technical support for the speed and convenience of surgical navigation.

[0067] Specifically, the unsupervised registration method for pelvic reduction surgery navigation proposed in the present invention discloses an unsupervised 2D / 3D registration algorithm based on key points, which is mainly aimed at X-ray images and CT images in pelvic reduction surgery. It can directly predict the posture of two-dimensional DRR or X-ray. The algorithm understands the 2D / 3D registration problem as a 3D-3D registration problem.

[0068] During the implementation process, the three-dimensional reconstruction network model first realizes sparse perspective three-dimensional reconstruction of the acquired intraoperative 2D X-ray images. For the subsequent unsupervised key point registration, the reconstructed three-dimensional CT images are then segmented into the pelvic part using the latest medical MedSAM 3D segmentation technology. Finally, key points are randomly sampled on the segmented pelvic CT, and key points are extracted from the two groups of segmented CT data before and during surgery respectively. The closed-form function is used to minimize the key point distance to obtain the registration matrix and solve the registration parameters.

[0069] like Figure 2 As shown, in the above implementation, preferably, a 3D reconstruction network model is constructed using an encoding module, a cross-dimensional connection module and a decoding module, and training is completed based on a generative adversarial loss function. The specific process includes:

[0070] The encoder-decoder structure is used to build a 3D reconstruction network model, and the encoder and decoder are connected through a cross-dimensional connection module;

[0071] The encoder is composed of a preset number of dense blocks DenseBlock, the decoder is composed of a preset number of upsampling modules and a three-dimensional feature extraction module Basic3D, and the cross-dimensional connection module includes two cross-dimensional connection modules Connection A and Connection B. Connection A is used to map the two-dimensional features of the encoder to the three-dimensional feature space, and Connection B is used to fuse the two-dimensional features and the three-dimensional features in the encoding and decoding stages;

[0072] The three-dimensional reconstruction network model is trained using two-dimensional X-ray images and the corresponding three-dimensional CT images, and based on the generative adversarial loss function, the training is completed when the three-dimensional reconstruction network model can reach the preset loss value.

[0073] In this embodiment, a deep learning network framework for generating three-dimensional CT images based on two-dimensional X-ray images adopts an encoder-decoder structure as a whole model, which can effectively extract information from two-dimensional features and generate high-quality three-dimensional reconstruction results. The model framework mainly includes an encoding module, a cross-dimensional connection module and a decoding module, and the specific design is as follows: the input of the model is a two-dimensional X-ray image, and the output is a corresponding three-dimensional CT image. The model completes the mapping from two-dimensional to three-dimensional by fully extracting the features of the two-dimensional image and combining cross-dimensional feature interactions.

[0074] Specifically, the encoder module consists of multiple Dense Blocks, which are used to extract deep features of the input two-dimensional image. Dense connections are used inside each Dense Block to improve feature reuse and enhance information fluidity. At the same time, in order to gradually reduce the spatial resolution of the feature map, a convolution (Compress layer) with a step size of 2 is used after each Dense Block to implement downsampling operations. In the process of feature extraction, the Basic 2D module is also combined after the Dense Block to further refine the two-dimensional features. The Basic 2D module contains two-dimensional convolution (Conv2D), ReLU activation function, and instance normalization (IN), which can effectively improve the feature expression capability.

[0075] like Figure 3As shown in the figure, further, in order to realize the conversion from two-dimensional features to three-dimensional features, the model designs two cross-dimensional connection modules. Among them, Connection A is mainly used to map the two-dimensional features of the encoder to the three-dimensional feature space. The specific process is as follows: Flatten flattens the input two-dimensional features into a one-dimensional vector for processing in the fully connected layer. The features are linearly transformed through the fully connected layer, combined with Dropout to prevent overfitting, and the ReLU activation function is used to improve the nonlinear expression ability. Reshape reshapes the one-dimensional vector output by the fully connected layer into a three-dimensional feature map to complete the conversion from two-dimensional to three-dimensional. This design effectively maps the two-dimensional features into global features in three-dimensional space by making full use of the modeling ability of the fully connected layer for global information. Connection B is used to fuse two-dimensional and three-dimensional features in the encoding and decoding stages to realize the gradual feature conversion from two-dimensional to three-dimensional. The specific process is as follows: The Basic 2D module first extracts two-dimensional features through a two-dimensional convolution module and further optimizes them. Expand expands the two-dimensional features into three-dimensional features as the input for subsequent three-dimensional feature extraction. The Basic 3D module extracts high-dimensional features through three-dimensional convolution and enhances the expressiveness of three-dimensional features.

[0076] The design of Connection B fully considers the correlation between 2D and 3D features, and can effectively fuse multi-level feature information to provide richer features for subsequent decoding. In general, Connection A focuses more on the conversion of global features and is suitable for direct mapping from 2D to 3D, while Connection B emphasizes gradual feature fusion and the expression of multi-scale features.

[0077] The core goal of the decoder module is to reconstruct a 3D image by gradually upsampling. Specifically: The decoder consists of multiple upsampling modules (UP). Each upsampling module first uses transposed convolution (Deconv3D) to gradually restore the spatial resolution of the feature map, and further optimizes the features in combination with ReLU and IN. During the decoding process, the high-resolution features of the encoder are introduced into the decoder through skip connections (SkipConnection) to better preserve detail information. The 3D feature extraction module Basic 3D in the decoder consists of 3D convolution (Conv3D), ReLU and IN, which further enhances the extraction and reconstruction capabilities of 3D features during the decoding process.

[0078] Through layer-by-layer upsampling of the decoder, the model finally generates a high-resolution three-dimensional CT image. This design can effectively retain the structural information of the input two-dimensional image and complete the mapping from two-dimensional to three-dimensional. Through dense feature extraction, cross-dimensional interaction of two-dimensional and three-dimensional features, and jump connection design, the network can not only make full use of the global and local information in the two-dimensional image, but also effectively realize the reconstruction of the three-dimensional structure. It is particularly suitable for medical image reconstruction tasks of complex anatomical structures such as the pelvis.

[0079] In order to train the reconstruction network, the present invention uses a generative adversarial loss function, in which the loss of the generator consists of three parts, namely, a generative adversarial loss function, a loss function between the reconstruction result and the real result, and a loss function between the projection image and the real projection image along three axes.

[0080] like Figure 4 As shown, in the above embodiment, preferably, the preoperative CT image and the intraoperative CT image are segmented using the MedSAM segmentation model respectively to obtain the preoperative segmentation image and the intraoperative segmentation image of the pelvic region, and the specific process includes:

[0081] The MedSAM segmentation model is constructed using an image encoder, a prompt encoder, and a mask decoder. The image encoder is used to extract the global features of the CT image and generate image embeddings. The prompt encoder is used to generate corresponding feature vectors according to the prompt box information input by the user. The mask decoder is used to fuse the output features of the image encoder and the prompt encoder, and generate the segmentation result of the target area through decoding operations.

[0082] The preoperative CT image and the intraoperative CT image are input into the MedSAM segmentation model to obtain the preoperative segmentation image mask and the intraoperative segmentation image mask for the pelvic region.

[0083] In this embodiment, since the subsequent registration requires key point acquisition on CT, the CT image has interference parts such as soft tissue and air in addition to the pelvis, which is not conducive to key point acquisition, and the use of segmentation results is more convenient for the calculation of registration loss, which can accelerate the convergence speed of the registration network. Therefore, it is very important to perform pelvic segmentation on CT. Due to the small amount of known real data, the cost of manual annotation of medical images is high, especially in pelvic CT images, the anatomical structure is complex and the regional differences are obvious. Therefore, the present invention uses the Medsam medical image large model to increase the sample size of training data. The MedSAM large model is an automatic segmentation method based on the extension of the Segment Anything Model (SAM). MedSAM achieves efficient segmentation mask generation by combining image embedding with prompts.

[0084] The MedSAM model consists of three parts: Image Encoder, Prompt Encoder and Mask Decoder. Its specific functions are as follows: Image Encoder The image encoder is responsible for extracting the global features of the input medical image and generating image embedding. In an embodiment of the present invention, the pelvic CT image is used as the model input and encoded by the pre-trained Vision Transformer (ViT). The encoder not only captures the spatial context information of the image, but also retains the fine-grained anatomical details, providing sufficient feature support for subsequent segmentation.

[0085] The prompt encoder mainly processes the prompt information input by the user. This implementation uses a prompt box. By embedding the prompt box into the encoder, the prompt encoder generates a feature vector related to the prompt. These prompts can be the approximate range or key points of the anatomical region to guide the generation of the segmentation mask. The introduction of prompt information provides flexibility to the model, allowing users to annotate specific areas according to their needs.

[0086] The task of the mask decoder is to combine the image embedding and the hint embedding to generate a segmentation mask. Specifically, the decoder first fuses the output features of the image encoder and the hint encoder, and then generates the segmentation result of the target area through a series of decoding operations. In the pelvic CT segmentation task, the mask decoder can accurately extract the boundaries and details of the pelvic area to meet the needs of high-precision segmentation.

[0087] When expanding the dataset, the pelvic CT image is input to the image encoder to extract global features and generate image embedding. The user provides prompt box information, and the prompt encoder generates prompt embedding. The mask decoder combines the above embedding information to generate a mask of the target segmentation area.

[0088] The expanded dataset is used for the MedSAM segmentation model of CT images. Since the output of CTmask should be a fully automatic process without human intervention, the prompt box should be set to the maximum when predicting the segmentation result, that is, the model should think that the key part is the global CT rather than the local CT.

[0089] like Figure 5 As shown, in the above implementation, preferably, the key points of the preoperative segmentation image and the intraoperative segmentation image are determined by a convolutional neural network, and the distance of the key points is minimized according to a closed-form function to derive a rigid transformation matrix. The specific process includes:

[0090] The preoperative segmentation image and the intraoperative segmentation image are input into the convolutional neural network Conv-COM. After multi-layer convolution, normalization, ReLU activation and downsampling operations, the features are output to the centroid layer.

[0091] Calculating the centroids of a preset number of activation maps and converting the activation maps into coordinates of key points, wherein each key point corresponds to a specific location in the anatomical structure;

[0092] According to the key point set of the preoperative segmentation image and the key point set of the intraoperative segmentation image, the closed-form function that minimizes the key point distance is derived using derivatives, and the optimal solution of the rigid transformation matrix is ​​derived.

[0093] In this embodiment, the 3D pelvic segmentation data during and before the operation are recorded as floating images X and X respectively. m and a fixed graph X f The purpose of registration is to find a rigid transformation matrix A such that X f =X m *A. The method of calculating the matrix A in this embodiment is to use X m and X f The corresponding key points are calculated using the Conv-COM convolutional neural network to obtain X m and X f The key point.

[0094] During the implementation process, in order to obtain the key points of the two images, this implementation uses the convolutional neural network Conv-COM calculation. After multiple layers of convolution, instance-level normalization, ReLU activation, and downsampling, the features are finally output to the center of mess layer to calculate the center of mass of each activation map in the K activation maps, and convert the high-dimensional activation map into the precise coordinates of the key points. Each key point corresponds to a specific position in the anatomical structure, where K is the number of key points. Compared with other layers, the CoM layer has (approximately) translation equivariance, which means that when the input image is translated, the key point position will also be translated accordingly. This feature ensures that the key point positioning remains consistent at different positions in the input image. The network inputs a fixed image X and finally outputs a key point P. In this implementation, K is selected as 64.

[0095] Based on the obtained K corresponding key points, a variable closed-form expression of rigid transformation that aligns the key points is derived. Assume that the key point set of the floating graph is p m , the set of key points of the fixed graph is p f , the optimal solution rigid transformation matrix is ​​expected to be A The distance between key points can be minimized:

[0096] ;

[0097] The minimum value can be derived through derivatives, and after derivation, a closed-form solution is obtained:

[0098] ;

[0099] In the above implementation, preferably, the intraoperative CT image is registered according to the rigid transformation matrix and the preoperative CT image, and the specific process includes:

[0100] Based on the preoperative CT image, the rigid transformation matrix of the key point set is used to perform registration transformation to obtain the registration result image;

[0101] The registration effect of the registration result image is judged according to the DICE loss between the registration result image and the intraoperative segmentation image.

[0102] Specifically, during the model training process, the overall goal is to determine the difference between the floating image and the fixed image after the obtained A is applied. The mask of the pelvic segmentation area in the intraoperative segmentation image is used to calculate the dice loss to determine the registration effect. The specific formula is:

[0103] ;

[0104] in, x m For floating graphs, x f is a fixed graph, The key point network trained for this implementation is, A is a rigid transformation matrix, so that the loss between the floating image and the fixed image after the optimal rigid transformation matrix is ​​minimized.

[0105] The present invention further proposes an unsupervised registration system for pelvic reduction surgery navigation, which uses the unsupervised registration method for pelvic reduction surgery navigation disclosed in any one of the above embodiments, including:

[0106] A reconstruction model building module is used to build a 3D reconstruction network model using an encoding module, a cross-dimensional connection module, and a decoding module, and complete training based on a generative adversarial loss function;

[0107] A three-dimensional image mapping module is used to input the intraoperative two-dimensional X-ray image into the trained three-dimensional reconstruction network model to obtain a phase-mapped three-dimensional intraoperative CT image;

[0108] A pelvic image segmentation module is used to segment the preoperative CT image and the intraoperative CT image using the MedSAM segmentation model to obtain a preoperative segmentation image and an intraoperative segmentation image of the pelvic region;

[0109] A registration matrix derivation module is used to determine the key points of the preoperative segmentation image and the intraoperative segmentation image through a convolutional neural network, and to minimize the distance between the key points according to a closed-form function to derive a rigid transformation matrix;

[0110] The pelvic image registration module is used to register the intraoperative CT image based on the rigid transformation matrix and the preoperative CT image.

[0111] In this embodiment, the two-dimensional X-ray image is converted into a three-dimensional CT image by reconstructing the model to fully utilize the depth information of the image, and the 2D / 3D problem can be converted into a 3D / 3D registration problem, which serves as a bridge for the subsequent registration process. The MedSAM segmentation model is used for image segmentation, which is used as data expansion and segmentation network to improve the registration accuracy. The unsupervised registration framework effectively combines the key information of preoperative CT and has generalization ability. It can predict the position and posture of all patients without training the network for each patient, and the prediction range of the position and posture is wider, which provides good technical support for the speed and convenience of surgical navigation.

[0112] In the above implementation, preferably, the reconstruction model building module is specifically used to:

[0113] The encoder-decoder structure is used to build a 3D reconstruction network model, and the encoder and decoder are connected through a cross-dimensional connection module;

[0114] The encoder is composed of a preset number of dense blocks DenseBlock, the decoder is composed of a preset number of upsampling modules and a three-dimensional feature extraction module Basic3D, and the cross-dimensional connection module includes two cross-dimensional connection modules Connection A and Connection B. Connection A is used to map the two-dimensional features of the encoder to the three-dimensional feature space, and Connection B is used to fuse the two-dimensional features and the three-dimensional features in the encoding and decoding stages;

[0115] The three-dimensional reconstruction network model is trained using two-dimensional X-ray images and the corresponding three-dimensional CT images, and based on the generative adversarial loss function, the training is completed when the three-dimensional reconstruction network model can reach the preset loss value.

[0116] In the above implementation, preferably, the three-dimensional image mapping module is specifically used for:

[0117] The MedSAM segmentation model is constructed using an image encoder, a prompt encoder, and a mask decoder. The image encoder is used to extract the global features of the CT image and generate image embeddings. The prompt encoder is used to generate corresponding feature vectors according to the prompt box information input by the user. The mask decoder is used to fuse the output features of the image encoder and the prompt encoder, and generate the segmentation result of the target area through decoding operations.

[0118] The preoperative CT image and the intraoperative CT image are input into the MedSAM segmentation model to obtain the preoperative segmentation image mask and the intraoperative segmentation image mask for the pelvic region.

[0119] In the above implementation, preferably, the registration matrix derivation module is specifically used for:

[0120] The preoperative segmentation image and the intraoperative segmentation image are input into the convolutional neural network Conv-COM. After multi-layer convolution, normalization, ReLU activation and downsampling operations, the features are output to the centroid layer.

[0121] Calculating the centroids of a preset number of activation maps and converting the activation maps into coordinates of key points, wherein each key point corresponds to a specific location in the anatomical structure;

[0122] According to the key point set of the preoperative segmentation image and the key point set of the intraoperative segmentation image, the closed-form function that minimizes the key point distance is derived using derivatives, and the optimal solution of the rigid transformation matrix is ​​derived.

[0123] In the above embodiment, preferably, the pelvic image registration module is specifically used for:

[0124] Based on the preoperative CT image, the rigid transformation matrix of the key point set is used to perform registration transformation to obtain the registration result image;

[0125] The registration effect of the registration result image is judged according to the DICE loss between the registration result image and the intraoperative segmentation image.

[0126] According to the unsupervised registration system for pelvic reduction surgery navigation disclosed in the above-mentioned embodiment, the functions to be implemented by each module thereof correspond to the steps of the unsupervised registration method for pelvic reduction surgery navigation disclosed in the above-mentioned embodiment. During the implementation process, the operations are performed with reference to the above-mentioned embodiment, which will not be repeated here.

[0127] According to the unsupervised registration method and system for pelvic reduction surgery navigation disclosed in the above embodiment, during the implementation process, the intraoperative two-dimensional X-ray image is input, and the reconstruction result is displayed as follows Figure 6 As shown in Figure 2, AP and LA are the DRR projection results in the frontal and lateral positions respectively. The 3D segmentation results are shown in Figure 2. Figure 7 The registration result is shown in Figure 8 As shown, it can be seen that the warped image obtained after the fixed image is transformed is basically close to the fixed one.

[0128] According to the above implementation process and results, the above method of the present invention can convert 2D / 3D problems into 3D / 3D registration problems, plays a bridging role in the subsequent registration process, and improves the registration accuracy. The unsupervised registration framework effectively combines the key information of preoperative CT and has generalization capabilities. It can predict the position and posture of all patients without training the network for each patient, and the prediction range of the position and posture is wider, which can provide good technical support for the speed and convenience of surgical navigation.

[0129] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An unsupervised registration method for pelvic reduction surgery navigation, characterized in that: include: The 3D reconstruction network model is constructed using the encoding module, cross-dimensional connection module and decoding module, and the training is completed based on the generative adversarial loss function; Inputting the intraoperative two-dimensional X-ray image into the trained three-dimensional reconstruction network model to obtain a phase-mapped three-dimensional intraoperative CT image; The preoperative CT image and the intraoperative CT image are segmented using the MedSAM segmentation model to obtain a preoperative segmentation image and an intraoperative segmentation image of the pelvic region. The specific process includes: The MedSAM segmentation model is constructed by using an image encoder, a prompt encoder and a mask decoder, wherein the image encoder is used to extract the global features of the CT image and generate an image embedding, the prompt encoder is used to generate a corresponding feature vector according to the prompt box information input by the user, and the mask decoder is used to fuse the output features of the image encoder and the prompt encoder, and generate a segmentation result of the target area through a decoding operation; Inputting the preoperative CT image and the intraoperative CT image into the MedSAM segmentation model to obtain a preoperative segmentation image mask and an intraoperative segmentation image mask for the pelvic region; The key points of the preoperative segmentation image and the intraoperative segmentation image are determined by a convolutional neural network, and the distances of the key points are minimized according to a closed-form function to derive a rigid transformation matrix. The specific process includes: The preoperative segmentation image and the intraoperative segmentation image are input into the convolutional neural network Conv-COM, and after multi-layer convolution, normalization, ReLU activation and downsampling operations, the features are output to the centroid layer; Calculating the centroids of a preset number of activation maps and converting the activation maps into coordinates of key points, wherein each key point corresponds to a specific location in the anatomical structure; According to the key point set of the preoperative segmentation image and the key point set of the intraoperative segmentation image, a closed-form function that minimizes the key point distance is derived using derivatives to derive an optimal solution of the rigid transformation matrix; The intraoperative CT image is registered according to the rigid transformation matrix and the preoperative CT image.

2. The unsupervised registration method for pelvic reduction surgery navigation according to claim 1, characterized in that: The encoding module, the cross-dimensional connection module and the decoding module are used to construct a 3D reconstruction network model, and the training is completed based on the generative adversarial loss function. The specific process includes: The encoder-decoder structure is used to build a 3D reconstruction network model, and the encoder and decoder are connected through a cross-dimensional connection module; The encoder is composed of a preset number of dense blocks DenseBlock, the decoder is composed of a preset number of upsampling modules and a three-dimensional feature extraction module Basic3D, the cross-dimensional connection module includes two cross-dimensional connection modules Connection A and Connection B, the Connection A is used to map the two-dimensional features of the encoder to a three-dimensional feature space, and the Connection B is used to fuse the two-dimensional features and the three-dimensional features in the encoding and decoding stages; The three-dimensional reconstruction network model is trained using two-dimensional X-ray images and corresponding three-dimensional CT images, and based on the generation of an adversarial loss function, the training is completed when the three-dimensional reconstruction network model can reach a preset loss value.

3. The unsupervised registration method for pelvic reduction surgery navigation according to claim 1, characterized in that: The intraoperative CT image is registered according to the rigid transformation matrix and the preoperative CT image. The specific process includes: Based on the preoperative CT image, a registration transformation is performed using a rigid transformation matrix of a key point set to obtain a registration result image; The registration effect of the registration result image is judged according to the DICE loss between the registration result image and the intraoperative segmentation image.

4. An unsupervised registration system for pelvic reduction surgery navigation, characterized in that: The unsupervised registration method for pelvic reduction surgery navigation according to any one of claims 1 to 3 is applied, comprising: A reconstruction model building module is used to build a 3D reconstruction network model using an encoding module, a cross-dimensional connection module, and a decoding module, and complete training based on a generative adversarial loss function; A three-dimensional image mapping module, used for inputting the intraoperative two-dimensional X-ray image into the trained three-dimensional reconstruction network model to obtain a phase-mapped three-dimensional intraoperative CT image; A pelvic image segmentation module, used to segment the preoperative CT image and the intraoperative CT image using the MedSAM segmentation model to obtain a preoperative segmentation image and an intraoperative segmentation image of the pelvic region; A registration matrix derivation module, used for determining key points of the preoperative segmentation image and the intraoperative segmentation image through a convolutional neural network, and performing minimization calculation on the distance of the key points according to a closed-form function to derive a rigid transformation matrix; The pelvic image registration module is used to register the intraoperative CT image according to the rigid transformation matrix and the preoperative CT image.

5. The unsupervised registration system for pelvic reduction surgery navigation according to claim 4, characterized in that: The reconstruction model building module is specifically used for: The encoder-decoder structure is used to build a 3D reconstruction network model, and the encoder and decoder are connected through a cross-dimensional connection module; The encoder is composed of a preset number of dense blocks DenseBlock, the decoder is composed of a preset number of upsampling modules and a three-dimensional feature extraction module Basic3D, the cross-dimensional connection module includes two cross-dimensional connection modules Connection A and Connection B, the Connection A is used to map the two-dimensional features of the encoder to a three-dimensional feature space, and the Connection B is used to fuse the two-dimensional features and the three-dimensional features in the encoding and decoding stages; The three-dimensional reconstruction network model is trained using two-dimensional X-ray images and corresponding three-dimensional CT images, and based on the generation of an adversarial loss function, the training is completed when the three-dimensional reconstruction network model can reach a preset loss value.

6. The unsupervised registration system for pelvic reduction surgery navigation according to claim 5, characterized in that: The three-dimensional image mapping module is specifically used for: The MedSAM segmentation model is constructed by using an image encoder, a prompt encoder and a mask decoder, wherein the image encoder is used to extract the global features of the CT image and generate an image embedding, the prompt encoder is used to generate a corresponding feature vector according to the prompt box information input by the user, and the mask decoder is used to fuse the output features of the image encoder and the prompt encoder, and generate a segmentation result of the target area through a decoding operation; The preoperative CT image and the intraoperative CT image are input into the MedSAM segmentation model to obtain a preoperative segmentation image mask and an intraoperative segmentation image mask for the pelvic region.

7. The unsupervised registration system for pelvic reduction surgery navigation according to claim 6, characterized in that: The registration matrix derivation module is specifically used for: The preoperative segmentation image and the intraoperative segmentation image are input into the convolutional neural network Conv-COM, and after multi-layer convolution, normalization, ReLU activation and downsampling operations, the features are output to the centroid layer; Calculating the centroids of a preset number of activation maps and converting the activation maps into coordinates of key points, wherein each key point corresponds to a specific location in the anatomical structure; According to the key point set of the preoperative segmentation image and the key point set of the intraoperative segmentation image, a closed function that minimizes the key point distance is derived using derivatives to derive an optimal solution of the rigid transformation matrix.

8. The unsupervised registration system for pelvic reduction surgery navigation according to claim 7, characterized in that: The pelvic image registration module is specifically used for: Based on the preoperative CT image, a registration transformation is performed using a rigid transformation matrix of a key point set to obtain a registration result image; The registration effect of the registration result image is judged according to the DICE loss between the registration result image and the intraoperative segmentation image.

Citation Information

Patent Citations

  • Method and device and equipment for generating CT image and storage medium

    CN109745062A

  • Augmented reality fusion method in thoracoscope lung tumor resection surgical navigation

    CN116421313A