Medical image registration method and device, computer equipment and endoscope system
Patent Information
- Application Number
- CN202610906919.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]相关的配准方法中,普遍存在配准精度不高,且配准速度慢的问题
[0011]本申请实施例的技术效果包括:通过将术前三维医学图像转换为三维高斯模型,利用3DGS的可微渲染能力实现端到端的配准优化,无需手工设计特征提取和匹配步骤,配准精度高、泛化能力强。三维高斯模型以显式高斯点云表征肝脏表面细节,重建速度快、渲染质量高,术中渲染迭代速度快。由此,本申请实施例可以实现对三维模型和二维图像的高精度且实时的配准。
Smart Images

Figure CN122820779A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a medical image registration method, apparatus, computer equipment, and endoscope system. Background Technology
[0002] Endoscopic minimally invasive surgery has become the mainstream method for the treatment of soft tissue organs, such as laparoscopic liver tumor resection. However, due to the limited field of vision of the endoscope during minimally invasive surgery, the surgeon's perception of the surgical environment is limited, making it difficult for the surgeon to locate key tissue structures inside organs such as tumors.
[0003] The development of Augmented Reality (AR) navigation technology has provided a new solution to this problem—by registering the three-dimensional model of an organ to the endoscopic field of view, it can provide surgical guidance for doctors. The most critical technology lies in the registration of the three-dimensional model with the two-dimensional image during the operation.
[0004] Existing registration methods generally suffer from low registration accuracy and slow registration speed. Therefore, there is a need in the art for a medical image registration method that can improve both the registration accuracy and speed of registering 3D models and 2D images. Summary of the Invention
[0005] Therefore, it is necessary to provide a method, apparatus, computer equipment, and endoscope system for high-precision and real-time medical image registration of three-dimensional models and two-dimensional images, addressing the aforementioned technical problems.
[0006] This application provides a medical image registration method, the method comprising: Acquire preoperative 3D medical images and convert them into 3D Gaussian models. Preoperative 3D medical images are images of the target organ reconstructed in 3D before surgery.
[0007] The initial pose is used as the target camera pose to perform differentiable rendering of the 3D Gaussian model, resulting in a 2D rendering image.
[0008] Acquiring intraoperative two-dimensional endoscopic images. Intraoperative two-dimensional endoscopic images are two-dimensional images of the target organ acquired during the operation via endoscopy.
[0009] Based on the image similarity between the 2D rendered image and the intraoperative 2D endoscopic image, the target camera pose is optimized through gradient backpropagation until the convergence condition is met.
[0010] Based on the target camera pose when the convergence condition is met, the preoperative three-dimensional medical image is registered to the spatial coordinate system where the intraoperative two-dimensional endoscopic image is located.
[0011] The technical advantages of this application's embodiments include: by converting preoperative 3D medical images into 3D Gaussian models, end-to-end registration optimization is achieved using the differentiable rendering capabilities of 3DGS, eliminating the need for manual feature extraction and matching steps, resulting in high registration accuracy and strong generalization ability. The 3D Gaussian model represents liver surface details using explicit Gaussian point clouds, enabling fast reconstruction, high rendering quality, and rapid intraoperative rendering iteration. Therefore, this application's embodiments can achieve high-precision and real-time registration of 3D models and 2D images.
[0012] In one embodiment, the convergence condition includes at least one of the following: the image similarity is greater than a preset similarity threshold, the number of optimization iterations reaches a maximum threshold, and the optimization time reaches a maximum threshold.
[0013] In one embodiment, the target organ is the liver. The initial pose is the pose corresponding to the maximum exposed view of the hepatodiaphragmatic surface of the target liver.
[0014] In one embodiment, converting preoperative three-dimensional medical images into a three-dimensional Gaussian model includes: Three-dimensional anatomical features are segmented from preoperative three-dimensional medical data to obtain a three-dimensional segmentation model with anatomical feature identifiers.
[0015] Multi-view virtual sampling is performed on the 3D segmentation model to generate multiple 2D sampled images with anatomical feature markers and corresponding sampling camera poses.
[0016] Based on the two-dimensional sampled image and the corresponding sampling camera pose, a three-dimensional Gaussian model is constructed using the 3D Gaussian splashing method.
[0017] In one embodiment, before optimizing the target camera pose through gradient backpropagation based on the similarity between the 2D rendered image and the intraoperative 2D endoscopic image, until the convergence condition is met, the method further includes: Two-dimensional anatomical feature segmentation was performed on the intraoperative two-dimensional endoscopic images to obtain intraoperative two-dimensional endoscopic images with anatomical feature labels.
[0018] Based on the anatomical feature identifiers of the 2D rendered image and the intraoperative 2D endoscopic image, the image similarity between the 2D rendered image and the intraoperative 2D endoscopic image is calculated.
[0019] In one embodiment, based on the anatomical feature identifiers of the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image, the image similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image is calculated, including: Calculate the global similarity between the 2D rendered image and the intraoperative 2D endoscopic image.
[0020] Based on the anatomical feature identifiers of the 2D rendered image and the intraoperative 2D endoscopic image, the feature similarity between the 2D rendered image and the intraoperative 2D endoscopic image is calculated.
[0021] Image similarity is calculated based on global similarity and feature similarity.
[0022] In one embodiment, multi-view virtual sampling is performed on the 3D segmentation model to generate multiple 2D sampled images with anatomical feature markers and corresponding sampling camera poses, including: Generate a regular icosahedron based on the radius of the spherical bounding box of the target organ.
[0023] The camera pose is constructed based on the vertices of a regular icosahedron, and it is determined whether the recursion requirement is met.
[0024] When the recursion requirement is not met, each triangular face of the regular icosahedron is divided into 4 triangular faces. The corresponding sampling camera pose is constructed at the vertices of the new polyhedron, and a two-dimensional sampling image of the viewpoint corresponding to each sampling camera pose is generated, resulting in multiple two-dimensional sampling images with anatomical feature labels and the corresponding sampling camera poses.
[0025] In one embodiment, the target camera pose is optimized through gradient backpropagation until a convergence condition is met, including: Using the initial pose as the initial value, a loss function is constructed based on image similarity. The gradient is calculated through backpropagation of the loss function, and the increment on the Lie algebra se(3) is mapped back to the SE(3) group space using exponential mapping to achieve iterative update of the target camera pose until the convergence condition is met. A medical image registration device, the device includes: The model conversion unit is used to acquire preoperative 3D medical images and convert them into 3D Gaussian models. Preoperative 3D medical images are images of the target organ reconstructed in 3D before surgery.
[0026] The 2D rendering unit is used to perform differentiable rendering of the 3D Gaussian model with the initial pose as the target camera pose, and obtain a 2D rendering image.
[0027] The image acquisition unit is used to acquire intraoperative two-dimensional endoscopic images. Intraoperative two-dimensional endoscopic images are two-dimensional images of the target organ acquired during the operation via endoscopy.
[0028] The pose matching unit is used to optimize the target camera pose through gradient backpropagation based on the image similarity between the 2D rendered image and the intraoperative 2D endoscopic image until the convergence condition is met.
[0029] The registration unit is used to register the preoperative three-dimensional medical image to the spatial coordinate system of the intraoperative two-dimensional endoscopic image based on the target camera pose when the convergence condition is met.
[0030] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the various medical image registration method embodiments described above.
[0031] An endoscope system includes: an endoscope, an image processing device, a light source device, a display device, and a computer device as described in the above embodiments.
[0032] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described embodiments of the medical image registration methods.
[0033] The beneficial effects of the above-mentioned medical image registration device, computer equipment, endoscope system, and storage medium can be seen in the relevant descriptions of the above-mentioned medical image registration method embodiments. Attached Figure Description
[0034] Figure 1 This is a schematic diagram of an endoscope system provided in one embodiment; Figure 2 This is a flowchart illustrating a medical image registration method in one embodiment; Figure 3 This is a flowchart illustrating a three-dimensional Gaussian model conversion method in one embodiment; Figure 4 This is a flowchart illustrating a multi-view virtual sampling method in one embodiment; Figure 5 This is a flowchart illustrating a medical image registration method in another embodiment; Figure 6 This is a flowchart illustrating the similarity calculation method in another embodiment; Figure 7 This is a structural block diagram of a medical image registration device in one embodiment; Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0036] Please refer to Figure 1 , Figure 1This is a schematic diagram of an endoscope system provided in an embodiment of this application. The endoscope system may include: a light source device 101, an endoscope 102 (or endoscope body), an image processing device 103, and a display device 104. The image processing device 103 can serve as the execution subject of the medical image registration method in the following embodiments, completing the medical image registration required by the embodiments of this application. The various devices are described in sequence below: For light source device 101: The light source device 101 may include a light source component and optical elements. The light source component is used to generate a light beam of a specific wavelength. The number of light source components can be one or more, such as one, two, three, four, five, or six. The type of light source component can be an LED light source or a laser light source, etc., without further limitation. The optical elements may include one or more lenses, such as, but not limited to, collimating lenses, condenser lenses, etc.
[0037] After the light source device 101 and the endoscope 102 are successfully connected through their respective connecting parts, the light source device 101 can be used to direct the combined light beam (light) onto the end face of the light guide of the endoscope 102, so that the light beam is emitted from the tip of the endoscope 102, providing illumination for the endoscope 102. For example, when using the endoscope 102 to examine the abdominal cavity, the light source device 101 directs light into the endoscope 102, which is then transmitted through the light guide inside the endoscope 102, and the light is irradiated onto the observed abdominal cavity area through the illumination window at the tip of the endoscope 102, achieving effective illumination of the abdominal cavity area.
[0038] Regarding endoscope 102: Endoscope 102 can be a flexible endoscope or a rigid endoscope. The specific structure of endoscope 102 will vary depending on the type of endoscope, and no further restrictions will be made here.
[0039] For image processing device 103: Image processing device 103 is a dedicated processing device specifically designed for endoscope systems. Image processing device 103 can perform image processing on images acquired by endoscope 102 and send the processing results to display device 104. Simultaneously, image processing device 103 can also perform image analysis to achieve certain preset functions. For example, in some optional embodiments, image processing device 103 can perform image enhancement processing (including image sharpening), physiological part recognition, etc. The specific image analysis functions supported by image processing device 103 are not limited here and can be determined according to the actual application.
[0040] For display device 104: Display device 104 can be a liquid crystal display (LCD), a light-emitting diode (LED) display, a VR headset, or other display devices with display capabilities. Display device 104 can display images based on image processing results provided by image processing device 103. In some examples, display device 104 can display in real time images captured by the imaging module of endoscope 102, as well as the recognition results (such as lesion recognition) of the images by image processing device 103. In other examples, display device 104 can also display a three-dimensional model, and simultaneously display the registered three-dimensional model and two-dimensional image.
[0041] The following describes the workflow of endoscopic systems in common clinical scenarios: During a patient examination, medical personnel (i.e., the user) first connect and start the various devices of the endoscope system, then insert the insertion part into the patient's body. At this time, the light source device 101 shines light into the endoscope 102, which is then emitted from the illumination window at the tip of the endoscope 102 to the physiological part inside the patient's body containing the object to be measured. After being illuminated by the light, the physiological part reflects the illuminated light into the imaging module of the endoscope 102. The imaging module generates an image of the physiological part based on the received reflected light and transmits the image via cable to the image processing device 103. Finally, the image processing device 103 processes the image and sends it to the display device 104 for output display.
[0042] It should also be noted that, Figure 1 This does not constitute a limitation on the endoscope system. In practical applications, the endoscope system may include more or fewer components than illustrated, or combine certain components, or use different components. For example, the endoscope system may also include a trolley, accessory equipment (such as a carbon dioxide pump), and image registration equipment specifically for performing image registration (such as an image registration host or remote server, etc.), etc., which can be determined according to the actual application situation, and this application does not impose any restrictions on this. In addition, the connection relationship between the various devices may also have certain differences. For example, in some alternative embodiments, when the endoscope system includes image registration, the image registration can be set between the image processing device 103 and the display device 104, thereby realizing secondary processing and re-display of the image output by the image processing device 103. Correspondingly, depending on the different connection relationships, the data transmission path of the endoscope 102 will also have adaptive differences, which will not be elaborated here.
[0043] The execution entity computer device in the following embodiments of this application can be Figure 1 The image processing device in the illustrated embodiment can also be in Figure 1The additional devices added to the illustrated embodiments, for example in some alternative embodiments, may be image registration devices dedicated to image registration.
[0044] The following describes some features that may be involved in the embodiments of this application: 3D Gaussian Model (3DGS): A three-dimensional representation model built based on 3D Gaussian splashing technology, which uses a large set of three-dimensional Gaussian distributions to represent the three-dimensional surface and appearance information of the target.
[0045] Differentiable rendering: a rendering technique in which the rendering process is differentiable with respect to input parameters (such as camera pose, 3D model parameters, etc.), thus enabling optimization through gradient backpropagation.
[0046] Anatomical features: Locations or structures on the target organ that have significant anatomical meaning, such as the diaphragmatic surface, visceral surface, falciform ligament, and hepatic ridge line of the liver.
[0047] SE(3): Special Euclidean group, representing the set of rigid body transformations (rotation + translation) in three-dimensional space.
[0048] se(3): The Lie algebra of SE(3) is a six-dimensional vector space containing three rotational degrees of freedom and three translational degrees of freedom.
[0049] To achieve registration between 3D models and 2D images, some possible methods include: 1. A method for ICP registration based on feature line keypoint matching. This method extracts feature points from the 3D model and the 2D image, finds the corresponding keypoints using scale normalization, and then performs ICP registration. This approach is essentially an improvement on traditional geometric feature matching; however, due to its heavy reliance on manually designed geometric feature extraction, the algorithm has poor adaptability, registration accuracy is greatly affected by the quality of feature matching, and the multiple manually defined geometric operations result in low registration efficiency.
[0050] 2. An end-to-end image registration method based on a monocular camera and differentiable rendering. This method extracts the structured boundary information of the target object in the input image through a contour segmentation model to generate a binary mask image, and then performs registration through a 6D pose estimation deep network that fuses a convolutional neural network (CNN) and a Transformer module. This approach belongs to deep learning and requires a large amount of training data to train the model. However, training data is difficult to obtain in endoscopic scenarios, resulting in poor practicality.
[0051] 3. A method applying Neural Rendering (NeRF) to registration. This method utilizes style transfer to process intraoperative model appearance information, estimating surgical camera pose by minimizing the difference between the rendered image and the intraoperative target image. However, style transfer in registration can lead to geometric misalignment—blurred edges, distorted anatomical structures, and the generation of false textures and structures, causing registration results to drift and become inaccurate. Furthermore, the NeRF renderer requires sampling multiple points along the light rays, resulting in a large computational load in the forward process. Both preoperative model scene reconstruction and intraoperative registration are time-consuming, making real-time registration difficult.
[0052] 4. A method for introducing neural radiation fields into endoscopic liver surgery navigation and registration. This method reconstructs the liver organ in the surgical scene based on NeRF, and then performs 3D-3D registration between preoperative and intraoperative liver model point clouds based on the Iterative Closest Point (ICP) scheme. However, the endoscopic field of view is relatively narrow in minimally invasive surgery, and even with the most extensive circumferential scanning around the liver organ, it is difficult to expose the complete liver organ, resulting in surface noise, artifacts, and incomplete structures during reconstruction. Registration errors accumulate in the subsequent ICP stage.
[0053] Therefore, there is an urgent need for a medical image registration scheme with high registration accuracy, good real-time performance, and adaptability to cross-modal scenarios. To achieve this goal, the embodiments of this application convert preoperative 3D medical images into 3D Gaussian models and use differentiable rendering technology to achieve end-to-end registration between the preoperative model and the intraoperative image. This avoids the multi-step manually designed geometric feature extraction operations in traditional registration methods, resulting in high registration accuracy and efficiency. Based on gradient backpropagation to optimize camera pose, the registration process is continuous and stable, requiring no large amount of training data, and has stronger practicality and generalization ability.
[0054] refer to Figure 2 This is a flowchart illustrating a medical image registration method provided in an embodiment of this application, detailed below: S101. Obtain preoperative three-dimensional medical images and convert them into a three-dimensional Gaussian model.
[0055] First, preoperative 3D medical images of the target organ are obtained. These images are obtained by reconstructing volumetric data or surface mesh models using 3D reconstruction technology after preoperative scanning of the patient's target organ with medical imaging equipment (such as CT scanners, MRI equipment, etc.). In other words, the preoperative 3D medical image is a model built based on the actual patient's own organ. The target organ can be any soft tissue organ of the patient, specifically determined according to the actual clinical scenario. The following examples use the patient's liver as the target organ for illustration.
[0056] For example, taking liver surgery as an example, in some specific embodiments, the patient's abdomen is scanned using a CT scanner before surgery to obtain CT volume data containing the liver. Then, image segmentation algorithms such as thresholding and region growing are used to extract the liver region from the CT volume data. Finally, algorithms such as Moving Cubes are used to reconstruct a three-dimensional mesh model of the liver. This three-dimensional mesh model is the preoperative three-dimensional medical image, which contains complete three-dimensional geometric information of the liver, including the diaphragmatic surface, visceral surface, and internal anatomical structures such as blood vessels and tumors.
[0057] Then, in this embodiment of the application, the acquired preoperative three-dimensional medical images are converted into a three-dimensional Gaussian model. The three-dimensional Gaussian model is an explicit three-dimensional representation constructed based on 3D Gaussian Splatting (3DGS) technology. It uses a large set of three-dimensional Gaussian distributions to represent the three-dimensional surface and appearance information of the target organ. Each three-dimensional Gaussian distribution contains parameters such as position (mean), scaling, rotation, color, and opacity.
[0058] As an optional embodiment of this application, refer to Figure 3 This is a flowchart illustrating a three-dimensional Gaussian model conversion method provided in an embodiment of this application, detailed below: S1011, Three-dimensional anatomical feature segmentation: Perform three-dimensional anatomical feature segmentation on preoperative three-dimensional medical images to obtain a three-dimensional segmentation model with anatomical feature labels.
[0059] Anatomical features refer to the parts or structures on the target organ that have obvious anatomical significance. Taking the liver as an example, anatomical features include, but are not limited to: the diaphragmatic surface of the liver, the visceral surface of the liver, the falciform ligament, the hepatic ridge line, the segmental boundaries of the liver, and the course of the portal vein branches.
[0060] Specifically, deep learning-based segmentation networks (such as U-Net and V-Net) can be used to perform semantic segmentation of the aforementioned anatomical features on the 3D model, assigning corresponding anatomical feature labels to each region on the model surface (e.g., red for the diaphragm, blue for the visceral surface, green for the falciform ligament, etc.). Alternatively, manual segmentation can be used, where doctors or technicians manually label the anatomical feature regions on the 3D model. After segmentation, in the 3D segmentation model, each vertex or voxel carries identification information of its associated anatomical feature in addition to its geometric coordinates.
[0061] S1012, Multi-view Virtual Sampling: Perform multi-view virtual sampling on the 3D segmentation model to generate multiple 2D sampling images with anatomical feature labels and corresponding sampling camera poses.
[0062] To reconstruct a complete 3D Gaussian model of the organ, it is necessary to obtain 2D views of the 3D model containing anatomical feature information from multiple perspectives. This step designs a sampling procedure based on recursive partitioning of regular polyhedra to ensure uniform coverage with fewer sampling perspectives, balancing reconstruction efficiency and quality. For a detailed description of the sampling strategy in Example 2, please refer to Example 2. Through this sampling procedure, multiple 2D sampling images (i.e., virtual endoscopic images) with anatomical feature markers are obtained. Each sampling image records the organ surface and its anatomical feature distribution from the corresponding perspective, and also records the camera pose (including camera position, orientation, field of view, and other parameters) corresponding to each sampling image.
[0063] In some embodiments, the sampling camera can be set as a fluoroscopic camera model, and its intrinsic parameter matrix (including focal length, principal point coordinates, etc.) is set according to the calibration parameters of the actual endoscope so that the generated virtual sampling image is consistent with the real intraoperative endoscopic image in terms of imaging parameters, thereby reducing the scale uncertainty of subsequent registration.
[0064] S103, 3D Gaussian Model Reconstruction: Based on the multiple 2D sampled images with anatomical feature markers generated in step S1-2 and the corresponding sampled camera poses, a 3D Gaussian model is constructed using the 3D Gaussian splashing method.
[0065] Specifically, each point or vertex in the preoperative 3D medical image is initialized as a 3D Gaussian distribution. Then, the initial set of 3D Gaussian distributions is rendered using the known sampling camera pose to obtain a rendered image. Image differences (such as L1 loss, SSIM loss, etc.) between the rendered image and the corresponding original 2D sampling image are calculated. The position, scaling, rotation, color, and opacity parameters of all 3D Gaussian distributions are optimized through gradient backpropagation. During optimization, an adaptive density control strategy is employed, performing cloning, splitting, and culling operations on the Gaussian distributions based on gradient information, adaptively matching the density distribution of the Gaussian distribution to the geometric complexity of the liver surface. After optimization convergence, the final 3D Gaussian model is obtained. This model represents liver surface details in the form of an explicit Gaussian point cloud, accurately capturing subtle anatomical structures such as liver lobe folds and vascular bulges.
[0066] S102. Using the initial pose as the target camera pose, perform differentiable rendering on the 3D Gaussian model to obtain a 2D rendering image.
[0067] An initial pose is set for the target camera (i.e., the virtual counterpart of the intraoperative endoscope). This initial pose serves as the starting point for subsequent iterative optimization processes. Taking liver surgery as an example, since the endoscope typically observes the liver from a specific angle during the liver exploration phase, the initial pose can be set to the pose corresponding to the maximum exposure angle of the hepatoseptal surface of the target liver (i.e., the AP position commonly referred to clinically). This angle maximizes the exposure of the hepatoseptal region and is the most commonly used angle for intraoperative endoscopic observation of the liver. Selecting this angle as the initial pose minimizes the difference between the initial rendered image and the actual intraoperative endoscopic image, thereby reducing the number of iterations and accelerating registration convergence.
[0068] In practice, the initial pose can be set as follows: A preset observation sphere radius is defined in a three-dimensional Gaussian model space. The target camera is placed on the surface of this sphere, facing the center of the liver, with the camera's optical axis roughly aligned with the normal direction of the hepatic diaphragm. This preset radius can be set based on clinical experience, for example, 30-50 cm from the center of the liver. The three-dimensional coordinates and orientation of the initial pose are pre-calibrated using typical intraoperative endoscope placement, or automatically calculated based on the patient's preoperative imaging data.
[0069] Based on the existing camera pose, and using the current initial pose (or the target camera pose updated in subsequent iterations), a 3D Gaussian model is rendered in a differentiable manner to generate a 2D rendered image. The rendering process employs a 3D Gaussian splashing differentiable rasterization algorithm, which projects the 3D Gaussian distribution from 3D space onto the 2D image plane. Alpha mixing is then performed on the Gaussian distribution at each pixel location to obtain the final rendered image.
[0070] S103. Obtain intraoperative two-dimensional endoscopic images.
[0071] Two-dimensional endoscopic images of the target organ are acquired in real time during surgery using an endoscope. Taking liver surgery as an example, during the operation, the surgeon inserts an endoscope into the abdominal cavity through a trocar in the patient's abdomen. A high-definition camera at the tip of the endoscope acquires real-time laparoscopic images of the liver surface. The acquired intraoperative two-dimensional endoscopic images contain information such as the true texture, color, and lighting of the liver surface, but may also contain interfering factors such as surgical instruments, high light reflection, blood, and smoke.
[0072] S104. Based on the image similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image, the target camera pose is optimized through gradient backpropagation until the convergence condition is met.
[0073] This application's embodiments model the registration problem as a gradient-based optimization problem: using the target camera pose as the optimization variable and the image similarity between the 2D rendered image and the intraoperative 2D endoscopic image as the objective function, the target camera pose is iteratively updated through gradient backpropagation, making the rendered image and the endoscopic image more consistent in image content, thereby finding the optimal camera pose that matches the intraoperative endoscopic viewpoint. Specifically, after rendering the 2D rendered image based on the initial pose, the image similarity between the 2D rendered image and the intraoperative 2D endoscopic image is calculated. If the similarity does not meet the similarity threshold, the camera pose is adjusted, and a new 2D rendered image is rendered based on the new target camera pose. The image similarity between the new 2D rendered image and the intraoperative 2D endoscopic image is then calculated again. This achieves continuous iterative adjustment of the camera pose, thereby finding a suitable target camera pose.
[0074] As an optional embodiment of this application, in order to realize gradient backpropagation in S104, S104 in this embodiment includes: S1041 to S1043, which are detailed below: S1041. Calculate image similarity: Calculate the image similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image. In one embodiment, the similarity can be calculated based on pixel-level differences (such as L1 loss, SSIM, etc.). In a preferred embodiment, it involves the segmentation and matching of anatomical features, as described in detail in Embodiment 3.
[0075] S1042. Gradient Calculation and Pose Update: Using the current target camera pose as the initial value, a loss function is constructed based on image similarity. The gradient of the target camera pose is backpropagated through the loss function to achieve iterative update of the target camera pose.
[0076] To ensure the continuity and stability of the pose optimization process, in a preferred embodiment, the camera pose update is modeled as an optimization problem on the Lie group SE(3). Using the current camera pose as the initial value, the gradient is calculated through backpropagation of the loss function, and the increment on the Lie algebra se(3) is mapped back to the SE(3) group space using an exponential mapping (expmap), achieving efficient iterative update of the camera pose. This method fully utilizes the geometric properties of the Lie group, avoids the singularity problem of Euler angle representation, and guarantees the numerical stability and convergence speed of the pose optimization.
[0077] S1043. Convergence Judgment: Determine whether the preset convergence conditions are met. The convergence conditions include at least one of the following: the image similarity is greater than the preset similarity threshold, the number of optimization iterations reaches the upper limit threshold, and the optimization time reaches the upper limit threshold.
[0078] The convergence criteria include at least one of the following: image similarity greater than a preset similarity threshold (e.g., Dice coefficient greater than 95%), the number of optimization iterations reaching the upper limit threshold (e.g., 500 iterations), and the optimization time reaching the upper limit threshold (e.g., 2 seconds). Specific thresholds can be selected according to actual needs; the foregoing is merely illustrative and not limited to specific embodiments in this application.
[0079] When any convergence condition is met, the iteration stops, and the current target camera pose is output. The above convergence conditions take into account both registration accuracy and real-time requirements: a similarity threshold ensures that registration accuracy reaches a clinically acceptable level; an upper limit on the number of iterations prevents the algorithm from running indefinitely when there is no solution; and an upper limit on the time ensures real-time feedback capability for intraoperative registration.
[0080] S105, based on the target camera pose when the convergence condition is met, register the preoperative three-dimensional medical image to the spatial coordinate system where the intraoperative two-dimensional endoscopic image is located.
[0081] After optimizing the target camera pose, this pose represents the rigid body transformation matrix (containing the rotation matrix R and translation vector t) from the preoperative 3D model space to the intraoperative endoscopic camera space. Then, this transformation matrix is applied to the preoperative 3D medical image (e.g., the original CT reconstructed 3D mesh model), transforming it from the preoperative 3D coordinate system to the 2D image coordinate system of the intraoperative endoscopic image, achieving accurate registration between the preoperative 3D model and the intraoperative 2D endoscopic image.
[0082] In this embodiment, the original preoperative 3D medical image (containing complete internal anatomical information, such as blood vessels and tumors) is used for final registration, rather than the 3D Gaussian model (containing only surface information). This is because the main function of the 3D Gaussian model is to provide differentiable rendering capabilities for pose optimization, while the original 3D model contains all the anatomical structural information required for surgical navigation. Since the 3D Gaussian model and the original 3D model are completely coincident in the spatial coordinate system, the target camera pose optimized by S104 can be directly reused in the original 3D model.
[0083] The technical advantages of this application's embodiments include: by converting preoperative 3D medical images into 3D Gaussian models, end-to-end registration optimization is achieved using the differentiable rendering capabilities of 3DGS, eliminating the need for manual feature extraction and matching steps, resulting in high registration accuracy and strong generalization ability. The 3D Gaussian model represents liver surface details using explicit Gaussian point clouds, enabling rapid reconstruction (completed within minutes), high rendering quality (sharper edges and textures), and fast intraoperative rendering iteration speed (down to the second level), without requiring pre-training data. The initial pose is set based on common intraoperative viewpoints, accelerating registration convergence. Therefore, this application's embodiments can achieve high-precision and real-time registration of 3D models and 2D images. After registration, the deep structural information in the preoperative model can be accurately superimposed onto the intraoperative endoscopic field of view, providing navigation reference for the surgeon.
[0084] As an optional embodiment of this application, refer to Figure 4 This is a flowchart illustrating a multi-view virtual sampling method provided in an embodiment of this application. In this embodiment, S1012 multi-view virtual sampling includes: S10121 to S10123, detailed below: S10121. Generate a regular icosahedron based on the radius of the spherical bounding box of the target organ.
[0085] The core of this embodiment is to design a perspective sampling strategy based on recursive partitioning of regular polyhedra for the irregular shapes of organs such as the liver. This strategy ensures that the observation directions of each perspective are distributed as evenly as possible in three-dimensional space with a small number of sampling perspectives. This reduces the number of sampling images that need to be generated and lowers the computational burden of reconstruction while ensuring the quality of the three-dimensional Gaussian model reconstruction.
[0086] Based on this, the embodiments of this application first calculate the bounding box of the three-dimensional model of the target organ (such as the liver) and obtain the geometric center and radius of the bounding box. Preferably, a spherical bounding box is used to approximate the spatial extent of the liver, that is, a sphere defined by the geometric center of the liver model as the center of the sphere and the distance from the point on the model farthest from the center of the sphere to the center of the sphere as the radius.
[0087] Then, using the center of the sphere as the geometric center and the radius of the bounding box as a reference, a regular icosahedron circumscribed by the sphere is generated. The regular icosahedron is the regular polyhedron with the most faces in Platonic solids, containing 20 equilateral triangular faces and 12 vertices.
[0088] S10122. Construct the corresponding camera pose based on the vertices of a regular icosahedron and identify whether the recursion requirement is met.
[0089] Virtual target cameras are positioned at the 12 vertices of a regular icosahedron, with each camera facing the center of the sphere (i.e., the center of the liver). The direction pointing towards the center of the sphere is used as the optical axis of the camera, constructing 12 initial sampling viewpoints. For each viewpoint, a two-dimensional sampled image is generated from that viewpoint.
[0090] Determine whether the current number of viewpoints meets the recursive requirement. The recursive requirement can be a preset target number of viewpoints, or a viewpoint requirement determined after analyzing the model's complexity. In one implementation, the recursive requirement is whether the number of viewpoints is sufficient to cover the organ's surface (e.g., by calculating whether the angular interval between adjacent viewpoints is less than a preset threshold).
[0091] S10123. When the recursion requirement is not met, divide each triangular face of the regular icosahedron into 4 triangular faces, construct the corresponding sampling camera pose at the vertices of the new polyhedron, and generate a two-dimensional sampling image of the viewpoint corresponding to each sampling camera pose.
[0092] When the current number of viewpoints is insufficient, a recursive partitioning operation is performed: for each triangular face of the icosahedron, the midpoint of each edge is taken, and the original triangular face is divided into 4 smaller triangles (i.e., each original triangular face is divided into 3 new triangular faces). In this way, the new polyhedron will contain more vertices, and each vertex corresponds to a new sampling viewpoint.
[0093] It is important to note that the newly added vertices after partitioning are located on the edges of the original triangle face, and these new vertices still lie on the sphere centered at the sphere's center. For each newly added vertex, the camera is also set at the vertex position and facing the sphere's center to construct the corresponding sampling camera pose and generate the corresponding 2D sampling image.
[0094] Repeat the above recursive division process until the number of generated viewpoints meets the preset recursive requirements (e.g., reaching the preset target number of 540 viewpoints), or the angular interval between adjacent viewpoints is less than the preset threshold.
[0095] The core advantage of this sampling strategy lies in its balance between efficiency and effectiveness. Without a uniform sampling strategy, simple dense sampling would require generating a large number of sampled images (e.g., thousands), significantly increasing the computational burden and training time for 3D Gaussian model reconstruction. Sparse sampling, on the other hand, could lead to uneven viewpoint coverage, with some areas (such as the concave areas of the liver or more hidden locations) lacking sufficient viewpoint information, affecting reconstruction quality. This embodiment uses a recursive partitioning of an icosahedron to ensure that each new viewpoint is uniformly distributed on the sphere, achieving the most uniform coverage possible with the fewest possible viewpoints.
[0096] The technical effect of this embodiment is that by adopting a view sampling strategy based on the recursive partitioning of a regular icosahedron, uniform coverage in three-dimensional space can be achieved with a small number of viewpoints, effectively balancing reconstruction efficiency (reducing computation) and reconstruction effect (ensuring coverage integrity), and providing high-quality input data for the rapid and high-precision reconstruction of the subsequent three-dimensional Gaussian model.
[0097] As an optional embodiment of this application, refer to Figure 5 ,exist Figure 1 and Figure 3 Based on the illustrated embodiment, after S103 and before S104, the embodiments of this application further include: S106. Anatomical feature segmentation of intraoperative two-dimensional endoscopic images: Two-dimensional anatomical feature segmentation is performed on intraoperative two-dimensional endoscopic images to obtain intraoperative two-dimensional endoscopic images with anatomical feature labels.
[0098] During the procedure, after acquiring real-time two-dimensional endoscopic images, the images are segmented into two-dimensional anatomical features to obtain intraoperative two-dimensional endoscopic images with anatomical feature labels.
[0099] Specifically, deep learning-based image segmentation models (such as SAMV2) can be used to segment liver anatomical features in intraoperative endoscopic images. The segmentation goal is consistent with the preoperative goal, that is, to segment regions such as the hepatic diaphragm, visceral surface, falciform ligament, and hepatic ridge line in the endoscopic image under the same anatomical feature classification system, and assign them the same anatomical feature labels as before the operation (such as the same color coding or label coding).
[0100] In one implementation, manual segmentation can be used, where the surgeon manually annotates the endoscopic images during the procedure. However, considering the real-time requirements of the surgery, an AI-based automatic segmentation scheme is preferred.
[0101] S107. Image similarity calculation based on anatomical features: Calculate the image similarity between the two-dimensional rendering image and the intraoperative two-dimensional endoscopic image based on the anatomical feature identifiers of the two-dimensional rendering image and the anatomical feature identifiers of the intraoperative two-dimensional endoscopic image.
[0102] After the anatomical features of the preoperative 3D Gaussian model's 2D rendering and the intraoperative 2D endoscopic image were segmented, both were transformed into a unified "anatomical feature representation" space, in which effective similarity measurement could be performed.
[0103] As an optional embodiment of this application, a multi-level similarity measurement method is preferred, including two levels: global similarity and feature similarity. (Reference) Figure 6 This is a flowchart illustrating a similarity calculation method provided in an embodiment of this application, detailed below: S1071. Calculate global similarity: Calculate the global similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image at the overall image level.
[0104] Global similarity measures the degree of alignment between two images in terms of overall structure. Preferably, a combination of L1 loss and SSIM (structural similarity) can be used to measure global similarity. L1 loss measures absolute differences at the pixel level, while SSIM measures the consistency of image structure, brightness, and contrast.
[0105] S1072. Calculate anatomical feature similarity: Based on the anatomical feature identifiers of the 2D rendered image and the anatomical feature identifiers of the intraoperative 2D endoscopic image, calculate the feature similarity between the two. Feature similarity is used to measure the alignment accuracy of the two images at the locations of key anatomical structures.
[0106] Specifically, feature regions corresponding to each anatomical feature are extracted from both the 2D rendered image and the intraoperative 2D endoscopic image with anatomical feature markers. Taking the liver as an example, the diaphragmatic region, the falciform ligament region, and the hepatic ridge region are extracted respectively. Then, the positional deviation of each feature region in the two images is calculated.
[0107] To ensure the differentiability of the loss function, it is preferable to use one-way chamfer distance to calculate the local similarity loss of each anatomical feature region separately. Chamfer distance can measure the proximity between two point sets and is differentiable, so it can participate in gradient backpropagation.
[0108] S1073. Calculate image similarity comprehensively. Based on global similarity and feature similarity, calculate the final image similarity.
[0109] Specifically, the global loss and feature loss can be weighted and summed according to preset weights to obtain the total loss function value. The smaller the total loss, the more similar the 2D rendered image is to the intraoperative 2D endoscopic image.
[0110] In one alternative embodiment, the total loss function can be designed as follows: L_total=w1·L_global+w2·L_feature; Where L_global is the global similarity loss, L_feature is the feature similarity loss, and w1 and w2 are weight coefficients.
[0111] In another alternative embodiment, the total loss function can be designed as follows: L_total=w1·L_global+(w3·L_feature1+w4·L_feature2); Where L_global represents the global similarity loss, L_feature1 represents the local feature similarity loss of the falciform ligament, L_feature2 represents the local feature similarity loss of the ridge line, and w1, w3, and w4 are weight coefficients. To ensure the differentiability of the loss function, unidirectional chamfered distance is used to calculate the local similarity loss between the falciform ligament and the ridge line respectively.
[0112] In this embodiment of the application, the feature similarity includes the local feature similarity of the falciform ligament and the local feature similarity of the ridge line. S1072 includes anatomical feature identifiers based on the two-dimensional rendered image and the anatomical feature identifiers of the intraoperative two-dimensional endoscopic image, and calculating the local feature similarity of the falciform ligament and the local feature similarity of the ridge line between the two.
[0113] The technical advantages of this application's embodiments are as follows: By unifying the preoperative CT model (texture-free) and intraoperative endoscopic images (textured, illuminated) into an "anatomical feature representation" space through anatomical feature segmentation, the modal difference problem in cross-modal registration is effectively solved, eliminating the need for preprocessing steps such as style transfer that may introduce geometric distortions. Multi-level similarity measurement balances overall structural alignment and precise matching of local anatomical structures, improving registration accuracy and stability.
[0114] As an optional embodiment of this application, considering that directly optimizing the camera pose matrix (including the rotation matrix R and the translation vector t) in the gradient-based pose optimization process faces some difficulties: the rotation matrix R must satisfy the constraints of orthogonality and determinant of 1, and it is difficult to guarantee that the updated matrix still satisfies these constraints by directly performing gradient updates in the matrix space. In order to solve this problem, this embodiment models the update of the camera pose as an optimization problem on the Lie group SE(3). The specific implementation is as follows: The current target camera pose (initial pose or pose updated in the previous iteration) is used as the starting point for optimization. In each iteration, the gradient with respect to the camera pose is first calculated through backpropagation of the loss function. However, instead of directly updating the rotation matrix R and translation vector t in the SE(3) group space, the gradient is mapped to the Lie algebra se(3) space.
[0115] se(3) is the Lie algebra of SE(3), which is a six-dimensional vector space. Each six-dimensional vector ξ=(ω1, ω2, ω3, ν1, ν2, ν3)^T in se(3) contains three rotation components (ω) and three translation components (ν). In the se(3) space, gradient updates are linear and unconstrained, so the standard gradient descent method can be applied directly.
[0116] Then, the increment ξ on se(3) is mapped back to the SE(3) group space using an exponential map, resulting in the updated pose matrix T_new=exp(ξ)·T_old. The exponential map ensures that the updated matrix T_new is still a valid SE(3) transformation matrix, automatically satisfying the orthogonality constraint of the rotation matrix.
[0117] By using the method of "gradient calculation is performed in the se(3) space and pose update is performed by exponential mapping regression to SE(3)", continuous and stable iterative optimization of camera pose is achieved.
[0118] The technical effects of this embodiment are as follows: by modeling camera pose optimization as an optimization problem on a Lie group SE(3), the constraint problems faced when directly optimizing the pose matrix are effectively avoided, ensuring the numerical stability and convergence speed of pose updates. The six-dimensional space representation of the Lie algebra se(3) also avoids the gimbal lock problem that may occur in Euler angle representation.
[0119] The embodiments of this application described above have at least the following beneficial effects: 1. Combining 3DGS reconstruction technology with differentiable rendering registration technology balances registration accuracy and real-time performance. The liver 3D Gaussian model is reconstructed using 3DGS technology. Compared to implicit reconstruction methods such as NeRF, 3DGS uses explicit Gaussian point clouds to represent liver surface details, accurately capturing subtle anatomical structures such as liver lobe folds, resulting in higher reconstruction accuracy. Furthermore, it requires no lengthy training time and can be quickly reconstructed based on preoperative multi-view virtual images. On the other hand, during surgery, 3DGS differentiable rendering projects the 3D liver Gaussian model into a virtual image. Combined with segmentation mask similarity measurement, end-to-end gradient optimization of camera pose is achieved, avoiding the error accumulation caused by the one-time solution of the PnP method. This results in high registration accuracy, meeting the requirements for high-precision intraoperative navigation. Simultaneously, 3DGS's differentiable rasterization rendering speed is fast, enabling real-time iterative optimization of intraoperative pose.
[0120] 2. Anatomical feature-driven registration logic solves the problem of cross-modal adaptation. This application breaks through the limitations of traditional registration that relies on sparse feature points or single-pixel matching, and uses liver anatomical features as the core throughout the entire registration process: before surgery, key anatomical features such as hepatic veins, portal veins, and liver capsule are extracted from the 3D model, and multi-view virtual images are generated through vertex shading rendering, providing accurate geometric and semantic constraints for 3DGS reconstruction, ensuring that the reconstructed 3D Gaussian model can accurately fit the liver anatomical structure. In terms of cross-modal adaptation, by measuring the similarity between the anatomical feature mask and the differentiable renderable image, the modal differences between the preoperative CT model (no texture, no lighting) and the intraoperative endoscopic image (with highlights, texture) can be effectively aligned without relying on additional steps such as style transfer, avoiding registration failure caused by modal mismatch.
[0121] As an embodiment of this application, after image registration is completed in S105, this embodiment of the application can also perform fusion display of the preoperative three-dimensional model and the intraoperative endoscopic image: The registered preoperative 3D model (containing complete internal anatomical information, such as vascular networks, tumor location, and liver segment boundaries) is overlaid with the intraoperative real-time endoscopic video. Surgeons can simultaneously see the real liver surface and the overlaid deep anatomical structures (such as the location and boundaries of the tumor, the course of the portal vein and hepatic veins) within the endoscopic view, thus obtaining an augmented reality navigation view similar to a "transparent" effect.
[0122] During the surgery, when the endoscope position changes or the liver morphology changes due to the surgical procedure, steps S102 to S1055 can be repeated to achieve continuous updates of real-time intraoperative registration. In one embodiment, the coarse positional change of the endoscope can be obtained through the tracking device of the surgical navigation system (such as an optical tracker or an electromagnetic tracker) as an update reference for the initial pose, thereby reducing the number of iterations for each registration.
[0123] The technical effect of this embodiment is that by fusing the preoperative 3D model with the intraoperative endoscopic image, doctors are provided with the ability to visualize deep anatomical information beyond the naked eye, which helps to accurately locate tumors, plan surgical paths, avoid important vascular structures, and reduce surgical risks.
[0124] In one embodiment, such as Figure 7 As shown, a medical image registration device is provided, comprising: The model conversion unit 71 is used to acquire preoperative three-dimensional medical images and convert them into three-dimensional Gaussian models. The preoperative three-dimensional medical images are images of the target organ reconstructed in three dimensions before surgery.
[0125] The 2D rendering unit 72 is used to perform differentiable rendering of the 3D Gaussian model with the initial pose as the target camera pose, and obtain a 2D rendering image.
[0126] Image acquisition unit 73 is used to acquire intraoperative two-dimensional endoscopic images. Intraoperative two-dimensional endoscopic images are two-dimensional images of the target organ acquired during the operation via endoscopy.
[0127] The pose matching unit 74 is used to optimize the target camera pose through gradient backpropagation based on the image similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image until the convergence condition is met.
[0128] The registration unit 75 is used to register the preoperative three-dimensional medical image to the spatial coordinate system where the intraoperative two-dimensional endoscopic image is located, based on the target camera pose when the convergence condition is met.
[0129] The medical image registration device in this application embodiment is similar to the one described above. Figures 2 to 6 The device embodiments corresponding to the various method embodiments shown are illustrated above. Therefore, the content of the above method embodiments can also be applied to the embodiments of this application, and will not be repeated here. Furthermore, for specific limitations regarding the medical image registration device, please refer to the limitations of the medical image registration method above, which will not be repeated here. Each module in the above-described medical image registration device can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0130] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, communication interface, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a medical image registration method. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0131] Those skilled in the art will understand that Figure 8The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0132] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the various medical image registration method embodiments described above.
[0133] An endoscope system includes: an endoscope, an image processing device, a light source device, a display device, and a computer device as described in the above embodiments.
[0134] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described embodiments of the medical image registration methods.
[0135] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0136] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0137] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A medical image registration method, characterized in that, The method includes: Acquire preoperative three-dimensional medical images and convert them into three-dimensional Gaussian models; the preoperative three-dimensional medical images are images of three-dimensional reconstruction of the target organ before surgery; The three-dimensional Gaussian model is rendered using the initial pose as the target camera pose to obtain a two-dimensional rendering image. Acquire intraoperative two-dimensional endoscopic images; the intraoperative two-dimensional endoscopic images are two-dimensional images of the target organ acquired during the operation via endoscopy; Based on the image similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image, the pose of the target camera is optimized through gradient backpropagation until the convergence condition is met. Based on the target camera pose when the convergence condition is met, the preoperative three-dimensional medical image is registered to the spatial coordinate system where the intraoperative two-dimensional endoscopic image is located.
2. The method according to claim 1, characterized in that, The target organ is the target liver; the initial pose is the pose corresponding to the maximum exposed angle of the hepatic diaphragm of the target liver.
3. The method according to claim 1 or 2, characterized in that, The process of converting the preoperative three-dimensional medical image into a three-dimensional Gaussian model includes: The preoperative three-dimensional medical image is segmented into three-dimensional anatomical features to obtain a three-dimensional segmentation model with anatomical feature identifiers; The three-dimensional segmentation model is subjected to multi-view virtual sampling to generate multiple two-dimensional sampling images with anatomical feature markers and corresponding sampling camera poses; Based on the two-dimensional sampled image and the corresponding sampling camera pose, the three-dimensional Gaussian model is constructed using the 3D Gaussian splashing method.
4. The method according to claim 3, characterized in that, Before optimizing the target camera pose through gradient backpropagation based on the similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image, until the convergence condition is met, the method further includes: Two-dimensional anatomical feature segmentation is performed on the intraoperative two-dimensional endoscopic image to obtain the intraoperative two-dimensional endoscopic image with anatomical feature labels; Based on the anatomical feature identifiers of the two-dimensional rendered image and the anatomical feature identifiers of the intraoperative two-dimensional endoscopic image, the image similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image is calculated.
5. The method according to claim 4, characterized in that, The calculation of the image similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image based on the anatomical feature identifiers of the two-dimensional rendered image and the anatomical feature identifiers of the intraoperative two-dimensional endoscopic image includes: Calculate the global similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image; Based on the anatomical feature identifiers of the two-dimensional rendering image and the anatomical feature identifiers of the intraoperative two-dimensional endoscopic image, the feature similarity between the two-dimensional rendering image and the intraoperative two-dimensional endoscopic image is calculated. The image similarity is calculated based on the global similarity and the feature similarity.
6. The method according to claim 3, characterized in that, The step of performing multi-view virtual sampling on the three-dimensional segmentation model to generate multiple two-dimensional sampling images with anatomical feature markers and corresponding sampling camera poses includes: Based on the radius of the spherical bounding box of the target organ, a regular icosahedron is generated; The camera pose is constructed based on the vertices of the icosahedron, and it is determined whether the recursion requirement is met. When the recursion requirement is not met, each triangular face of the regular icosahedron is divided into 4 triangular faces. The corresponding sampling camera pose is constructed at the vertices of the new polyhedron, and a two-dimensional sampling image of the viewpoint corresponding to each sampling camera pose is generated, resulting in multiple two-dimensional sampling images with anatomical feature labels and corresponding sampling camera poses.
7. The method according to claim 1 or 2, characterized in that, The process of optimizing the target camera pose through gradient backpropagation until the convergence condition is met includes: Using the initial pose as the initial value, a loss function is constructed based on the image similarity. The gradient is calculated through backpropagation of the loss function, and the increment on the Lie algebra se(3) is mapped back to the SE(3) group space using exponential mapping to achieve iterative update of the target camera pose until the convergence condition is met.
8. A medical image registration device, characterized in that, The device includes: The model conversion unit is used to acquire preoperative three-dimensional medical images and convert the preoperative three-dimensional medical images into three-dimensional Gaussian models; the preoperative three-dimensional medical images are images of three-dimensional reconstruction of the target organ before surgery; A two-dimensional rendering unit is used to perform differentiable rendering of the three-dimensional Gaussian model with the initial pose as the target camera pose, to obtain a two-dimensional rendering image. The image acquisition unit is used to acquire intraoperative two-dimensional endoscopic images; the intraoperative two-dimensional endoscopic images are two-dimensional images of the target organ acquired during the operation via endoscopy. The pose matching unit is used to optimize the pose of the target camera through gradient backpropagation based on the image similarity between the two-dimensional rendered image and the intraoperative two-dimensional endoscopic image until the convergence condition is met. The registration unit is used to register the preoperative three-dimensional medical image to the spatial coordinate system where the intraoperative two-dimensional endoscopic image is located, based on the target camera pose when the convergence condition is met.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. An endoscope system, characterized in that, include: Endoscopes, light source devices, image processing devices, display devices, and computer devices as described in claim 9.