Navigation error correction method and device based on three-dimensional reconstruction of bony structure
By using a three-dimensional reconstruction method based on bony structures, and employing a three-dimensional reconstruction neural network and endoscopic video sequences for model reconstruction and registration, the problem of inaccurate reconstruction results caused by monocular video feature matching is solved, achieving high-precision navigation error correction and surgical positioning.
Patent Information
- Application Number
- CN202410314132.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-03-19
AI Technical Summary
Existing 3D reconstruction methods based on feature matching from monocular video are difficult to accurately restore internal structures, resulting in low resolution reconstruction results and easy generation of structural holes or missing parts, which affects the accuracy of navigation information, especially in medical scenarios that require fine structural information.
A method based on 3D reconstruction of bony structures is adopted. Endoscopic video sequences and pose sequences of the surgical site are acquired, and a 3D reconstruction neural network is used to reconstruct the model. The model is then registered with a reference 3D model of bony structures to obtain a pose deviation transformation matrix for navigation error correction. A neural radiation field network or a 3D Gaussian sputtering network is used to complete the data with high spatial resolution and incomplete data.
A high-precision 3D reconstruction model was achieved, which can effectively handle the problems of weak texture and high reflectivity in endoscopic images, capture image detail information, and improve the accuracy of navigation error correction and surgical positioning.
Smart Images

Figure CN118135108B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology, and in particular to a navigation error correction method and device based on three-dimensional reconstruction of bony structures. Background Technology
[0002] Surgical navigation systems process preoperative patient data to create a 3D model. During surgery, the patient and the 3D model are registered. A navigation system tracks the position of surgical instruments, providing the surgeon with virtual surgical instruments and visual information of anatomical structures within the 3D model. This assists the surgeon in performing specific surgical procedures, offering better guidance and support to improve accuracy, safety, and success rates. Therefore, accurately correcting navigation errors to assist surgeons in performing more precise surgical procedures is a crucial issue that the industry urgently needs to address.
[0003] In related technologies, most 3D reconstruction is based on a multi-view approach using feature matching. Specifically, a monocular video is acquired directly using a medical endoscope or ultrasound, and keyframes are determined from the image sequence. Then, the pose parameters of the keyframes are obtained, and the depth map of the selected keyframes is estimated through feature matching. Subsequently, image reconstruction is performed to obtain a 3D point cloud. Finally, a 3D mesh model is constructed from the 3D point cloud to achieve navigation information registration based on the 3D mesh model.
[0004] However, this reconstruction method often relies solely on feature matching from monocular video, making it difficult to accurately reconstruct the internal structure. The reconstruction results have low resolution and are prone to structural holes or missing parts, resulting in unrealistic reconstruction results and inaccurate navigation information. This negatively impacts doctors' ability to make accurate diagnoses, especially in medical scenarios that require detailed structural information. It makes it difficult for doctors to obtain complete information, thus affecting the accuracy of medical diagnoses. Summary of the Invention
[0005] This invention provides a navigation error correction method and device based on 3D reconstruction of bony structures, which solves the defects of existing technologies that rely solely on monocular video for feature matching, resulting in low resolution of reconstruction results and easy generation of structural holes or missing parts, making the reconstruction results less realistic and the navigation information inaccurate. This invention provides a more accurate 3D reconstruction model for precise correction of navigation and positioning errors.
[0006] This invention provides a navigation error correction method based on three-dimensional reconstruction of bony structures, comprising:
[0007] Acquire the current endoscopic video sequence and current endoscopic pose sequence of the surgical site during surgery;
[0008] The target bony structure images of each frame in the current endoscopic video sequence and the current endoscopic pose sequence are input into the three-dimensional reconstruction module to reconstruct the three-dimensional model of the target bony structure corresponding to the current endoscopic video sequence.
[0009] The target bony structure 3D model is registered with the reference bony structure 3D model to obtain the pose deviation transformation matrix, and the navigation error of the navigation system is corrected according to the pose deviation transformation matrix.
[0010] The reference bone structure 3D model is reconstructed based on the reference medical image sequence of the surgical site acquired before surgery; the 3D reconstruction module is obtained by training a 3D reconstruction neural network based on the sample bone structure map of each frame in the sample endoscopic video sequence, the sample endoscopic pose sequence, and the color and depth labels of the sample bone structure map of each frame.
[0011] According to the present invention, a navigation error correction method based on three-dimensional reconstruction of bony structures is provided. In the case that the three-dimensional reconstruction neural network is a neural radiation field network, the three-dimensional reconstruction neural network includes a three-dimensional feature mesh construction module and an encoder.
[0012] The 3D reconstruction module is trained based on the following steps:
[0013] Obtain the bony structure diagram of each frame of the sample endoscope video sequence and the pose sequence of the sample endoscope;
[0014] Color-coded the sample skeletal structure diagrams of each frame to obtain color labels for the sample skeletal structure diagrams of each frame;
[0015] Depth labeling is performed on the sample skeletal structure map of each frame to obtain the depth label of the sample skeletal structure map of each frame;
[0016] The sample ossicular structure map of each frame in the sample endoscope video sequence and the sample endoscope pose sequence are used as sample inputs. The color label and depth label of each frame of the sample ossicular structure map are used as sample labels. The three-dimensional feature mesh construction module and the encoder are jointly iteratively trained.
[0017] The 3D reconstruction module is constructed based on the trained 3D feature mesh construction module.
[0018] According to the present invention, a navigation error correction method based on 3D reconstruction of bony structures is provided, wherein the method uses the sample bony structure map of each frame in the sample endoscope video sequence and the sample endoscope pose sequence as sample input, and uses the color label and depth label of each frame of the sample bony structure map as sample label, and iteratively trains the 3D feature mesh construction module and the encoder in conjunction, including:
[0019] The sample ossicular structure diagrams of each frame in the sample endoscope video sequence and the sample endoscope pose sequence are input into the three-dimensional feature mesh construction module to obtain the three-dimensional model of the sample ossicular structure corresponding to the sample endoscope video sequence.
[0020] The three-dimensional model of the sample bone structure and the pose information corresponding to each frame of the sample bone structure diagram in the sample endoscope pose sequence are input into the encoder to obtain the color estimate and depth estimate of each frame of the sample bone structure diagram;
[0021] Based on the deviation between the color estimate of the sample skeletal structure map in each frame and the color label, a color estimation loss function is constructed;
[0022] Based on the deviation between the depth estimate of the sample skeletal structure map in each frame and the depth label, a depth estimation loss function is constructed;
[0023] Based on the color estimation loss function and the depth estimation loss function, the 3D feature mesh construction module and the encoder are jointly subjected to iterative training.
[0024] According to a navigation error correction method based on three-dimensional reconstruction of bony structures provided by the present invention, the step of performing depth labeling on the sample bony structure images of each frame to obtain depth labels for the sample bony structure images of each frame includes:
[0025] Based on the monocular depth estimation algorithm, depth labels are applied to the sample skeletal structure images in each frame to obtain the depth labels of the sample skeletal structure images in each frame.
[0026] According to the navigation error correction method based on three-dimensional reconstruction of bony structures provided by the present invention, the target bony structure map of each frame is obtained based on the following steps:
[0027] Each frame of the endoscopic image in the current endoscopic video sequence is input into the segmentation model to obtain the target bony structure image of each frame;
[0028] The segmentation model is obtained by training the U-Net model based on the sample endoscopic video sequence and the skeletal structure map of each frame in the sample endoscopic video sequence.
[0029] According to the present invention, a navigation error correction method based on three-dimensional reconstruction of bony structures is provided, wherein the navigation error correction of the navigation system is performed according to the pose deviation transformation matrix, including:
[0030] The pose deviation transformation matrix is fed back to the navigation system so that the navigation system can correct the pose transformation matrix between the endoscope tip and the surgical site based on the pose deviation transformation matrix, the pose transformation matrix between the navigation system and the surgical site, and the pose transformation matrix between the endoscope tip and the navigation system, and correct the position of the endoscope tip based on the corrected pose transformation matrix.
[0031] The present invention also provides a navigation error correction device based on three-dimensional reconstruction of bony structures, comprising:
[0032] The acquisition unit is used to acquire the current endoscopic video sequence and the current endoscopic pose sequence of the surgical site during surgery;
[0033] The reconstruction unit is used to input the target bone structure map of each frame in the current endoscopic video sequence and the current endoscopic pose sequence into the three-dimensional reconstruction module to reconstruct the three-dimensional model of the target bone structure corresponding to the current endoscopic video sequence.
[0034] The correction unit is used to register the target bony structure three-dimensional model with the reference bony structure three-dimensional model to obtain the pose deviation transformation matrix, and to correct the navigation error of the navigation system according to the pose deviation transformation matrix.
[0035] The reference bone structure 3D model is reconstructed based on the reference medical image sequence of the surgical site acquired before surgery; the 3D reconstruction module is obtained by training a 3D reconstruction neural network based on the sample bone structure map of each frame in the sample endoscopic video sequence, the sample endoscopic pose sequence, and the color and depth labels of the sample bone structure map of each frame.
[0036] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the navigation error correction method based on three-dimensional reconstruction of bony structures as described above.
[0037] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the navigation error correction method based on three-dimensional reconstruction of bony structures as described above.
[0038] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the navigation error correction method based on three-dimensional reconstruction of bony structures as described above.
[0039] The present invention provides a navigation error correction method and device based on 3D reconstruction of bony structures. This method trains a 3D reconstruction neural network using the bony structure images of each frame in a sample endoscopic video sequence, the endoscopic pose sequence, and the color and depth labels of each frame's bony structure image. This trains the network to address uncertainties and missing local information in real-world scenarios. The network can better predict the position, features, and confidence of point clouds, effectively handling weak textures and high reflectivity in endoscopic images. It also better captures image details and achieves high spatial resolution, enabling the completion and reconstruction of incomplete data. This results in a precise 3D reconstruction module that accurately reconstructs a 3D model of the bony structure. This module allows for accurate and realistic modeling of the bony structure in endoscopic video sequences, leading to more efficient and accurate navigation error correction, reducing navigation errors, and improving the accuracy of surgical positioning guidance. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0041] Figure 1 This is one of the flowcharts of the navigation error correction method based on three-dimensional reconstruction of bony structures provided by the present invention;
[0042] Figure 2 This is the second flowchart of the navigation error correction method based on three-dimensional reconstruction of bony structures provided by the present invention;
[0043] Figure 3 This is a schematic diagram of the navigation error correction device based on three-dimensional reconstruction of bony structures provided by the present invention;
[0044] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0046] Correction of medical navigation systems typically includes a 3D reconstruction component and an intraoperative image feedback component.
[0047] In the field of 3D reconstruction of monocular endoscopic images, most existing technologies employ a multi-view approach based on feature matching. For example, some researchers have proposed 3D reconstruction methods that involve directly acquiring monocular video using a medical endoscope and determining keyframes from the image sequence; then obtaining the pose parameters of the selected keyframes, estimating the depth map of the selected keyframes through feature matching, and subsequently reconstructing the image to obtain a 3D point cloud. Finally, a 3D mesh model (hereinafter also referred to as a 3D mesh model) is constructed from the 3D point cloud.
[0048] In terms of intraoperative image feedback, most existing technologies utilize intraoperative ultrasound for real-time imaging, navigation, and correction of brain tissue deformation. For example, some researchers have proposed intraoperative navigation correction methods that involve acquiring a series of brain images, such as MRI (Magnetic Resonance Imaging) or CT (Computed Tomography), as reference images before surgery, showing the patient's brain structures and lesions. During the surgery, the surgeon uses intraoperative ultrasound to image the surgical area in real time. The ultrasound probe is placed on the patient's head and generates real-time images by emitting and receiving ultrasound waves. By comparing the intraoperative ultrasound images with the reference images, the surgeon can detect deformations in the brain tissue. These deformations can be caused by brain expansion, gravity, surgical manipulation, or other factors. Once brain tissue deformation is detected, the surgeon can use intraoperative ultrasound to correct it. This is achieved by marking reference points or structures on the ultrasound images, which correspond to the corresponding structures in the reference images. After correction, the navigation system updates the surgical navigation information based on the corrected ultrasound images to assist the surgeon in performing the corresponding surgical procedures.
[0049] The aforementioned navigation correction methods, which rely solely on monocular endoscopic video or ultrasound images for feature matching, often struggle to accurately reconstruct internal structures. The reconstruction results have low resolution and are prone to structural voids or omissions, leading to unrealistic reconstructions. This lack of fidelity can negatively impact doctors' accurate diagnoses, especially in medical scenarios requiring detailed structural information. It makes it difficult for doctors to obtain complete information, thus affecting the accuracy of medical diagnoses.
[0050] To address the aforementioned issues, this embodiment provides a navigation error correction method and device based on three-dimensional reconstruction of bony structures. This method uses image segmentation technology to segment the bony components in endoscopic images. Then, it utilizes a three-dimensional reconstruction neural network, such as a neural radiation field network or a three-dimensional Gaussian sputtering network, to perform real-time three-dimensional reconstruction of the segmented bony image sequence, obtaining a three-dimensional model of the target bony structure. This target bony structure three-dimensional model is then registered with a standard three-dimensional bony structure model reconstructed from preoperative images to correct navigation errors. By leveraging the powerful learning capabilities of neural networks, missing information can be inferred from limited observation data, enabling the completion and reconstruction of incomplete data, thereby improving the fidelity of the three-dimensional reconstruction model and significantly enhancing the three-dimensional reconstruction effect. This, in turn, improves the accuracy of navigation correction, assisting surgeons in performing more precise surgical operations.
[0051] Most existing navigation and positioning methods complete navigation and positioning during the navigation preparation stage. Few studies consider using intraoperative endoscopic images to process navigation errors as feedback information. The method presented in this application, drawing on the approach of using real-time ultrasound images for intraoperative lesion information feedback, is the first to propose a method for correcting navigation errors by matching the non-deformation characteristics of the three-dimensional bony structures captured intraoperatively with the bony structures in preoperative baseline medical images. In this method, the quality of the three-dimensional reconstruction of the endoscopic images directly affects the effectiveness of the correction scheme. The following section describes the innovative advantages of the method presented in this application, addressing the limitations of existing methods in terms of three-dimensional reconstruction technology.
[0052] Figure 1 This is one of the flowcharts of the navigation error correction method based on three-dimensional reconstruction of bony structures provided by the present invention; the method is applied in the field of assistive medical care to assist doctors in performing more precise surgical operations.
[0053] like Figure 1 As shown, the method includes:
[0054] Step 110: Acquire the current endoscopic video sequence and the current endoscopic pose sequence of the surgical site during the operation;
[0055] Optionally, during the surgery, the current endoscopic video sequence of the surgical site and the current endoscopic pose sequence can be acquired in real time using an endoscope.
[0056] It should be noted that after acquiring the current endoscopic video sequence and the current endoscopic pose sequence, the current endoscopic video sequence can be preprocessed before performing the following 3D model reconstruction operation. This improves the data quality of the current endoscopic video sequence, enabling more accurate and efficient navigation error correction in subsequent operations. The preprocessing includes, but is not limited to, noise reduction and image enhancement, such as sharpening and smoothing. This embodiment does not specifically limit the preprocessing method.
[0057] The surgical site refers to the area where the surgical operation is to be performed; it can be the bottom of the saddle or the forehead, etc. This embodiment does not specifically limit it. The following describes the navigation error correction method provided in this embodiment with the bottom of the saddle as the surgical site.
[0058] The so-called current endoscopic video sequence is the real-time video sequence acquired at the surgical site during the current cycle of the endoscope; the so-called endoscopic pose sequence is the pose sequence formed by the pose information of each frame of the image acquired at the surgical site during the current cycle of the endoscope.
[0059] Step 120: Input the target bone structure map of each frame in the current endoscopy video sequence and the current endoscopy pose sequence into the three-dimensional reconstruction module to reconstruct a three-dimensional model of the target bone structure corresponding to the current endoscopy video sequence; the three-dimensional reconstruction module is obtained by training a three-dimensional reconstruction neural network based on the sample bone structure map of each frame in the sample endoscopy video sequence, the sample endoscopy pose sequence, and the color and depth labels of the sample bone structure map of each frame;
[0060] The 3D reconstruction neural network here can be a neural radiation field network, or a 3D Gaussian sputtering network, or other neural networks that can be used for 3D model reconstruction. This embodiment does not specifically limit it.
[0061] The so-called neural radiation field network is a research hotspot in the field of 3D (Three Dimensional) vision. By training on discrete multi-view images, neural radiation fields can generate new perspective images with extremely high realism. This technological breakthrough not only contributes to the field of 3D vision but also has broad application prospects in 3D model reconstruction, and is expected to overcome the problems of weak texture and high reflectivity that are difficult to handle with current technology. Neural radiation field networks have the ability to model complex geometries and details, including surfaces, edges, and textures. One of its advantages is its continuous representation, which enables it to achieve high spatial resolution, thus providing a more realistic visual effect. Another advantage is that, with the powerful learning ability of neural networks, neural radiation field networks can infer missing information from limited observation data, achieving the completion and reconstruction of incomplete data, thereby significantly improving the 3D reconstruction effect. Therefore, when the 3D reconstruction neural network is a neural radiation field network, the 3D reconstruction module is trained based on the neural radiation field network. It can achieve high spatial resolution and model the modeling advantages of incomplete data completion and reconstruction by utilizing the neural radiation field network, thus enabling accurate modeling of endoscopic video sequences and efficient and accurate navigation error correction.
[0062] The so-called 3D Gaussian sputtering network is a technique that has emerged in recent years in the fields of explicit radiation fields and computer graphics, and is one of the latest 3D reconstruction network methods. It first uses depth maps and camera pose to calculate a 3D point cloud, converts each point cloud into a Gaussian distribution, and projects the points onto the image plane. Next, the point cloud is input into the Gaussian network for training to obtain Gaussian parameters, followed by differentiable rasterization and rendering to obtain an image. After obtaining the rendered image, it is compared with a real endoscope image to calculate the loss value, update the parameters in the Gaussian network, and perform adaptive density control to update the point cloud. Since each Gaussian point cloud possesses position, size, and optimized color and opacity parameters, when this information is combined, a complete 3D model of the endoscope can be predicted and rendered from any angle, effectively filling in holes and achieving the completion and reconstruction of incomplete data, thus significantly improving the 3D reconstruction effect.
[0063] Therefore, when the 3D reconstruction neural network is a 3D Gaussian sputtering network, the 3D reconstruction module is trained based on this network. Utilizing this network, it can predict and render a complete 3D model of the endoscope from any given angle, effectively filling in gaps and achieving accurate modeling of endoscopic video sequences. This allows for efficient and precise navigation error correction. Therefore, before executing step 120, this embodiment can pre-construct a 3D reconstruction module capable of accurately reconstructing the 3D model using other neural networks suitable for 3D model reconstruction, such as neural radiation field networks or 3D Gaussian sputtering networks. The specific construction steps are as follows:
[0064] First, sample endoscope video sequences and sample endoscope pose sequences are acquired, and skeletal structure maps, color labels, and depth labels are extracted from the sample endoscope video sequences to construct a sample dataset.
[0065] Secondly, based on the sample dataset, the 3D reconstruction neural network is iteratively trained to obtain a 3D reconstruction module that can cope with the uncertainties and missing local information in real-world scenarios. This module can better use neural networks to predict the position, features, and confidence of point clouds, effectively handle the weak texture and high reflectivity problems that may exist in endoscopic images, better capture image detail information, and achieve high spatial resolution and completion and reconstruction of incomplete data, so as to accurately reconstruct the 3D model of the bony structure.
[0066] The optimization here can be performed by using an optimization algorithm to perform a global search optimization based on the sample dataset to obtain the 3D reconstruction module; or it can be performed by first determining the initial parameters based on the optimization algorithm, and then optimizing the model parameters based on the initial parameters based on the sample dataset to obtain the 3D reconstruction module. The specific optimization steps of the 3D reconstruction module are not specifically limited here.
[0067] Optionally, after obtaining the pre-3D reconstruction module through iterative optimization, the target bone structure map of each frame in the current endoscopic video sequence and the current endoscopic pose sequence can be input into the 3D reconstruction module to reconstruct the 3D model of the current endoscopic video sequence, thereby outputting the 3D model of the target bone structure corresponding to the current endoscopic video sequence.
[0068] Step 130: Register the target bony structure 3D model with the reference bony structure 3D model to obtain the pose deviation transformation matrix, and correct the navigation error of the navigation system according to the pose deviation transformation matrix; wherein, the reference bony structure 3D model is reconstructed from the reference medical image sequence of the surgical site acquired before surgery.
[0069] The baseline three-dimensional model of the bony structure here is obtained by calibration and reconstruction using the preoperative calibration module based on the baseline medical image sequence of the surgical site acquired before surgery.
[0070] Figure 2 This is the second flowchart illustrating the navigation error correction method based on three-dimensional reconstruction of bony structures provided by the present invention; as shown below. Figure 2 As shown, after obtaining the three-dimensional model of the target bony structure under endoscopy through the three-dimensional reconstruction module, the intraoperative registration module can be used to register the three-dimensional model of the target bony structure with the reference three-dimensional model of the bony structure. By matching the point cloud features or surface features between the two models, the pose deviation transformation matrix of the point cloud between the three-dimensional model of the target bony structure and the reference three-dimensional model of the bony structure under the navigation control of the navigation system is calculated, thereby determining the navigation error of the navigation system.
[0071] Next, the obtained pose deviation transformation matrix is fed back to the navigation system so that the navigation system can continuously correct navigation errors based on the pose deviation transformation matrix and the transformation matrix between the navigation system and the surgical site until the navigation error is completely eliminated, so as to finally correct the real-time end-effector pose.
[0072] The method provided in this embodiment trains a 3D reconstruction neural network using the sample endoscopic video sequence's bone structure map, endoscopic pose sequence, and color and depth labels of each frame's bone structure map. This trains the network to address uncertainties and missing local information in real-world scenarios, enabling better prediction of point cloud positions, features, and confidence levels. It effectively handles weak textures and high reflectivity issues in endoscopic images, captures image details, achieves high spatial resolution, and completes and reconstructs incomplete data. This results in a precise 3D reconstruction module that accurately reconstructs a 3D model of the bone structure. This module enables accurate and realistic modeling of the bone structure in endoscopic video sequences, allowing for more efficient and accurate navigation error correction, reducing navigation errors, and improving the accuracy of surgical positioning guidance.
[0073] In some embodiments, the navigation error correction method provided in this embodiment is described below using a three-dimensional reconstruction neural network as an example of a neural radiation field network.
[0074] In the case where the three-dimensional reconstruction neural network is a neural radiation field network, the three-dimensional reconstruction neural network includes a three-dimensional feature mesh construction module and an encoder;
[0075] The 3D reconstruction module is trained based on the following steps:
[0076] Obtain the bony structure diagram of each frame of the sample endoscope video sequence and the pose sequence of the sample endoscope;
[0077] Color-coded the sample skeletal structure diagrams of each frame to obtain color labels for the sample skeletal structure diagrams of each frame;
[0078] Depth labeling is performed on the sample skeletal structure map of each frame to obtain the depth label of the sample skeletal structure map of each frame;
[0079] The sample ossicular structure map of each frame in the sample endoscope video sequence and the sample endoscope pose sequence are used as sample inputs. The color label and depth label of each frame of the sample ossicular structure map are used as sample labels. The three-dimensional feature mesh construction module and the encoder are jointly iteratively trained.
[0080] The 3D reconstruction module is constructed based on the trained 3D feature mesh construction module.
[0081] Optionally, the neural radiation field network includes at least a three-dimensional feature mesh construction module and an encoder; the three-dimensional feature mesh construction module is used to structure objects or scenes in the scene into a three-dimensional mesh form for three-dimensional scene model reconstruction, and the encoder is used to capture point cloud features, including but not limited to color and depth information of the skeletal structure map.
[0082] Optionally, the specific training steps for the 3D reconstruction module include:
[0083] First, to improve the accuracy of 3D reconstruction, it is necessary to acquire sample endoscopic video sequences and sample endoscopic pose sequences from various surgical sites and under different endoscopic poses. Then, bone structure maps are extracted from the sample endoscopic video sequences, along with color and depth labels. This ensures that the sample dataset, constructed using the sample bone structure maps and sample endoscopic pose sequences from each frame of the sample endoscopic video sequence as input, and the color and depth labels of each frame's sample bone structure map as sample labels, possesses sufficient depth and breadth. This allows the trained 3D reconstruction module to accurately reconstruct 3D models of bone structures in different scenarios.
[0084] The endoscopic video sequences and endoscopic pose sequences of different surgical sites and different endoscopic poses can be obtained from historical acquisitions or loaded from open-source databases.
[0085] Next, based on the sample dataset, the 3D feature mesh construction module and the encoder are jointly trained iteratively so that the trained 3D feature mesh construction module can accurately reconstruct the 3D model.
[0086] Finally, a 3D reconstruction module is constructed based on the trained 3D feature mesh construction module.
[0087] In some embodiments, the step of obtaining the sample bony structure map of each frame in the sample endoscopic video sequence further includes:
[0088] Image segmentation was performed on each frame of the endoscopic video sequence of the sample to separate the bony structure from each frame of the endoscopic video sequence of the sample, thereby obtaining the bony structure map of each frame of the sample.
[0089] In some embodiments, the step of depth-marking the sample bony structure map of each frame further includes:
[0090] Based on the monocular depth estimation algorithm, depth labels are applied to the sample skeletal structure images in each frame to obtain the depth labels of the sample skeletal structure images in each frame.
[0091] Optionally, a monocular depth estimation algorithm is used to estimate the depth map of the skeletal structure map of each frame sample in order to obtain the depth label of the skeletal structure map of each frame sample.
[0092] The so-called monocular depth estimation algorithm can be implemented by learning the mapping relationship between the input image and the corresponding depth map through a deep learning model.
[0093] In some embodiments, the step of jointly iteratively training the 3D feature mesh construction module and the encoder further includes:
[0094] The sample ossicular structure diagrams of each frame in the sample endoscope video sequence and the sample endoscope pose sequence are input into the three-dimensional feature mesh construction module to obtain the three-dimensional model of the sample ossicular structure corresponding to the sample endoscope video sequence.
[0095] The three-dimensional model of the sample bone structure and the pose information corresponding to each frame of the sample bone structure diagram in the sample endoscope pose sequence are input into the encoder to obtain the color estimate and depth estimate of each frame of the sample bone structure diagram;
[0096] Based on the deviation between the color estimate of the sample skeletal structure map in each frame and the color label, a color estimation loss function is constructed;
[0097] Based on the deviation between the depth estimate of the sample skeletal structure map in each frame and the depth label, a depth estimation loss function is constructed;
[0098] Based on the color estimation loss function and the depth estimation loss function, the 3D feature mesh construction module and the encoder are jointly subjected to iterative training.
[0099] Optionally, the iterative training steps specifically include:
[0100] The skeletal structure map of each frame of the segmented endoscopic video sequence and the pose sequence of the endoscopic sample are input into the 3D feature mesh construction module. Based on the pose information of the endoscope at different time points and the skeletal structure map of each frame of the sample collected at different time points, the 3D feature mesh construction module reconstructs the 3D model of the skeletal structure of the sample corresponding to the endoscopic video sequence.
[0101] Subsequently, the three-dimensional model of the sample bone structure corresponding to the sample endoscope video sequence and the pose information corresponding to the sample bone structure map of each frame in the sample endoscope pose sequence are input into the encoder so that the encoder can perform point cloud feature estimation to obtain the color estimate and depth estimate of the sample bone structure map of each frame under different poses.
[0102] Subsequently, the color estimates of the skeletal structure images of each frame sample under different poses are compared with the color labels to obtain the deviation between the color estimates and the color labels, thereby constructing a color estimation loss function.
[0103] Furthermore, the depth estimates of the skeletal structure images of each frame sample under different poses are compared with the depth labels to obtain the deviation between the depth estimates and the depth labels, thereby constructing a depth estimation loss function.
[0104] Subsequently, the color estimation loss function and the depth estimation loss function are fused to obtain the final target loss function. The target loss function is then backpropagated along the differentiable encoder and the 3D feature mesh construction module. Through continuous optimization, the rendering effect and 3D modeling performance are improved, ultimately resulting in a 3D reconstruction module that can accurately construct 3D modules, thereby achieving accurate reconstruction of the skeletal structure 3D model.
[0105] The method provided in this embodiment achieves high-fidelity endoscopic 3D reconstruction by training a 3D reconstruction module through a neural radiation field network. On the one hand, since the neural radiation field is a continuous representation, it can achieve high spatial resolution and complete and reconstruct incomplete data, thus providing a more realistic 3D model reconstruction. On the other hand, the point cloud generator (i.e., encoder) uses a neural network to predict the position, features, and confidence of the point cloud, which can effectively handle the weak texture and high reflectivity problems that may exist in endoscopic images. Moreover, the continuous representation of the point cloud helps to better capture detailed information, thus providing a more accurate 3D model reconstruction.
[0106] In some embodiments, the target bony structure map in each frame of step 120 is obtained based on the following steps:
[0107] Each frame of the endoscopic image in the current endoscopic video sequence is input into the segmentation model to obtain the target bony structure image of each frame;
[0108] The segmentation model is obtained by training the U-Net model based on the sample endoscopic video sequence and the skeletal structure map of each frame in the sample endoscopic video sequence.
[0109] Optionally, when acquiring the bone structure map, the real-time video stream is processed using libraries such as OpenCV (Open Source Computer Vision) in the real-time video processing system. Each frame of the image is input into a segmentation model constructed by a trained U-Net model, so that the segmentation model can segment the bone structure map from each frame of the endoscopic image to obtain the target bone structure map for each frame.
[0110] Here, the training steps for the segmentation model include: collecting a dataset containing sample endoscopic video sequences and their corresponding annotations (i.e., labels of the bony structure regions to be segmented in the surgical images) before surgery; building a U-Net (Convolutional Networks for Biomedical Image Segmentation) model using a deep learning framework; and training the U-Net model using the dataset to enable it to accurately segment bony structure maps in endoscopic video sequences.
[0111] The method provided in this embodiment uses a convolutional neural network to segment bone structure images, which can accurately and efficiently mine and segment the bone structure images of each frame in the endoscopic video sequence, thereby improving the accuracy of three-dimensional reconstruction and thus improving the accuracy of navigation error correction.
[0112] In some embodiments, step 130, which corrects the navigation error of the navigation system based on the pose deviation transformation matrix, includes:
[0113] The pose deviation transformation matrix is fed back to the navigation system so that the navigation system can correct the pose transformation matrix between the endoscope tip and the surgical site based on the pose deviation transformation matrix, the pose transformation matrix between the navigation system and the surgical site, and the pose transformation matrix between the endoscope tip and the navigation system, and correct the position of the endoscope tip based on the corrected pose transformation matrix.
[0114] Optionally, the implementation of navigation error correction in step 130 based on the pose deviation transformation matrix fed back from the intraoperative three-dimensional reconstruction of bony structures aims to correct the pose error of the real-time endoscope tip (relative to the surgical site). This error may be generated by multiple factors. To simplify calculations, this embodiment assumes that the pose transformation matrix between the endoscope tip and the navigation system provided by the intraoperative tracking module is an accurate value. Therefore, the error mainly originates from the pose transformation matrix between the navigation system and the surgical site, which is calculated jointly by the preoperative calibration module and the preoperative registration module. Thus, the bony structure model obtained through intraoperative three-dimensional reconstruction of the bony structure can be registered with the bony structure model obtained from preoperative image reconstruction to obtain the pose deviation transformation matrix. This matrix is then fed back to the pose transformation matrix between the navigation system and the surgical site to ultimately correct the real-time endoscope tip pose, thereby achieving navigation error correction. The specific implementation steps are as follows:
[0115] Based on the preoperative calibration module and the preoperative registration module, the pose transformation matrix between the navigation system and the surgical site is calculated and labeled as H0.
[0116] The pose transformation matrix between the endoscope tip and the navigation system is calculated based on the intraoperative tracking module and labeled as H1;
[0117] Based on the intraoperative registration module, the pose deviation transformation matrix obtained by registering the actual point cloud to the estimated point cloud in the coordinate system at the end of the endoscope is calculated and denoted as H2.
[0118] Based on the pose deviation transformation matrix H2, pose transformation matrix H0, and pose transformation matrix H1, the pose transformation matrix between the endoscope tip and the surgical site is jointly corrected to obtain the corrected pose transformation matrix between the endoscope tip and the surgical site.
[0119] The correction here can be achieved based on the following method:
[0120] The pose transformation matrix H1 between the endoscope tip and the navigation system is combined with the pose transformation matrix H0 between the navigation system and the surgical site to obtain the initial pose transformation matrix H_init = H1 * H0 between the endoscope tip and the surgical site. The pose deviation transformation matrix H2 is then combined with the initial pose transformation matrix between the endoscope tip and the surgical site to obtain the corrected pose transformation matrix H_corrected = H_2 * H_init between the endoscope tip and the surgical site.
[0121] Then, based on the corrected pose transformation matrix, the position of the endoscope tip is corrected so that the intraoperative navigation module can correct the navigation information in real time, thereby eliminating navigation errors.
[0122] The method provided in this embodiment uses a pose deviation transformation matrix to provide feedback to the navigation system, and combines the pose transformation matrix between the navigation system and the surgical site, as well as the pose transformation matrix between the endoscope tip and the navigation system, to continuously correct the pose transformation matrix between the endoscope tip and the surgical site. This achieves the correction of the endoscope tip position, compensates for the error in the navigation system, breaks through the limit of preoperative registration and calibration accuracy, and thus improves the precision and accuracy of intraoperative navigation.
[0123] The navigation error correction device based on three-dimensional reconstruction of bony structure provided by the present invention will be described below. The navigation error correction device based on three-dimensional reconstruction of bony structure described below can be referred to in correspondence with the navigation error correction method based on three-dimensional reconstruction of bony structure described above.
[0124] Figure 3 This is a schematic diagram of the navigation error correction device based on three-dimensional reconstruction of bony structures provided by the present invention; as shown. Figure 3 As shown, the device includes:
[0125] The acquisition unit 310 is used to acquire the current endoscopic video sequence and the current endoscopic pose sequence of the surgical site during surgery;
[0126] The reconstruction unit 320 is used to input the target bone structure map of each frame in the current endoscope video sequence and the current endoscope pose sequence into the three-dimensional reconstruction module to reconstruct the three-dimensional model of the target bone structure corresponding to the current endoscope video sequence;
[0127] The correction unit 330 is used to register the target bony structure three-dimensional model with the reference bony structure three-dimensional model to obtain the pose deviation transformation matrix, and to correct the navigation error of the navigation system according to the pose deviation transformation matrix.
[0128] The reference bone structure 3D model is reconstructed based on the reference medical image sequence of the surgical site acquired before surgery; the 3D reconstruction module is obtained by training a 3D reconstruction neural network based on the sample bone structure map of each frame in the sample endoscopic video sequence, the sample endoscopic pose sequence, and the color and depth labels of the sample bone structure map of each frame.
[0129] The device provided by this invention trains a three-dimensional reconstruction neural network using the sample endoscopic video sequence's bone structure map, endoscopic pose sequence, and color and depth labels of each frame's bone structure map. This allows the network to better handle uncertainties and missing local information in real-world scenarios, predict the position, features, and confidence of point clouds, effectively address weak texture and high reflectivity issues in endoscopic images, capture image details, achieve high spatial resolution, and complete and reconstruct incomplete data. This results in a precise three-dimensional reconstruction module that accurately reconstructs a three-dimensional model of the bone structure. This module enables accurate and realistic modeling of the bone structure from endoscopic video sequences, allowing for more efficient and accurate navigation error correction, reducing navigation errors, and improving the accuracy of surgical positioning guidance.
[0130] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: processor 410, communication interface 420, memory 430 and communication bus 440, wherein the processor 810, communication interface 420 and memory 430 communicate with each other through communication bus 440. The processor 410 can call logic instructions in the memory 430 to execute the navigation error correction method based on three-dimensional reconstruction of bony structures provided in the above embodiments. The method includes: acquiring the current endoscopic video sequence and the current endoscopic pose sequence of the surgical site during surgery; inputting the target bony structure map of each frame in the current endoscopic video sequence and the current endoscopic pose sequence to the three-dimensional reconstruction module to reconstruct the target bony structure three-dimensional model corresponding to the current endoscopic video sequence; registering the target bony structure three-dimensional model with the reference bony structure three-dimensional model to obtain the pose deviation transformation matrix, and correcting the navigation error of the navigation system according to the pose deviation transformation matrix; wherein, the reference bony structure three-dimensional model is reconstructed based on the reference medical image sequence of the surgical site acquired before surgery; the three-dimensional reconstruction module is trained on the three-dimensional reconstruction neural network based on the sample bony structure map of each frame in the sample endoscopic video sequence, the sample endoscopic pose sequence, and the color and depth labels of each frame of the sample bony structure map.
[0131] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0132] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the navigation error correction method based on three-dimensional reconstruction of bony structures provided by the above methods. This method includes: acquiring the current endoscopic video sequence and the current endoscopic pose sequence of the surgical site during surgery; inputting the target bony structure map of each frame in the current endoscopic video sequence and the current endoscopic pose sequence to a three-dimensional reconstruction module to reconstruct the current endoscopic video sequence pair. The target bony structure 3D model is obtained; the target bony structure 3D model is registered with the reference bony structure 3D model to obtain the pose deviation transformation matrix, and the navigation error of the navigation system is corrected according to the pose deviation transformation matrix; wherein, the reference bony structure 3D model is reconstructed from the reference medical image sequence of the surgical site acquired before surgery; the 3D reconstruction module is obtained by training the 3D reconstruction neural network based on the sample bony structure map of each frame in the sample endoscopic video sequence, the sample endoscopic pose sequence, and the color and depth labels of each frame of the sample bony structure map.
[0133] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the navigation error correction method based on three-dimensional reconstruction of bony structures provided by the above methods. The method includes: acquiring a current endoscopic video sequence and a current endoscopic pose sequence of the surgical site during surgery; inputting each frame of the target bony structure image in the current endoscopic video sequence and the current endoscopic pose sequence into a three-dimensional reconstruction module to reconstruct a three-dimensional model of the target bony structure corresponding to the current endoscopic video sequence; registering the target bony structure three-dimensional model with a reference bony structure three-dimensional model to obtain a pose deviation transformation matrix, and correcting the navigation error of the navigation system according to the pose deviation transformation matrix; wherein the reference bony structure three-dimensional model is reconstructed based on a reference medical image sequence of the surgical site acquired before surgery; the three-dimensional reconstruction module is trained on a three-dimensional reconstruction neural network based on each frame of the sample bony structure image, the sample endoscopic pose sequence, and the color and depth labels of each frame of the sample bony structure image in the sample endoscopic video sequence.
[0134] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0135] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A navigation error correction method based on three-dimensional reconstruction of bony structures, characterized in that, The method comprises the following steps: Collecting a current endoscope video sequence and a current endoscope pose sequence of a surgical site during surgery, inputting each frame of endoscope image in the current endoscope video sequence into a segmentation model respectively, and segmenting to obtain a target bony structure image of each frame; Inputting each frame of target bony structure image in the current endoscope video sequence and the current endoscope pose sequence into a three-dimensional reconstruction module to reconstruct a target bony structure three-dimensional model corresponding to the current endoscope video sequence; Registering the target bony structure three-dimensional model with a reference bony structure three-dimensional model to obtain a pose deviation transformation matrix, and correcting the navigation error of a navigation system according to the pose deviation transformation matrix; Wherein, the reference bony structure three-dimensional model is reconstructed according to a reference medical image sequence of the surgical site collected before surgery; the three-dimensional reconstruction module is trained from a three-dimensional reconstruction neural network; The three-dimensional reconstruction neural network is a neural radiation field network, and the three-dimensional reconstruction neural network comprises a three-dimensional feature grid construction module and an encoder; The three-dimensional reconstruction module is trained based on the following steps: Obtaining each frame of sample bony structure image in a sample endoscope video sequence and the sample endoscope pose sequence; Color labeling each frame of the sample bony structure image to obtain a color label of each frame of the sample bony structure image; Depth labeling each frame of the sample bony structure image to obtain a depth label of each frame of the sample bony structure image; Inputting each frame of sample bony structure image in the sample endoscope video sequence and the sample endoscope pose sequence as a sample, inputting the color label and the depth label of each frame of the sample bony structure image as a sample label, and iteratively training the three-dimensional feature grid construction module and the encoder jointly; According to the trained three-dimensional feature grid construction module, the three-dimensional reconstruction module is constructed.
2. The method of claim 1, wherein, The inputting each frame of sample bony structure image in the sample endoscope video sequence and the sample endoscope pose sequence as a sample, inputting the color label and the depth label of each frame of the sample bony structure image as a sample label, and iteratively training the three-dimensional feature grid construction module and the encoder jointly comprises: Inputting each frame of sample bony structure image in the sample endoscope video sequence and the sample endoscope pose sequence into the three-dimensional feature grid construction module to obtain a sample bony structure three-dimensional model corresponding to the sample endoscope video sequence; Inputting the sample bony structure three-dimensional model and the pose information corresponding to each frame of the sample bony structure image in the sample endoscope pose sequence into the encoder to obtain a color estimation value and a depth estimation value of each frame of the sample bony structure image; According to the deviation between the color estimation value and the color label of each frame of the sample bony structure image, a color estimation loss function is constructed; According to the deviation between the depth estimation value and the depth label of each frame of the sample bony structure image, a depth estimation loss function is constructed; According to the color estimation loss function and the depth estimation loss function, the three-dimensional feature grid construction module and the encoder are iteratively trained.
3. The method of claim 1, wherein, The depth labeling of each frame of the sample bone structure graph comprises: The depth labeling of each frame of the sample bone structure graph is performed based on a monocular depth estimation algorithm, to obtain the depth label of each frame of the sample bone structure graph.
4. The method of claim 1-3, wherein, The segmentation model is obtained by training a U-Net model based on the sample endoscopic video sequence and each frame of the sample bone structure graph in the sample endoscopic video sequence.
5. The method of claim 1-3, wherein, The error correction of the navigation error of the navigation system according to the pose deviation transformation matrix comprises: The pose deviation transformation matrix is fed back to the navigation system, so that the navigation system corrects the pose transformation matrix between the endoscope tip and the surgical site according to the pose deviation transformation matrix, a pose transformation matrix between the navigation system and the surgical site, and a pose transformation matrix between the endoscope tip and the navigation system, and corrects the position of the endoscope tip according to the corrected pose transformation matrix.
6. A navigation error correction device based on three-dimensional reconstruction of bony structures, characterized by, Comprise: The acquisition unit is configured to acquire a current endoscopic video sequence and a current endoscopic pose sequence of a surgical site during surgery, and input each frame of endoscopic image in the current endoscopic video sequence into a segmentation model respectively to obtain each frame of target bone structure graph by segmentation; The reconstruction unit is configured to input each frame of target bone structure graph in the current endoscopic video sequence and the current endoscopic pose sequence into a three-dimensional reconstruction module to reconstruct a target bone structure three-dimensional model corresponding to the current endoscopic video sequence; The correction unit is configured to register the target bone structure three-dimensional model with a reference bone structure three-dimensional model to obtain a pose deviation transformation matrix, and correct the navigation error of the navigation system according to the pose deviation transformation matrix; The reference bone structure three-dimensional model is reconstructed according to a reference medical image sequence of the surgical site acquired before surgery; and the three-dimensional reconstruction module is trained from a three-dimensional reconstruction neural network; The three-dimensional reconstruction neural network is a neural radiation field network, and the three-dimensional reconstruction neural network comprises a three-dimensional feature grid construction module and an encoder; The training unit is specifically configured to: Obtain each frame of sample bone structure graph in a sample endoscopic video sequence and the sample endoscopic pose sequence; Color label each frame of the sample bone structure graph to obtain a color label of each frame of the sample bone structure graph; Depth label each frame of the sample bone structure graph to obtain a depth label of each frame of the sample bone structure graph; Input each frame of sample bone structure graph in the sample endoscopic video sequence and the sample endoscopic pose sequence as a sample, and input the color label and the depth label of each frame of the sample bone structure graph as a sample label, and iteratively train the three-dimensional feature grid construction module and the encoder jointly; According to the trained three-dimensional feature grid construction module, the three-dimensional reconstruction module is constructed.
7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the bone structure three-dimensional reconstruction-based navigation error correction method according to any one of claims 1 to 5 when executing the program.
8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the navigation error correction method based on three-dimensional reconstruction of bony structures according to any one of claims 1 to 5.
9. A computer program product comprising a computer program, characterized in that, The computer program is executed by a processor to implement the navigation error correction method based on three-dimensional reconstruction of bony structures according to any one of claims 1 to 5.
Citation Information
Patent Citations
Navigation method, device and electronic equipment for laparoscopic augmented reality surgery
CN113143459A
Monocular endoscope new view angle image generation method based on deformable nerve radiation field
CN117392312A