Endoscope navigation method, device, equipment and medium
By acquiring and matching images of the target 3D model and endoscopic images, the problem of narrow field of view and lack of depth information in minimally invasive surgery with flexible endoscopes is solved, enabling accurate positioning and navigation of the endoscope and improving the precision and efficiency of the surgery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU YUANHE MEDICAL TECHNOLOGY CO LTD
- Filing Date
- 2025-11-28
- Publication Date
- 2026-04-17
AI Technical Summary
Existing flexible endoscopes in minimally invasive surgery suffer from a narrow field of view, provide only a two-dimensional image, and lack depth information, which can easily lead to spatial disorientation and distance perception errors for surgeons during operation. Furthermore, navigation solutions that rely on additional hardware increase complexity and cost.
By acquiring the target object's 3D model and endoscopic images, the endoscopic images are acquired in real time using a visual camera, and image detection and matching are performed to determine the position of the endoscope. Navigation is then performed in conjunction with the 3D reconstruction model.
It enables accurate positioning and navigation of the endoscope, ensuring navigation accuracy, avoiding additional hardware impacts on the size and reliability of the endoscope, and improving the precision and efficiency of the surgery.
Smart Images

Figure CN121867942A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to an endoscopic navigation method, apparatus, device, and medium. Background Technology
[0002] Flexible endoscopes (such as uroscopes, bronchoscopes, and cystoscopes) are widely used in minimally invasive surgery. However, because these endoscopes have a narrow field of view and provide only a two-dimensional image with a lack of depth information, surgeons are prone to spatial disorientation and distance perception errors during operation.
[0003] Real-time acquisition of the three-dimensional structure of the surgical area provides surgeons with valuable spatial positioning references. For example, real-time 3D reconstruction can be used to accurately measure the distance between two points within the body, aiding in determining the relative positions of surgical instruments and tissues. However, achieving real-time 3D reconstruction and navigation in flexible endoscopic surgery is no easy feat. Even with successful 3D reconstruction in flexible endoscopic surgery, the complexity of internal tissue cavities means that surgeons may still be unable to pinpoint the exact location of the endoscope after obtaining the 3D reconstruction model. Therefore, to achieve accurate positioning, some existing endoscopic navigation solutions rely on additional hardware, such as electromagnetic locators. However, these solutions require additional hardware within the endoscope, impacting its size and reliability. Summary of the Invention
[0004] Therefore, it is necessary to provide an endoscopic navigation method, device, equipment, and medium to address the above problems, which enables endoscopic navigation through a visual camera and ensures navigation accuracy.
[0005] This application provides an endoscopic navigation method, the method comprising: Obtain a target 3D model of the target object; the target 3D model is obtained based on target images from at least two orientations; the target images are obtained by scanning the target area of the target object; Acquire endoscopic images of the target object; the endoscopic images are acquired in real time by a visual camera on the endoscope when the endoscope is inserted into the target area; Image detection is performed on the endoscopic image to obtain the target tissue image in the endoscopic image; The target tissue image is compared with the target 3D model to determine the target location of the endoscope, so as to navigate based on the target location.
[0006] In one alternative implementation, multiple tissues to be detected exist in the target area; The step of comparing the target tissue image with the target 3D model to determine the target location of the endoscope includes: Based on the target 3D model, a simulated view library is obtained; the simulated view library includes simulated images of each tissue to be detected from multiple perspectives. The target tissue image is matched with the simulated images in the simulated view library to determine the target simulated image; The target location of the endoscope is determined based on the position of the tissue to be detected in the target 3D model of the target simulation image and the viewpoint corresponding to the target simulation image.
[0007] In one optional implementation, the navigation based on the target location includes: The image display interface shows the exploration status of each tissue to be detected; the exploration status indicates whether the tissue to be detected has been explored or not. The method further includes: Based on the endoscopic images of the target object, it is determined whether the endoscope has entered the target tissue in each tissue to be detected; If the target organization is entered, the exploration status of the target organization will be updated to "explored".
[0008] In one optional implementation, the navigation based on the target location includes: Three-dimensional reconstruction is performed based on the endoscopic images to obtain a three-dimensional reconstruction model of the target region; the three-dimensional reconstruction model is used to indicate path information in the target region; Based on the target location, the 3D reconstruction model is registered with the target 3D model to indicate the path to each of the tissues to be detected.
[0009] In one optional implementation, the step of performing three-dimensional reconstruction based on the endoscopic image to obtain a three-dimensional reconstruction model of the target region includes: Obtain the view of the i-th window; the i-th window view includes several consecutive frames of endoscopic images within the i-th time window; where i is any positive integer from 1 to N, and N≥1; Input the i-th window view into the 3D reconstruction network to obtain the global point cloud corresponding to the i-th window view; The global point cloud of each window view is transformed to a specified coordinate system and fused to obtain a three-dimensional reconstruction model of the target area.
[0010] In one alternative implementation, the specified coordinate system is the coordinate system of the global map for the first time window.
[0011] In an alternative implementation, before inputting the i-th window view into the 3D reconstruction network, the method further includes: Obtain a training sample set; the training sample set includes several training samples; each training sample includes several consecutive sample images, global point cloud ground truth, and local point cloud ground truth corresponding to each sample image; For each training sample, the continuous sample image is input into the 3D reconstruction network to obtain the global point cloud prediction value and the local point cloud prediction value corresponding to each frame of sample image; A first error is generated based on the predicted global point cloud value and the true global point cloud value. For each frame of sample image, a second error is generated between the predicted value of the local point cloud and the true value of the local point cloud. The parameters in the 3D reconstruction network are updated based on the first error and the second error of each frame of sample image.
[0012] This application also provides an endoscope navigation device, the device comprising: A 3D model acquisition module is used to acquire a target 3D model of a target object; the target 3D model is obtained based on target images from at least two orientations; the target images are obtained by scanning the target area of the target object; An endoscopic image acquisition module is used to acquire endoscopic images of a target object; the endoscopic images are acquired in real time by a visual camera on the endoscope when the endoscope is inserted into the target area; The image detection module is used to perform image detection on the endoscopic image and obtain the target tissue image in the endoscopic image; The navigation module is used to compare the target tissue image with the target 3D model to determine the target position of the endoscope, so as to navigate based on the target position.
[0013] In another aspect, this application also provides an electronic device, including: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the above-described endoscopic navigation method.
[0014] This application also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the above-described endoscopic navigation method.
[0015] Compared with the prior art, the technical solution provided in this application has the following advantages: This application first obtains a 3D model of the target patient undergoing surgery, obtained through scanning from at least two perspectives, which fully reflects the physiological characteristics of the target area. Then, during surgery, after real-time endoscopic images of the target patient are acquired via a visual camera, image detection is performed on the endoscopic images to obtain the target tissue image. The computer then compares the target tissue image with the 3D model to intuitively determine the specific target location of the endoscope, allowing for subsequent navigation operations based on this location. This scheme pre-generates a 3D model of the target patient using pre-operative scan images, and then compares the real-time acquired endoscopic images with the 3D model to directly determine the specific location of the endoscope. Endoscopic navigation can then be achieved using a visual camera, ensuring accuracy during endoscopic navigation. Attached Figure Description
[0016] Figure 1 A flowchart of an endoscopic navigation method according to an embodiment of the present invention is shown.
[0017] Figure 2 A flowchart illustrating an endoscopic navigation method according to an embodiment of the present invention is shown.
[0018] Figure 3 The diagram illustrates the training logic of a three-dimensional reconstruction network according to an embodiment of this application.
[0019] Figure 4 This is a schematic diagram of the structure of an endoscope navigation device provided in an embodiment of this application.
[0020] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an optional embodiment of the present invention. Detailed Implementation
[0021] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without being bound by the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0022] Flexible endoscopes (such as uroscopes, bronchoscopes, and cystoscopes) are widely used in minimally invasive surgery. However, due to their narrow field of view, two-dimensional imaging, and lack of depth information, surgeons are prone to spatial disorientation and distance perception errors during operation. Real-time acquisition of the three-dimensional structure of the surgical area can provide surgeons with valuable spatial positioning references. For example, real-time 3D reconstruction can be used to accurately measure the distance between two points within the body, assisting in determining the relative positions of surgical instruments and tissues.
[0023] However, achieving real-time 3D reconstruction and navigation in flexible endoscopic surgery is no easy task. The surgical environment under endoscopy is complex and variable: tissues may deform, and both the endoscope and the tissue can move simultaneously, failing to meet the rigid "static world" assumptions commonly used in traditional computer vision. Furthermore, endoscopic images often contain texture-sparse tissue surfaces, complex lighting and shadows, and interference from fluids such as blood, all of which pose challenges to feature-matching-based 3D reconstruction algorithms. Traditional multi-view geometry-based structured light, stereo vision, or SLAM (Simultaneous Localization and Mapping) methods are prone to failure or accuracy degradation in these endoscopic scenarios, requiring extensive engineering optimization to guarantee robust real-time performance. Some existing endoscopic navigation solutions rely on additional hardware, such as electromagnetic locators, but these methods often increase the complexity and cost of navigation.
[0024] In order to achieve endoscope navigation based on a visual camera, this application provides an endoscope navigation method. Figure 1 A flowchart of an endoscopic navigation method according to an embodiment of the present invention is shown, as follows: Figure 1 As shown, the method flow includes: Step 101: Obtain the target 3D model of the target object.
[0025] The target's three-dimensional model is obtained from target images from at least two orientations; the target images are obtained by scanning the target area of the target object.
[0026] In one possible implementation of this application, a computed tomography (CT) device can be used to scan the target area of the target object before surgery to obtain at least two original tomographic image sequences with different projection directions.
[0027] In one alternative implementation, at least two orientations, including but not limited to coronal, sagittal, transverse, or helical scan data from the device at different angles, are used to enhance the structural integrity of the model.
[0028] After obtaining the original tomographic image sequence, it can be preprocessed to improve its image quality. Then, based on the preprocessed tomographic image sequence, a 3D model of the target object can be generated using a 3D reconstruction algorithm.
[0029] In another possible implementation of this application, the target image can be an ultrasound scan image of a target area of the target object. For example, the target object can undergo an ultrasound scan before surgery, and the operator can manually and in real time change the position and angle of the probe to obtain different sections of the ultrasound image, thereby obtaining ultrasound images from different orientations.
[0030] Furthermore, in this embodiment of the application, in order to reconstruct a three-dimensional model from an ultrasound image sequence, it is necessary to obtain the position and orientation information of each frame of the image in three-dimensional space. This embodiment of the application can use visual algorithms (such as feature-point-based image registration and visual SLAM technology) or neural networks to estimate the motion trajectory of the ultrasound probe during the scanning process, thereby recovering the corresponding three-dimensional spatial coordinates of each frame of the two-dimensional ultrasound image, so that a target three-dimensional model of the target object can be generated subsequently through a three-dimensional reconstruction algorithm.
[0031] Optionally, the aforementioned 3D reconstruction algorithms can be Marching Cubes, voxel-based surface extraction algorithms, or implicit surface reconstruction methods based on deep learning. The target 3D model includes the spatial geometry, topological relationships, and surface texture information of the target region.
[0032] Step 102: Obtain endoscopic images of the target object.
[0033] The endoscopic images are acquired in real time by a visual camera on the endoscope when the endoscope is inserted into the target area. Specifically, during the surgery, after the flexible endoscope enters the surgical area, the visual camera acquires endoscopic images in real time, which include a continuous sequence of two-dimensional color image frames.
[0034] Optionally, the visual camera in this application can be any one of a monocular camera, a binocular camera, or a multi-view camera. When the visual camera is a monocular camera, its hardware integration difficulty is minimal, and its impact on the flexibility and size of the endoscope is minimal. However, when the size requirements of the endoscope are relatively relaxed, a binocular camera or a multi-view camera can be used, thereby enabling the acquired endoscopic images to have more realistic depth information, or even directly form a dense three-dimensional point cloud, thereby improving the accuracy of subsequent image processing.
[0035] Step 103: Perform image detection on the endoscopic image to obtain the target tissue image in the endoscopic image.
[0036] In this embodiment of the application, after acquiring the endoscopic image, the real-time acquired endoscopic image can be input into a trained tissue recognition model, such as YOLO, Mask R-CNN, Swin Transformer, etc., to automatically detect the target tissue region.
[0037] Step 104: Compare the target tissue image with the target 3D model to determine the target location of the endoscope, so as to navigate based on the target location.
[0038] In this embodiment of the application, the target tissue image obtained in step 103 can be matched with the surface features of the target three-dimensional model generated in step 101.
[0039] Optionally, in this embodiment, the target tissue image can be pixel-matched with the target 3D model. For example, a computer device can acquire images of the target 3D model from various directions (e.g., sample them) to obtain candidate images corresponding to different positions and directions. Then, the candidate images are matched with the target tissue image for similarity. The position and orientation of the endoscope are determined based on the position and orientation of the candidate image with the highest similarity.
[0040] In another alternative implementation, after acquiring the target tissue image in the endoscopic image, the feature point cloud in the endoscopic image can also be acquired. Then, the feature point cloud is registered with the feature point cloud in the target 3D model, and the pose of the endoscopic camera coordinate system relative to the 3D model coordinate system is solved, thereby determining the position and orientation of the endoscope at this time.
[0041] Once the target location of the endoscope is determined, since the target area's three-dimensional model has already been constructed using preoperative CT images, the operator of the endoscope can autonomously select the subsequent operation method based on the real-time acquired position of the endoscope (i.e., the target location) and the various tissue structures marked on the target three-dimensional model, thereby achieving endoscope navigation.
[0042] In summary, this application first obtains a three-dimensional model of the target object to be operated on, which is obtained through scanning from at least two perspectives, fully reflecting the physiological characteristics of the target area. Then, during the operation, after acquiring endoscopic images of the target object in real time via a visual camera, image detection is performed on the endoscopic images to obtain the target tissue image. The computer then compares the target tissue image with the target three-dimensional model, allowing for a direct determination of the endoscope's specific target location, enabling subsequent navigation operations based on this location. This scheme pre-generates a target three-dimensional model from pre-operative scan images of the target object, and then compares the real-time acquired endoscopic images with the target three-dimensional model to directly determine the endoscope's specific location. Endoscopic navigation can then be achieved using a visual camera, ensuring accuracy during endoscopic navigation.
[0043] Figure 2This illustration shows a flowchart of an endoscopic navigation method according to an embodiment of the present invention. In this embodiment, the target region is the kidney region of the human body, and the tissue to be detected within the target region is the renal calyx. Figure 2 As shown, the method includes: Step 201: Obtain the target 3D model of the target object.
[0044] In this embodiment, a target three-dimensional model of the renal pelvis and calyces collecting system can be reconstructed based on the preoperative renal CT image data of the target object. Then, medical image segmentation and three-dimensional reconstruction software (such as 3D Slicer) is used to extract the anatomical structure of the renal collecting system to obtain the spatial morphology of each level of renal calyces. The method for generating this target three-dimensional model can be found in step 101, and will not be repeated here.
[0045] Step 202: Obtain endoscopic images of the target object.
[0046] After completing the preoperative preparations as described in step 201, continuous endoscopic images can be acquired in real time via the visual camera in the endoscope during the operation or examination of the target subject.
[0047] In this embodiment of the application, the endoscope is equipped with a visual camera. When the endoscope intervenes in the target area of the target object, the visual camera can collect images inside the target area, such as continuously collecting images at a certain frequency, that is, collecting real-time endoscopic images of the target area.
[0048] Step 203: Perform image detection on the endoscopic image to obtain the target tissue image in the endoscopic image.
[0049] In this embodiment, real-time endoscopic images acquired during flexible endoscopic surgery can be input into an image recognition module. A deep learning-based target detection model (such as the YOLO series algorithms) combined with traditional computer vision algorithms is used to detect and identify the visible renal calyx openings within the renal pelvis. Specifically, a YOLO model is pre-annotated using a large number of flexible endoscopic surgical video frames to train the "renal calyx opening" target, enabling it to detect the location of the renal calyx inlet appearing in the field of view in real time within complex endoscopic images.
[0050] Step 204: Based on the target 3D model, obtain a simulation view library; the simulation view library includes simulation images of each tissue to be detected from multiple perspectives.
[0051] In this embodiment, multiple virtual camera positions and poses can be preset within the spatial range of the target 3D model (e.g., distributed in a regular grid or generated based on the normal direction of the model surface). Then, the 3D model is visualized and rendered using a computer graphics rendering engine (such as OpenGL, Unity, Unreal Engine, or volume rendering algorithm) to generate simulated images from each virtual perspective. When generating each simulated image, the 3D coordinate position and perspective pose parameters (rotation matrix, translation vector, etc.) of the target tissue in the simulated image are recorded for subsequent pose calculation.
[0052] The simulated view library formed by the above scheme contains simulated images from multiple perspectives, the corresponding three-dimensional positions of tissues, and virtual camera pose information.
[0053] Taking the kidney as an example, in this embodiment, the anatomical number and attribute information of each renal calyx are identified in the target 3D model, including the number of renal calyxes, their location distribution, and their grouping (e.g., upper calyx group, middle calyx group, lower calyx group). The relative positional relationship of each renal calyx opening within the renal pelvis can be intuitively obtained through the target 3D model. For example, one renal calyx opening is located slightly above the inner wall of the renal pelvis, while another is located posteroinferiorly. Then, by setting up a virtual camera to render each renal calyx of the target 3D model, the multi-view simulation view obtained can characterize the image features that should be exhibited when observing different renal calyxes from various angles.
[0054] Step 205: Match the target tissue image with the simulated images in the simulated view library to determine the target simulated image.
[0055] In this embodiment of the application, when the target tissue image is detected, that is, the renal calyx is detected, its visual feature parameters are further extracted through image processing, such as the diameter or area of the renal calyx (reflecting the size of the renal calyx), the relative position of the renal calyx in the image (the degree to which the center point is biased towards the top, bottom, left, or right side of the image), and the angle or orientation of the renal calyx relative to the endoscopic visual axis.
[0056] The visual feature parameters obtained from the intraoperative images are matched with the multi-view simulation view generated from the preoperative target 3D model to identify the renal calyx currently being seen and observe its specific orientation in real time, thereby determining the current position of the flexible endoscope in the kidney.
[0057] Step 206: Determine the target location of the endoscope based on the position of the tissue to be detected in the target 3D model of the target simulation image and the viewpoint corresponding to the target simulation image.
[0058] In this embodiment, a simulated image library stores simulated images from multiple viewpoints, the 3D position of each simulated image, and the pose information of the virtual camera. Simply put, since the simulated images in the simulated image library are obtained by rendering with a virtual camera set up in the target 3D model, the 3D position and pose information of the virtual camera can be recorded simultaneously when a simulated image is acquired.
[0059] If the target tissue image is matched with the simulated view library to obtain the target simulated image corresponding to the target tissue image, the three-dimensional position and virtual camera pose information corresponding to the target simulated image can be directly used as the position and pose information of the endoscope.
[0060] Specifically, during the procedure, the semantic features (size, position, angle, etc.) of the renal calyx openings detected by YOLO in the endoscopic image can be matched with the simulation view library to find the most matching renal calyx model identifier, thereby knowing which renal calyx the flexible endoscope is currently pointing at or entering, and thus determining the target position corresponding to the current endoscope.
[0061] Furthermore, when matching the target tissue image with the simulated view library, the resulting simulated target image may not be completely identical to the target tissue image; it may only be the simulated image with the highest similarity in the simulated view library. In other words, the position and pose of the endoscope when acquiring the target tissue image may not be completely consistent with a pre-set virtual camera. In this case, to improve the positioning accuracy of the endoscope, it is necessary to adaptively adjust the three-dimensional position corresponding to the target simulated image and the pose of the virtual camera.
[0062] In one optional implementation, when the present application embodiment obtains a target virtual image that matches the target tissue image, it is first necessary to extract feature points from the target tissue image to obtain first feature point information; simultaneously, feature points are extracted from the target virtual image to obtain second feature point information; then, the first feature point information and the second feature point information are registered to obtain mapping parameters between the first feature point information and the second feature point information (mapping parameters include rotation parameters, translation parameters, and scaling parameters); then, the three-dimensional position of the target simulated image and the pose of the virtual camera are corrected according to the mapping parameters to obtain the position information and pose information of the endoscope.
[0063] Furthermore, since some tissue images may be highly similar but belong to different locations (for example, the entrances of different renal calyces may be very similar, and the accuracy of matching by images or feature points alone may not be high enough, and there is a risk of positioning errors), in other words, in certain specific scenarios, computer equipment may have difficulty distinguishing the accurate location of the endoscope based on a single frame image.
[0064] Therefore, in this embodiment of the application, the position information determined by the endoscope each time can be recorded in real time. Whenever the endoscope needs to locate itself based on the currently acquired endoscopic image, the target tissue image can be matched with the simulated view library to obtain several candidate images with the highest similarity. Then, based on the position indicated by the candidate image and the position determined by the endoscope last time, a virtual image is selected from several candidate images.
[0065] In practice, a state machine module (such as a Markov chain) can be added to the visual matching process to assist in decision-making. At the start of the surgery, the endoscope is inserted into the target area of the target object, and the state machine module initializes the state machine based on the entry position.
[0066] Then, after the endoscope acquires a tissue image for the first time, it will match it with the simulated view library. At this time, the computer device will select the candidate position that is connected to the previous position and has a reasonable motion trajectory from the matched candidate images (for example, based on the anatomical structure of the target area to be acquired and the motion estimation of the endoscope) as the selected virtual image. Then, the position corresponding to the virtual image is taken as the current position of the endoscope and is recorded in the state machine module.
[0067] In the subsequent process, it is essentially an iterative process described above. That is, whenever an tissue image is acquired and matched with the simulated view library to obtain candidate images with similarity that meet the threshold, the previous position recorded in the state machine is used as a reference to select the candidate images that are connected to it and have a reasonable motion trajectory as the final selected virtual image. The location is then determined based on the position corresponding to the virtual image.
[0068] Step 207: Perform 3D reconstruction based on the endoscopic image to obtain a 3D reconstruction model of the target area; the 3D reconstruction model is used to indicate path information in the target area.
[0069] In this embodiment of the application, when an endoscopic image is acquired, it can not only be used to determine the specific location of the endoscope, but also to perform three-dimensional reconstruction based on the endoscopic image to obtain a specific map of the target area, so as to guide the operator of the endoscope on how to operate the endoscope.
[0070] When performing 3D reconstruction, the acquired endoscopic images first need to be preprocessed. Specifically, the raw video frames acquired by the endoscope first enter the system's frame buffer module. Since the endoscope is a flexible endoscope, considering the potential interference from lens distortion, background noise, and other factors in the flexible endoscope image, the system performs preprocessing on each frame, including lens distortion correction, white balance and brightness equalization, and noise reduction, to improve the robustness of subsequent algorithms. The preprocessed image frames are then entered into a waiting queue in chronological order.
[0071] Optionally, in this embodiment of the application, the i-th window view is obtained; the i-th window view includes several frames of continuous endoscopic images within the i-th time window; where i is any positive integer from 1 to N, and N≥1; Input the i-th window view into the 3D reconstruction network to obtain the global point cloud corresponding to the i-th window view; The global point cloud of each window view is transformed to a specified coordinate system and fused to obtain a 3D reconstruction model of the target area.
[0072] Specifically, to fully utilize the parallel multi-view capability of the 3D reconstruction network, the system employs a sliding window mechanism to organize input frames. For example, the time window size can be set to a number of the most recent frames (generally 5-10 frames, selected based on hardware performance and scene dynamics). Whenever a new frame arrives, it is grouped with the preceding frames to form a window view and input into the 3D reconstruction network.
[0073] Inputting a window view into a 3D reconstruction network allows the network to process a small segment of video within the most recent time period each time, thereby utilizing parallax information between multiple viewpoints to improve reconstruction quality and robustness. At the same time, the window moves forward continuously, enabling the 3D reconstruction network to continuously monitor the latest changes in the scene.
[0074] Once each window view group is ready, it is fed into the 3D reconstruction network for forward inference. Due to the use of the Transformer to fuse multi-view features, the 3D reconstruction network comprehensively considers the information from these frames at once, outputting the local 3D point cloud corresponding to that window. Specifically, the decoder of the 3D reconstruction network outputs a global point cloud in a set reference coordinate system, such as the coordinate system of the first frame in the window (optionally, the decoder also predicts a local point cloud for each input image in the coordinate system of that view camera; this local point cloud is not used for subsequent image processing but is only used to update the network parameters in the 3D reconstruction network, so that the 3D reconstruction network can focus on both global features and local details in each frame).
[0075] In addition, the 3D reconstruction network can also output corresponding camera pose estimation and other information. Since the images in the window view in this embodiment are consecutive frames arranged in time, the first frame can be approximated as the reference coordinate origin of the window. Therefore, the final global point cloud can actually be regarded as the point cloud obtained by aligning all video frames in the time window.
[0076] After obtaining the global point cloud for each window, the system transforms it to the surgical global coordinate system for fusion. For example, for the first time window, the global point cloud is directly used as the initial global map (that is, the specified coordinate system mentioned above is the coordinate system of the global map for the first time window). For subsequent windows, the pose changes relative to the previous window need to be applied to the point cloud before merging it with the accumulated map. Furthermore, the global point cloud output by the aforementioned 3D reconstruction network has already solved the multi-view alignment problem to a certain extent. Therefore, each batch of new point clouds only needs minor pose adjustments based on the endoscopic motion increment to match the existing map. The scene fusion module performs filtering on the accumulated point cloud to remove outliers and can further eliminate minor errors using the ICP (Iterative Closest Point) fine registration algorithm to ensure map consistency and accuracy. The fused global point cloud is stored in the scene cache and supports fast querying and updating.
[0077] The latest point cloud map in the scene cache is then fed into the rendering module to generate an intuitive 3D display. Where hardware allows, the rendering module can convert the point cloud into a mesh surface in real time on the GPU: for example, using incremental Poisson reconstruction or Marching Cubes algorithms to generate a smooth triangular mesh from the accumulated point cloud for a more continuous and realistic surface visualization. If computational resources are limited, it can also be rendered directly as a point cloud (through volume rendering or point-by-point rendering). During rendering, camera intrinsics can be used to color the point cloud, giving each point a color texture from the original image, ensuring the 3D reconstructed model matches the actual endoscopic field of view. The rendering view is synchronized with the current endoscopic view by default, but the operator can also rotate and zoom the 3D view in the GUI to view different angles. The entire rendering process is updated in real time with each new frame, ensuring that the 3D model seen by the operator changes almost synchronously with the endoscopic operation.
[0078] Existing 3D reconstruction methods typically have very high computational requirements. During surgery or examination, the ability to quickly and in real-time complete 3D reconstruction places extremely high demands on the computing power of the equipment. However, by pre-training a 3D reconstruction network using the above method, and directly processing images in continuous time windows through the 3D reconstruction network, the amount of registration computation between multiple viewpoints can be greatly simplified. While ensuring the 3D reconstruction effect, the amount of 3D reconstruction computation is reduced and the generation speed of the 3D reconstruction model is improved.
[0079] Furthermore, Figure 3 A training logic diagram of a three-dimensional reconstruction network according to an embodiment of this application is shown. Figure 3 As shown, the 3D reconstruction network in this embodiment can be trained in the following way: Step 301: Obtain the training sample set.
[0080] The training sample set includes several training samples; each training sample includes several consecutive sample images, global point cloud ground truth, and local point cloud ground truth corresponding to each sample image.
[0081] Optionally, the training sample set can be data collected from intraoperative endoscopic videos of multiple different patients or in a simulated surgical environment, forming several training samples consisting of several consecutive frames of images; each sample contains temporally adjacent consecutive frames to characterize the visual changes of the endoscope moving continuously in actual operation.
[0082] For each training sample, its global 3D ground truth point cloud can be generated based on a 3D model generated by preoperative CT / MRI, through registration and calibration transformation to generate the corresponding global point cloud, or it can be obtained based on a high-quality SLAM or structured light scanning system. The local point cloud ground truth of each frame image can be obtained by calibrating the camera parameters to convert the global point cloud into the local view point cloud of the frame image for each sample image; or it can be obtained directly using external depth measurement equipment (such as ToF probe, structured light depth camera) to obtain the local depth information of the corresponding frame and then convert it into point cloud form.
[0083] Step 302: For each training sample, input the continuous sample image into the 3D reconstruction network to obtain the global point cloud prediction value and the local point cloud prediction value corresponding to each frame of the sample image.
[0084] In this embodiment, the 3D reconstruction network has a shared image encoder (ViT-L / CroCo / DINOv2 / DINOv3) that extracts visual features for each frame. Features from all frames are input together into a fusion Transformer for cross-frame self-attention interaction, allowing features from each frame to interact with information from other frames. The 3D reconstruction network contains two independent but structurally identical decoders (DPT decoders). One decoder outputs the local point cloud (and confidence score); the other decoder outputs the global point map (and confidence score).
[0085] Step 303: Generate a first error based on the global point cloud prediction value and the global point cloud ground truth value, and generate a second error between the local point cloud prediction value and the local point cloud ground truth value for each frame of sample image.
[0086] After obtaining the global point cloud prediction value and the local point cloud prediction value, the corresponding loss function can be input according to the two and their corresponding ground truth values to calculate their respective errors (i.e., the first error and the second error), so that the parameters in the 3D reconstruction network can be updated in the future through backpropagation and other methods.
[0087] Step 304: Update the parameters in the 3D reconstruction network based on the first error and the second error of each frame of sample image.
[0088] In this embodiment, the first error and the second error can be combined proportionally to generate the total loss value for updating the 3D reconstruction network. Optionally, after obtaining the first error and the second error corresponding to each frame of sample image, they can be weighted and summed to obtain the total error value. The weights of the first error and the second error can be preset to determine whether the 3D reconstruction network is more inclined towards the global accuracy of 3D reconstruction or the accuracy of local details in 3D reconstruction.
[0089] After obtaining the total error value, the parameters in the 3D reconstruction network can be updated by gradient backpropagation using optimization methods such as stochastic gradient descent.
[0090] Repeat steps 302-304 above until training converges or the preset training termination condition is met (e.g., the number of training iterations reaches a threshold), so that the 3D reconstruction network can stably output high-precision global point clouds.
[0091] In general, before the system is put into use in surgery, the training and verification of the three-dimensional reconstruction network need to be completed first. The training data can come from two sources: (1) video sequences of flexible endoscopy examinations and corresponding patient image data (such as CT / MRI) obtained before clinical surgery. The real three-dimensional reconstruction model is obtained through preoperative image reconstruction, and the camera pose of the video frame in the target three-dimensional model is obtained by optical tracking or manual labeling; (2) flexible endoscopy videos of ex vivo organs or high-precision simulation models taken under laboratory conditions, as well as high-precision three-dimensional ground truth values obtained by methods such as structured light scanning. Based on the above data, a training sample set containing a large number of multi-angle and multi-scene data can be established. Each sample consists of a set of flexible endoscopy images and their corresponding three-dimensional point cloud ground truth values.
[0092] During training, feature representations are first extracted for each training image. This invention employs an image encoder (e.g., a ViT-based CNN-Transformer hybrid coding network) to encode consecutive sample images in blocks, obtaining multi-scale features for each sample image.
[0093] Then, the multi-scale features are fed into the Transformer module, where image index location encoding is added to enable the 3D reconstruction network to identify the viewpoint from which different features originate. Simply put, consecutive images come from different times and different camera orientations. If the 3D reconstruction network doesn't know "which frame and viewpoint this feature comes from," it cannot perform cross-frame 3D inference. Therefore, image index location encoding is needed, for example, by processing the multi-scale features with encoding vectors set according to time and camera orientation, so that the 3D reconstruction network can retain the time, location, and pose information when extracting features from the image.
[0094] It should be noted that this embodiment employs a three-plane feature encoding strategy to represent three-dimensional spatial information: the cavity space is divided into three mutually perpendicular orthogonal planes, and a multi-channel feature map is learned on each plane. This allows the 3D reconstruction network to implicitly represent the position encoding and geometric texture information of any point in space by querying the feature values of these three planes. This significantly reduces storage and computational requirements while maintaining expressive power, thereby accelerating the training and inference speed of the 3D reconstruction network.
[0095] The 3D reconstruction network is then trained end-to-end, outputting a local point cloud for each input viewpoint and a global point cloud in the first viewpoint coordinate system (i.e., a 3D coordinate set fused from multiple views). The loss function employs a point-by-point regression error and confidence-weighted mechanism. Specifically, for each pixel in a training sample, the 3D reconstruction network predicts its 3D coordinates and confidence in the camera coordinate system, with the known ground truth coordinates as given. First, the normalized point distance error is calculated. The error of all pixels is accumulated using metrics such as Chamfer distance to obtain the overall point cloud loss, and the error is weighted by confidence to reduce gradient interference caused by a few points with large errors. Simultaneously, to ensure the density and accurate coverage of the output point cloud, the loss function also includes constraints on the prediction confidence, encouraging the 3D reconstruction network to output low confidence for background regions without ground truth supervision and high confidence for foreground cavity regions. Through iterative training with the Adam optimizer, the 3D reconstruction network gradually learns to directly infer stable 3D structural representations from multi-view images.
[0096] After training is completed, the 3D reconstruction network is tested on a validation dataset to confirm that the reconstruction accuracy meets the requirements (e.g., the average distance error between the reconstructed point cloud and the CT model is in the sub-millimeter range). Then, the model parameters can be fixed and deployed to the inference device for later use.
[0097] Step 208: Based on the target location, register the 3D reconstruction model with the target 3D model to indicate the path to each tissue to be detected.
[0098] In this embodiment of the application, with the help of a real-time updated 3D reconstruction model, the computer device can synchronously calculate a series of navigation assistance information.
[0099] First, the endoscopic camera pose: Through model prediction and point cloud registration, the computer device can estimate the current position and orientation of the endoscopic lens in the 3D reconstruction model (6-DOF pose) so as to mark the "current position of the camera" in the 3D reconstruction model, usually represented by a small camera icon or a view frustum.
[0100] Secondly, there is coverage area analysis: the computer equipment spatially rasterizes the accumulated point cloud, divides the inner surface of the cavity into several small areas, marks which areas are covered by the point cloud and which are still blank, and thus generates untraversed area prompts, such as highlighting the uncovered areas with different colors.
[0101] Furthermore, the computer equipment combines known anatomical structure models (i.e., the target 3D model) with preoperative planning data to match the reconstructed map, thereby indicating to the surgeon the current position of the endoscope relative to anatomical orientation (such as the upper and lower calyx of the kidney, the order of the bronchial branches, etc.).
[0102] In an optional implementation, the embodiments of this application may also display the exploration status of each tissue to be detected in the image display interface; the exploration status is whether the tissue to be detected has been explored or not.
[0103] The status of each organization to be tested can be updated in the following ways: Based on the endoscopic images of the target object, determine whether the endoscope has entered the target tissue in each tissue to be detected; if it has entered the target tissue, update the exploration status of the target tissue to "explored".
[0104] Specifically, all the above information is presented to the surgeon through a GUI. For example, in this embodiment, the computer device's GUI includes two main windows: the left side displays a real-time video feed of a traditional endoscope, allowing the surgeon to directly observe tissue details; the right side displays a 3D reconstructed map view, where the surgeon can see the overall structural model of the cavity and the location markers of the endoscope. When the surgeon moves the endoscope, the camera icon in the right-hand 3D view moves synchronously, serving as virtual navigation. The GUI also overlays the endoscope's trajectory—usually represented by a polyline indicating the path the endoscope lens has traveled—to help review and locate previously visited locations. For untraversed areas, the system displays them in semi-transparent red on the 3D map, indicating the need for further examination. The surgeon can use the interactive controls provided by the GUI to manipulate the 3D view (rotate, zoom, pan) to better observe a particular structure, and can also turn certain auxiliary layers on / off (such as hiding the trajectory lines or indicating uncovered areas). Upon detection of a target lesion (via an integrated AI module), the computer device will simultaneously provide alerts in two views: the target will be highlighted with a border or pseudo-color in the video feed, and a marker (such as a small flag) will be placed at the corresponding spatial location on the 3D map to indicate its spatial orientation, facilitating the surgeon's planning of the approach path. If connected to a preoperative navigation system, the GUI can also display the relative direction of the predetermined target (such as a stone or tumor), guiding the endoscope to gradually approach the target area, thus achieving a navigation loop.
[0105] In summary, this application first obtains a three-dimensional model of the target object to be operated on, which is obtained through scanning from at least two perspectives, fully reflecting the physiological characteristics of the target area. Then, during the operation, after acquiring endoscopic images of the target object in real time via a visual camera, image detection is performed on the endoscopic images to obtain the target tissue image. The computer then compares the target tissue image with the target three-dimensional model, allowing for a direct determination of the endoscope's specific target location, enabling subsequent navigation operations based on this location. This scheme pre-generates a target three-dimensional model from pre-operative scan images of the target object, and then compares the real-time acquired endoscopic images with the target three-dimensional model to directly determine the endoscope's specific location. Endoscopic navigation can then be achieved using a visual camera, ensuring accuracy during endoscopic navigation.
[0106] Furthermore, in practical use, this system will provide surgeons with a significant and intuitive navigation enhancement effect. Surgeons can not only view the traditional endoscopic field of view on the monitor but also simultaneously refer to a 3D map to understand the overall anatomical structure. Especially in cavities with complex anatomical structures (such as the renal collecting system and small bronchial branches), the system's path backtracking and unobserved area prompts help surgeons systematically examine every corner without overlooking anything. When encountering a suspected lesion, surgeons can quickly confirm its spatial location on the 3D map, greatly improving positioning efficiency by considering its depth and direction relative to the entry point. If multiple entries and exits along the same path are required (e.g., repeated nephrolithotomy in the renal calyx), 3D navigation can guide surgeons to the target along the optimal path, reducing unnecessary repeated exploration time and improving the predictability and safety of the operation.
[0107] This application also provides an endoscope navigation device for implementing the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0108] This application provides an endoscope navigation device. Figure 4 This is a schematic diagram of the structure of an endoscope navigation device provided in an embodiment of this application. The device includes: The 3D model acquisition module 401 is used to acquire a target 3D model of the target object; the target 3D model is obtained based on target images from at least two orientations; the target images are obtained by scanning the target area of the target object; Endoscopic image acquisition module 402 is used to acquire endoscopic images of a target object; the endoscopic images are acquired in real time by a visual camera on the endoscope when the endoscope is inserted into the target area; Image detection module 403 is used to perform image detection on the endoscopic image and obtain the target tissue image in the endoscopic image; The navigation module 404 is used to compare the target tissue image with the target three-dimensional model to determine the target position of the endoscope, so as to navigate based on the target position.
[0109] In summary, this application first obtains a three-dimensional model of the target object to be operated on, which is obtained through scanning from at least two perspectives, fully reflecting the physiological characteristics of the target area. Then, during the operation, after acquiring endoscopic images of the target object in real time via a visual camera, image detection is performed on the endoscopic images to obtain the target tissue image. The computer then compares the target tissue image with the target three-dimensional model, allowing for a direct determination of the endoscope's specific target location, enabling subsequent navigation operations based on this location. This scheme pre-generates a target three-dimensional model from pre-operative scan images of the target object, and then compares the real-time acquired endoscopic images with the target three-dimensional model to directly determine the endoscope's specific location. Endoscopic navigation can then be achieved using a visual camera, ensuring accuracy during endoscopic navigation.
[0110] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an optional embodiment of the present invention. This electronic device can be a computer device used to execute the above-described method. Figure 5 As shown, the electronic device includes one or more processors 10, a memory 20, and interfaces for connecting the various components, including high-speed interfaces and low-speed interfaces. The various components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processor can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces).
[0111] The processor 10 may further include a hardware chip. This hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0112] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.
[0113] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the use of the electronic device based on the display of a mini-program landing page. Furthermore, the memory 20 may include high-speed random access memory (RAM), and may also include non-transient memory, such as at least one disk storage device, flash memory device, or other non-transient solid-state storage device. The memory 20 may include volatile memory, such as RAM; the memory may also include non-volatile memory, such as flash memory, hard disk, or solid-state drive; the memory 20 may also include combinations of the above types of memory.
[0114] The electronic device also includes a communication interface 30 for communicating with other devices or communication networks.
[0115] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0116] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0117] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. An endoscope navigation method characterized by, The method includes: Obtain a target 3D model of the target object; the target 3D model is obtained based on target images from at least two orientations; the target images are obtained by scanning the target area of the target object; Acquire endoscopic images of the target object; the endoscopic images are acquired in real time by a visual camera on the endoscope when the endoscope is inserted into the target area; Image detection is performed on the endoscopic image to obtain the target tissue image in the endoscopic image; The target tissue image is compared with the target 3D model to determine the target location of the endoscope, so as to navigate based on the target location.
2. The method of claim 1, wherein, There are multiple tissues to be detected in the target area; The step of comparing the target tissue image with the target 3D model to determine the target location of the endoscope includes: Based on the target 3D model, a simulated view library is obtained; the simulated view library includes simulated images of each tissue to be detected from multiple perspectives. The target tissue image is matched with the simulated images in the simulated view library to determine the target simulated image; The target location of the endoscope is determined based on the position of the tissue to be detected in the target 3D model of the target simulation image and the viewpoint corresponding to the target simulation image.
3. The method of claim 2, wherein, The navigation based on the target location includes: The image display interface shows the exploration status of each tissue to be detected; the exploration status indicates whether the tissue to be detected has been explored or not. The method further includes: Based on the endoscopic images of the target object, it is determined whether the endoscope has entered the target tissue in each tissue to be detected; If the target organization is entered, the exploration status of the target organization will be updated to "explored".
4. The method according to claim 2 or 3, characterized in that, The navigation based on the target location includes: Three-dimensional reconstruction is performed based on the endoscopic images to obtain a three-dimensional reconstruction model of the target region; the three-dimensional reconstruction model is used to indicate path information in the target region; Based on the target location, the 3D reconstruction model is registered with the target 3D model to indicate the path to each of the tissues to be detected.
5. The method according to claim 4, characterized in that, The step of performing three-dimensional reconstruction based on the endoscopic image to obtain a three-dimensional reconstruction model of the target region includes: Obtain the view of the i-th window; the i-th window view includes several consecutive frames of endoscopic images within the i-th time window; where i is any positive integer from 1 to N, and N≥1; Input the i-th window view into the 3D reconstruction network to obtain the global point cloud corresponding to the i-th window view; The global point cloud of each window view is transformed to a specified coordinate system and fused to obtain a three-dimensional reconstruction model of the target area.
6. The method according to claim 5, characterized in that, The specified coordinate system is the coordinate system of the global map in the first time window.
7. The method according to claim 5, characterized in that, Before inputting the i-th window view into the 3D reconstruction network, the method further includes: Obtain a training sample set; the training sample set includes several training samples; each training sample includes several consecutive sample images, global point cloud ground truth, and local point cloud ground truth corresponding to each sample image; For each training sample, the continuous sample image is input into the 3D reconstruction network to obtain the global point cloud prediction value and the local point cloud prediction value corresponding to each frame of sample image; A first error is generated based on the predicted global point cloud value and the true global point cloud value. For each frame of sample image, a second error is generated between the predicted value of the local point cloud and the true value of the local point cloud. The parameters in the 3D reconstruction network are updated based on the first error and the second error of each frame of sample image.
8. An endoscope navigation device, characterized in that, The device includes: A 3D model acquisition module is used to acquire a target 3D model of a target object; the target 3D model is obtained based on target images from at least two orientations; the target images are obtained by scanning the target area of the target object; An endoscopic image acquisition module is used to acquire endoscopic images of a target object; the endoscopic images are acquired in real time by a visual camera on the endoscope when the endoscope is inserted into the target area; The image detection module is used to perform image detection on the endoscopic image and obtain the target tissue image in the endoscopic image; The navigation module is used to compare the target tissue image with the target 3D model to determine the target position of the endoscope, so as to navigate based on the target position.
9. An electronic device, characterized in that, include: The device includes a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the endoscopic navigation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the endoscopic navigation method according to any one of claims 1 to 7.