Pose tracking method and system
Patent Information
- Application Number
- CN202610817820.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]本发明提供一种位姿追踪方法和系统,用以解决现有技术难以在远距离手术场景下实现稳定、高鲁棒性的位姿追踪的缺陷,实现远距离下导航棒六自由度位姿的稳定追踪
[0015]本发明通过采用多个相互不共面外表面布设可见标记的导航棒,搭配第一相机与作为定位基准的第二相机组成的双相机架构,先通过第一相机在预设远距离成像范围内采集包含导航棒的图像,识别可见标记并获取其二维位置坐标,结合第一相机成像参数解算可见标记在第一相机坐标系下的第一位姿信息,再通过预先标定的双相机空间变换关系,将第一位姿信息转换至第二相机基准坐标系下得到第二位姿信息,最终基于可见标记相对导航棒的预设相对位姿信息,结合第二位姿信息解算并确定导航棒的最终位姿;解决了现有单相机视觉导航在远距离手术场景下小尺寸标记成像不足、单平面标记视角受限易失效的技术问题,实现了远距离下导航棒六自由度位姿的稳定追踪,规避了单标记遮挡导致的追踪中断风险,提升了手术导航定位的鲁棒性,同时无需依赖昂贵的红外定位设备,降低了系统成本与部署难度。
Smart Images

Figure CN122827784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pose tracking technology, and more particularly to a pose tracking method and system. Background Technology
[0002] In high-precision surgical scenarios such as neurosurgery, surgical navigation systems are core tools for achieving precise correspondence between preoperative medical images and intraoperative procedures. To avoid interfering with surgical operations, navigation sensors typically need to be deployed at a distance of 1.5 to 2 meters from the surgical area. Currently, existing visual navigation solutions usually employ single-camera systems with a standard field of view, estimating pose by identifying single-planar visual markers on surgical instruments such as navigation rods and combining this with camera imaging parameters or depth data.
[0003] However, due to the limitations of conventional lens imaging resolution and the observation angle of single-plane markers, existing visual navigation solutions have extremely low effective pixel ratios at distances of 1.5 to 2 meters for small-sized markers. Furthermore, when surgical instruments such as navigation rods undergo significant posture changes or partial occlusion, tracking features are easily lost, making it difficult for existing technologies to achieve stable and robust pose tracking in long-distance surgical scenarios. Summary of the Invention
[0004] This invention provides a pose tracking method and system to address the shortcomings of existing technologies in achieving stable and robust pose tracking in remote surgical scenarios, and to achieve stable tracking of the six-degree-of-freedom pose of a navigation rod at long distances.
[0005] This invention provides a pose tracking method, comprising the following steps: Acquire an image captured by a first camera within a preset imaging distance range. The image contains a navigation stick located within the target scene. Visible marks are distributed on multiple non-coplanar outer surfaces of the navigation stick. Identify at least one visible marker in the image and obtain the two-dimensional position coordinates of each visible marker in the image; Based on the two-dimensional position coordinates of each visible marker and the imaging parameters of the first camera, determine the first pose information of each visible marker in the coordinate system corresponding to the first camera; Based on the spatial transformation relationship between the coordinate system corresponding to the first camera and the coordinate system corresponding to the second camera, each first pose information is converted into second pose information in the coordinate system corresponding to the second camera, wherein the second camera is used to establish a reference coordinate system for pose tracking. Based on the preset relative pose information of each visible marker relative to the navigation stick, and combined with the second pose information corresponding to each visible marker, at least one candidate pose information of the navigation stick is determined; The pose information of the navigation stick is determined based on at least one candidate pose information of the navigation stick.
[0006] According to the pose tracking method provided by the present invention, each candidate pose information includes: position information and attitude information; When there are multiple candidate pose information, determining the pose information of the navigation stick based on at least one candidate pose information includes: Based on the position information and the attitude information, calculate the translation consistency error and rotation consistency error between each candidate pose information; Based on the translation consistency error and the rotation consistency error, target candidate pose information that meets the preset consistency condition is determined from the plurality of candidate pose information; The target candidate pose information is fused to obtain the pose information of the navigation stick.
[0007] According to a pose tracking method provided by the present invention, the target candidate poses are fused to obtain the pose information of the navigation stick, including: Robust statistical fusion is performed on the position information of the target candidate pose information to obtain fused position information; Quaternion pose fusion is performed on the pose information of the target candidate pose information to obtain fused pose information; Based on the fused location information and the fused attitude information, the target pose information of the navigation stick is determined.
[0008] According to a pose tracking method provided by the present invention, determining the pose information of the navigation stick based on at least one candidate pose information of the navigation stick includes: Based on the at least one candidate pose information, determine the initial pose information of the navigation stick; Obtain the pose information of the navigation stick from the previous moment as reference pose information; Based on the reference pose information, the position information in the initial pose information is subjected to temporal domain filtering, and the rotation information in the initial pose information is subjected to temporal domain interpolation to obtain smoothed position information and rotation information. The pose information of the navigation stick is obtained based on the smoothed position and rotation information.
[0009] According to a pose tracking method provided by the present invention, the method further includes: Based on the pose information of the navigation stick and the preset relative pose information of each visible mark relative to the navigation stick, calculate the three-dimensional coordinates of each visible mark in the coordinate system corresponding to the first camera; Based on the imaging parameters of the first camera, the three-dimensional coordinates of each visible mark are projected onto the image to obtain the projection position coordinates of each visible mark; Calculate the reprojection error of each visible mark based on the projected position coordinates and the two-dimensional position coordinates of each visible mark; The pose verification result is determined based on the reprojection error of each visible marker.
[0010] According to a pose tracking method provided by the present invention, the pose verification result includes: verification failure; The method further includes: If the pose verification result is a failure, the image captured by the first camera within the preset imaging distance range is reacquired to redetermine the pose information of the navigation stick.
[0011] The present invention also provides a pose tracking system, comprising: The navigation stick has a head comprising multiple non-coplanar outer surfaces, each of which is provided with visible markings. The first camera is used to acquire images within a preset imaging distance range; The second camera is used to acquire depth information of the target scene, and the depth information is used to construct a reference coordinate system for pose tracking. A processing unit is used to implement the pose tracking method as described in any of the above.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the pose tracking method as described above.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the pose tracking method as described above.
[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the pose tracking method as described above.
[0015] This invention employs a navigation rod with visible markers on multiple non-coplanar outer surfaces, combined with a first camera and a second camera serving as a positioning reference, forming a dual-camera architecture. First, the first camera acquires images containing the navigation rod within a preset long-distance imaging range, identifies the visible markers, and obtains their two-dimensional position coordinates. The first camera's imaging parameters are then used to calculate the first pose information of the visible markers in the first camera's coordinate system. Next, through a pre-calibrated dual-camera spatial transformation relationship, the first pose information is transformed to the second camera's reference coordinate system to obtain the second pose information. Finally, based on the preset relative pose information of the visible markers relative to the navigation rod, combined with the second pose information, the final pose of the navigation rod is calculated and determined. This invention solves the technical problems of insufficient imaging of small-sized markers and the limited viewing angle of single-plane markers leading to easy failure in long-distance surgical scenarios in existing single-camera visual navigation systems. It achieves stable tracking of the six-degree-of-freedom pose of the navigation rod at long distances, avoids the risk of tracking interruption caused by single marker occlusion, improves the robustness of surgical navigation and positioning, and eliminates the need for expensive infrared positioning equipment, reducing system costs and deployment difficulty. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is one of the flowcharts of the pose tracking method provided by the present invention.
[0018] Figure 2 This is one of the example flowcharts of the pose tracking method provided by the present invention.
[0019] Figure 3 This is one of the structural schematic diagrams of the pose tracking system provided by the present invention.
[0020] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0022] In high-precision surgeries such as neurosurgery, to ensure the sterility of the surgical area and to avoid the navigation equipment interfering with the surgeon's operating space, navigation sensors typically need to be deployed at a distance of 1.5 to 2 meters from the surgical target. At this specific imaging distance, to avoid interfering with the operation, the visible markers on the navigation stick used for pose tracking are often designed to be small.
[0023] In current high-precision surgical scenarios such as neurosurgery, pose tracking is mainly achieved using the following two types of technical solutions: 1. Infrared optical positioning system: This system uses multiple infrared cameras to capture infrared reflective balls deployed on the navigation device and achieves pose tracking based on the principle of triangulation.
[0024] 2. Traditional visual navigation solution: A single-camera system with a standard field of view (such as a regular RGB or RGBD camera) is used to estimate pose by recognizing a single planar visual marker on the navigation device and combining the camera's intrinsic parameters or single-point depth data.
[0025] The above-mentioned technical solutions have the following technical problems in long-distance surgical navigation scenarios: 1. Lack of long-distance imaging accuracy: Due to the small size of the navigation markers, when a standard field-of-view camera images at a distance of 1.5 meters or more, the effective pixel ratio of the visible markers on the photosensitive element is extremely low. This leads to a significant decrease in the recognition stability and positioning accuracy of the visible markers, making it difficult to meet the accuracy requirements in surgical scenarios.
[0026] 2. Poor robustness and susceptibility to occlusion: Because traditional methods rely on a single planar marker, during surgical procedures, if the navigation rod undergoes a significant attitude deflection (resulting in an excessively large angle between the visible marker surface normal and the camera axis) or if the surgical area is partially occluded, the camera will be unable to acquire sufficient geometric constraint features. This will cause immediate interruption of tracking, affecting the continuity of the surgical procedure.
[0027] To address the aforementioned technical issues, this application employs a first camera to perform long-distance imaging of a navigation stick with multiple non-coplanar visible markers and identifies the two-dimensional position coordinates of the visible markers. After determining the first pose of each visible marker in the first camera coordinate system based on the imaging parameters of the first camera, the first pose is transformed to a coordinate system with the second camera as the reference using the spatial transformation relationship between the first and second cameras. At least one candidate pose is determined by combining the preset relative pose information of each visible marker with respect to the navigation stick. Finally, the pose information of the navigation stick is determined based on the candidate poses.
[0028] The following is combined with Figures 1 to 2 The pose tracking method of the present invention is described.
[0029] Figure 1This is one of the flowcharts illustrating the pose tracking method provided by the present invention, such as... Figure 1 As shown, the method includes the following: S101. Acquire an image captured by the first camera within a preset imaging distance range. The image contains a navigation stick located in the target scene. Visible marks are distributed on multiple non-coplanar outer surfaces of the navigation stick.
[0030] Specifically, the first camera refers to an imaging device used to acquire two-dimensional images containing visible markers; the preset imaging distance range refers to a pre-set object distance range that can avoid interference areas in the surgical area and allows the first camera to effectively image the target visible markers; the target scene refers to the physical environment that includes the surgical operation space for pose tracking; the navigation stick refers to a rigid instrument that serves as a carrier for pose tracking; the visible markers refer to planar markers with visually identifiable coded features used for pose tracking; and the non-coplanar outer surfaces refer to multiple rigid outer surfaces on the navigation stick whose normal directions are not parallel and do not exist in the same plane.
[0031] In this embodiment, the first camera continuously acquires images of the target scene at a preset acquisition frame rate within a pre-calibrated effective imaging distance range. The acquired images must completely cover the observable area of the navigation stick, and the images must include at least one outer surface of the navigation stick with visible markings, providing raw image data for subsequent mark recognition and pose tracking.
[0032] Figure 2 The following is an exemplary flowchart of the pose tracking method provided by the present invention, such as... Figure 2 As shown, in the application scenario of neurosurgical navigation, the first camera adopts a telephoto RGB (Red Green) sensor. The camera (using the red, green, and blue primary color light mode) has a 25mm lens focal length, a horizontal field of view of 12°, and a preset imaging distance range of 1.5 to 2 meters, corresponding to the safe operating distance between the navigation device and the patient's surgical area in the operating room. The target scene is the surgical operating area of the operating room. The navigation stick is a neurosurgical navigation probe with a head that is a regular square polyhedron or an irregular polyhedron structure. Its sides are non-coplanar outer surfaces, and the angle between the normals of adjacent outer surfaces is 90°. Each outer surface is affixed with a visible AprilTag code with a unique ID. Each visible tag has a physical size of 10mm × 10mm. The first camera is fixed to the head of the operating table by a rigid bracket, with a straight-line distance of 1.8 meters from the patient's surgical area, within the preset imaging distance range. It synchronously acquires color images at a fixed frame rate of 30fps, with a single frame resolution of 1920 × 1080. The acquired images contain two visible AprilTag tags on the head of the navigation stick, without motion blur or overexposure issues.
[0033] By limiting the preset imaging distance range of the first camera to match the distance operation requirements of the target scene, and by using the structure of visible marks arranged on the multiple non-coplanar outer surfaces of the navigation stick, the problem of limited viewing angle of a single planar mark in the prior art is solved. This provides the hardware foundation and raw data support for stable mark recognition at long distances, and avoids the interruption of pose tracking due to the object distance exceeding the effective imaging range or the limited viewing angle of a single visible mark.
[0034] S102. Identify at least one visible marker in the image and obtain the two-dimensional position coordinates of each visible marker in the image.
[0035] Specifically, identifying at least one visible marker in an image refers to the process of extracting the encoded features of the visible marker from the image and completing identity matching through a visual detection algorithm; two-dimensional position coordinates refer to the two-dimensional coordinate values of the feature points of the visible marker in the image pixel coordinate system.
[0036] In this embodiment, after preprocessing the acquired image, feature extraction and identification are performed on visible markers in the image to distinguish different visible markers. Simultaneously, key feature points of each identified visible marker are extracted, and the two-dimensional position coordinates of each feature point in the image pixel coordinate system are output. The method for identifying visible markers in the image can employ conventional coded visual marker recognition algorithms in the art; any algorithm capable of matching the identity of visible markers and extracting feature point coordinates is applicable, and this embodiment of the invention does not impose a unique limitation on this.
[0037] For example, such as Figure 2 As shown, following the aforementioned neurosurgical navigation scenario, after Gaussian denoising preprocessing is performed on the acquired color image, the corner points of the two AprilTag visible marks are extracted by identifying the visible marks in the image, and the two-dimensional position coordinates of the four corner points of each visible mark with sub-pixel accuracy in the image pixel coordinate system are obtained respectively.
[0038] By accurately identifying visible markers in the image and extracting their two-dimensional position coordinates, accurate input data is provided for subsequent pose tracking. This solves the problem of insufficient feature extraction accuracy for small markers at long distances in existing technologies, and ensures the reliability of the basic data for pose calculation.
[0039] S103. Based on the two-dimensional position coordinates of each visible marker and the imaging parameters of the first camera, determine the first pose information of each visible marker in the coordinate system corresponding to the first camera.
[0040] Specifically, the imaging parameters of the first camera refer to the pre-calibrated intrinsic parameter matrix used to describe the imaging geometry of the first camera; the coordinate system corresponding to the first camera refers to the three-dimensional camera coordinate system established with the optical center of the first camera as the origin; the first pose information refers to the six-degree-of-freedom pose data of the visible markers in the image relative to the first camera coordinate system, which includes three-dimensional position information and three-dimensional pose information.
[0041] Optionally, in this embodiment, the pre-calibrated intrinsic parameter matrix of the first camera is used as the imaging constraint. By combining the actual physical size of the visible marker with the extracted two-dimensional position coordinates, the first pose information corresponding to each visible marker is obtained by solving the six-degree-of-freedom pose of each visible marker in the first camera coordinate system.
[0042] For example, such as Figure 2 As shown, following the aforementioned scenario of neurosurgical navigation, firstly, 25 images of a 12×9 checkerboard calibration board at different angles are pre-acquired using a telephoto RGB camera. The grid size is 20mm×20mm, and the intrinsic parameter matrix of the telephoto RGB camera is obtained for calibration. Next, after determining the two-dimensional position coordinates of the four corner points of two visible markers, based on the physical size of each visible marker, the two-dimensional position coordinates of the four corner points of each visible marker, and the intrinsic parameter matrix of the telephoto RGB camera, the rotation matrix and translation vector of each visible marker in the coordinate system corresponding to the telephoto RGB camera are obtained using the PnP (Perspective-n-Point) algorithm. These are combined into a three-dimensional special Euclidean group matrix, which serves as the first pose information. The formula for calculating the first pose information is shown below: in, This represents the first pose information of the i-th visible marker. Let i be the rotation matrix of the i-th visible marker. Let be the translation vector of the i-th visible marker.
[0043] This embodiment calculates the pose of the visible marker in the first camera coordinate system based on the imaging parameters of the first camera and the two-dimensional position coordinates of the visible marker, providing accurate data support for subsequent pose transformation.
[0044] S104. Based on the spatial transformation relationship between the coordinate system corresponding to the first camera and the coordinate system corresponding to the second camera, convert each first pose information into second pose information in the coordinate system corresponding to the second camera, wherein the second camera is used to establish the reference coordinate system for pose tracking.
[0045] Specifically, the second camera refers to the depth imaging device used to establish a spatial positioning reference and acquire scene depth information; the spatial transformation relationship refers to the rigid body transformation matrix between the first camera coordinate system and the second camera coordinate system obtained by pre-calibration, which includes rotation matrix and translation vector, and belongs to the SE(3) special Euclidean group transformation; the reference coordinate system refers to the world coordinate system used as a unified reference standard in the spatial positioning process, that is, the coordinate system established by the scene depth information collected by the second camera; the second pose information refers to the six-degree-of-freedom pose data of the visible marker relative to the second camera coordinate system.
[0046] In this embodiment, the pre-calibrated spatial transformation relationship is used as the transformation constraint. The first pose information of each visible mark obtained in step S103 is multiplied by the rigid body transformation matrix. The reference coordinate system of the first pose information is transformed from the first camera coordinate system to the second camera coordinate system to obtain the second pose information of each visible mark in the reference coordinate system.
[0047] Optionally, the calibration process for the spatial transformation relationship between the coordinate system corresponding to the first camera and the coordinate system corresponding to the second camera adopts a dual-camera calibration process. After the imaging parameters of the first camera and the second camera are calibrated respectively, a planar checkerboard calibration board is first used to simultaneously acquire images of the calibration board under different poses of the two cameras, ensuring that the calibration board is simultaneously within the complete field of view of both cameras. Then, for each set of synchronized images, the corner coordinates of the calibration board are extracted, a one-to-one correspondence between the corner points under the perspectives of the two cameras is established, and the fundamental matrix and essential matrix are solved based on the corresponding corner points to decompose the rotation matrix and translation vector between the two cameras. By globally optimizing the rotation matrix and translation vector and minimizing the reprojection error, the extrinsic transformation matrix between the first camera and the second camera is obtained, which serves as the spatial transformation relationship between the coordinate system corresponding to the first camera and the coordinate system corresponding to the second camera.
[0048] For example, such as Figure 2 As shown, following the aforementioned scenario, the second camera is a structured light depth camera, whose coordinate system serves as the reference coordinate system for surgical navigation pose tracking. Extrinsic parameter calibration of the two cameras is pre-done through a stereo calibration process, obtaining the spatial transformation relationship from the telephoto RGB camera coordinate system to the depth camera coordinate system. The second pose information of the two visible markers in the depth camera coordinate system is calculated using the following formula: in, The second pose information for the i-th visible marker. This represents the spatial transformation relationship between the coordinate system corresponding to the first camera and the coordinate system corresponding to the second camera. This is the first pose information of the i-th visible marker.
[0049] This embodiment completes the unified transformation of pose from the first camera coordinate system to the reference coordinate system through pre-calibrated spatial transformation relationships, realizing the decoupling of marker tracking and spatial positioning reference. It solves the problem of accuracy and stability conflict caused by a single camera simultaneously undertaking marker tracking and navigation registration in the prior art, ensuring the consistency of pose data and reference coordinate system, and avoiding subsequent pose tracking reference confusion caused by inconsistency in coordinate systems.
[0050] S105. Based on the preset relative pose information of each visible marker relative to the navigation stick, and combined with the second pose information corresponding to each visible marker, determine at least one candidate pose information of the navigation stick.
[0051] Specifically, the preset relative pose information refers to the fixed rigid body transformation relationship between the coordinate system of each visible marker and the coordinate system of the navigation stick body, which is obtained through pre-calibration; the coordinate system of the navigation stick body refers to the rigid coordinate system established with the preset reference point of the navigation stick as the origin; and the candidate pose information refers to the six-degree-of-freedom pose data of the navigation stick body relative to the reference coordinate system, which is obtained by solving the pose data of a single visible marker.
[0052] In this embodiment, the preset relative pose information corresponding to each visible marker is used as a geometric constraint. Combined with the second pose information of the visible marker in the reference coordinate system, the six-degree-of-freedom pose of the navigation stick body in the reference coordinate system is calculated through the inverse operation of rigid body transformation. A set of candidate pose information for the navigation stick is calculated for each visible marker. The candidate pose information corresponding to each visible marker is determined according to the following formula: in, This represents the candidate pose information corresponding to the i-th visible marker. The second pose information for the i-th visible marker. This represents the preset relative pose information of the i-th visible marker relative to the navigation stick.
[0053] For example, such as Figure 2 As shown, following the aforementioned scenario, during the manufacturing and assembly stage of the navigation stick, a fixed rigid body transformation matrix of the coordinate system of each visible marker relative to the coordinate system of the navigation stick body is predetermined through CAD design and calibration processes. This matrix represents the preset relative pose information corresponding to each marker. For each visible marker, its corresponding second pose information is multiplied by the inverse matrix of the preset relative pose information of that marker. This multiplication yields the six-degree-of-freedom pose of the navigation stick body corresponding to that visible marker in the depth camera coordinate system, which is a set of candidate pose information. Two sets of candidate pose information for the navigation stick are obtained for each of the two visible markers.
[0054] Based on the pre-established visible markers and fixed geometric constraints of the navigation stick, the overall pose of the navigation stick can be obtained through any visible marker. This solves the problem that pose tracking cannot be achieved when a single planar marker is occluded or the viewpoint is deflected in the prior art, and ensures the continuity and robustness of the navigation stick pose tracking.
[0055] S106. Determine the pose information of the navigation stick based on at least one candidate pose information of the navigation stick.
[0056] Specifically, the pose information of the navigation stick refers to the final output of the six-degree-of-freedom pose data of the navigation stick body relative to the reference coordinate system, which serves as the final output result of surgical navigation.
[0057] For example, such as Figure 2 As shown, following the aforementioned scenario, after verifying the consistency of the two sets of candidate pose information obtained from the calculation, the final six-degree-of-freedom pose information of the navigation stick body in the depth camera coordinate system is obtained through weighted fusion processing, which serves as the output of the spatial positioning result for neurosurgical navigation.
[0058] This invention employs a navigation rod with multiple visible markers on non-coplanar outer surfaces, combined with a first camera and a second camera serving as a positioning reference, forming a dual-camera architecture. First, the first camera acquires images containing the navigation rod within a preset long-distance imaging range, identifies the visible markers, and obtains their two-dimensional position coordinates. The first camera's imaging parameters are then used to calculate the first pose information of the visible markers in the first camera's coordinate system. Next, through a pre-calibrated dual-camera spatial transformation relationship, the first pose information is transformed to the second camera's reference coordinate system to obtain the second pose information. Finally, based on the preset relative pose information of the visible markers relative to the navigation rod, and combined with the second pose information, the final pose of the navigation rod is calculated and determined. This invention solves the technical problems of insufficient imaging of small-sized markers and the limited viewing angle of single-plane markers leading to easy failure in long-distance surgical scenarios using existing single-camera visual navigation. It achieves stable calculation of the six-degree-of-freedom pose of the navigation rod at long distances, avoids the risk of tracking interruption caused by single marker occlusion, improves the robustness of surgical navigation and positioning, and eliminates the need for expensive infrared positioning equipment, reducing system cost and deployment difficulty.
[0059] Optionally, the methods for determining the pose information of the navigation stick provided in the embodiments of the present invention are mainly the following two, and the first method for determining the pose information of the navigation stick will be described below.
[0060] In some embodiments, each candidate pose information includes: position information and attitude information.
[0061] When there are multiple candidate pose information, S106, determine the pose information of the navigation stick based on at least one candidate pose information, including: S201. Based on the position information and attitude information, calculate the translation consistency error and rotation consistency error between each candidate pose information.
[0062] S202. Based on translational consistency error and rotational consistency error, determine the target candidate pose information that meets the preset consistency conditions from multiple candidate pose information.
[0063] S203. The target candidate pose information is fused to obtain the pose information of the navigation stick.
[0064] Specifically, translational consistency error refers to the spatial distance deviation between the positional information of two sets of candidate pose information; rotational consistency error refers to the angular deviation between the attitude information of two sets of candidate pose information; preset consistency conditions refer to the pre-set translational error threshold and rotational error threshold used to filter valid candidate pose information; target candidate pose information refers to valid candidate pose information that meets the preset consistency conditions; fusion processing refers to the process of statistically optimizing multiple sets of valid pose data to obtain a comprehensive pose result. The specific values of the preset consistency conditions can be flexibly set according to the actual application scenario, navigation accuracy requirements, and sensor characteristics; this embodiment of the invention does not impose specific limitations.
[0065] In this embodiment, after acquiring multiple sets of candidate pose information, the translational consistency error and rotational consistency error between any two sets of candidate pose information are first calculated. Based on preset translational and rotational error thresholds, target candidate pose information that meets the consistency conditions is selected from the multiple sets of candidate pose information. Then, all the selected target candidate pose information is fused to obtain the final navigation stick pose information. The calculation formulas for translational consistency error and rotational consistency error are shown below: in, The translation consistency error between the two sets of candidate pose information. It is the Euclidean distance function. The position information of candidate pose information a. The position information of candidate pose information b. The rotational consistency error between the two sets of candidate pose information. It is an inverse cosine function. Matrix trace operation function, Let a be the pose information of candidate pose information a. The pose information of candidate pose information b.
[0066] For example, in the aforementioned neurosurgical navigation scenario, for the two sets of candidate pose information obtained, the translational consistency error between the two sets of candidate pose information is calculated to be 0.12 mm, and the rotational consistency error is 0.08°. The preset consistency conditions are set as translational error thresholds less than or equal to 0.2 mm and rotational error thresholds less than or equal to 0.15°. Both sets of candidate pose information meet the preset consistency conditions and are determined as target candidate pose information. By performing equal weight fusion processing on the two sets of target candidate pose information, the final navigation rod pose information is obtained.
[0067] The embodiments of the present invention can effectively eliminate abnormal pose results caused by factors such as marker detection noise and local reflection by verifying and screening the consistency of multiple candidate poses. Furthermore, the accuracy and stability of pose tracking are improved through fusion processing, which solves the problems of easy interference and insufficient robustness of pose results that rely on single markers in the prior art.
[0068] In some embodiments, S203, the target candidate poses are fused to obtain the pose information of the navigation stick, including: S301. Robustly statistically fuse the position information of the target candidate pose information to obtain fused position information.
[0069] S302. Perform quaternion pose fusion on the pose information of the target candidate pose information to obtain fused pose information.
[0070] S303. Based on the fused position information and fused attitude information, determine the target pose information of the navigation stick.
[0071] Specifically, robust statistical fusion refers to a processing method that optimizes and fuses multiple sets of position data based on robust statistical strategies. The robust statistical strategy is preferably a statistically robust weighted average or median strategy to suppress the interference of outlier data on the fusion result. This embodiment of the invention does not impose specific limitations on this. Quaternion attitude fusion refers to a fusion processing method that converts attitude information into unit quaternion form and performs weighted averaging and normalization on multiple sets of attitude data in quaternion space.
[0072] In this embodiment, the position information of all the selected target candidate pose information is robustly statistically fused to obtain the fused position information; at the same time, the attitude information of all the target candidate pose information is converted into unit quaternion form, and quaternion attitude fusion and normalization processing is performed to obtain the fused attitude information; finally, the fused position information and fused attitude information are combined to obtain the target pose information of the navigation stick.
[0073] For example, following the aforementioned scenario, the three-dimensional position information of the two sets of target candidate pose information is robustly statistically fused using the median method to obtain the fused three-dimensional position coordinates; the rotation matrices of the two sets of target candidate pose information are converted into unit quaternions respectively, and then normalized after equal weighted averaging to obtain the fused unit quaternions, which are then converted into rotation matrices, i.e., the fused attitude information; the fused position information is combined with the attitude information to obtain the final target pose information of the navigation stick.
[0074] This embodiment uses robust statistical fusion to process position information, effectively suppressing outlier interference in position data. It uses quaternion space fusion to process pose information, avoiding singularity issues in the rotation matrix fusion process. This further improves the accuracy and numerical stability of pose fusion results and solves the problems of numerical anomalies and accuracy loss that easily occur in multi-set pose fusion in the prior art.
[0075] The second method for determining the pose information of the navigation stick is explained below.
[0076] In some embodiments, S106, determining the pose information of the navigation stick based on at least one candidate pose information of the navigation stick includes: S401. Determine the initial pose information of the navigation stick based on at least one candidate pose information.
[0077] S402. Obtain the pose information of the navigation stick at the previous moment as reference pose information.
[0078] S403. Based on the reference pose information, perform temporal filtering on the position information in the initial pose information and temporal interpolation on the rotation information in the initial pose information to obtain smoothed position and rotation information.
[0079] S404. Based on the smoothed position and rotation information, obtain the pose information of the navigation stick.
[0080] Specifically, initial pose information refers to the navigation stick pose data initially determined based on candidate pose information; reference pose information refers to the verified navigation stick pose data output at the previous moment; temporal filtering processing refers to the processing method of smoothing the position information based on the continuity of the time series; temporal interpolation processing refers to the processing method of smoothing the rotation information based on the continuity of the time series.
[0081] This embodiment determines the initial pose information of the navigation stick based on at least one candidate pose information, obtains the navigation stick pose information output at the previous time step as reference pose information, uses the reference pose information as a smoothing benchmark, performs temporal filtering on the position information in the initial pose information, and performs temporal interpolation on the rotation information in the initial pose information to obtain smoothed position information and rotation information, which are finally combined to obtain the final pose information of the navigation stick. The method for determining the initial pose information of the navigation stick has been described in detail in the preceding embodiments, specifically in steps S201-S203 and S301-S303, and will not be repeated here.
[0082] For example, following the aforementioned neurosurgical navigation scenario, the initial pose information of the navigation rod is obtained by fusing two sets of candidate pose information, and the navigation rod pose information output from the previous frame is used as the reference pose information. The three-dimensional position information in the initial pose information is subjected to temporal filtering using a first-order exponential weighted moving average algorithm, with the filtering coefficient set to 0.7. The rotation matrices in the initial pose information and the rotation matrices in the reference pose information are both converted into unit quaternions, and temporal interpolation is performed using a spherical linear interpolation algorithm, with the interpolation coefficient set to 0.3. Finally, the smoothed position information and rotation information are combined to obtain the final output navigation rod pose information for the current frame.
[0083] This embodiment smooths the initial pose information based on time continuity, effectively suppressing the instantaneous pose jitter caused by marker detection noise without increasing computational overhead. This ensures the continuity and smoothness of the navigation stick pose output, solving the problem of pose output jitter and inability to meet the requirements of real-time intraoperative navigation display in the prior art.
[0084] The following describes the method for verifying the pose information of the navigation stick and the anomaly handling mechanism provided by the present invention.
[0085] In some embodiments, it also includes: S501. Based on the pose information of the navigation stick and the preset relative pose information of each visible marker relative to the navigation stick, calculate the three-dimensional coordinates of each visible marker in the coordinate system corresponding to the first camera.
[0086] S502. Based on the imaging parameters of the first camera, project the three-dimensional coordinates of each visible mark onto the image to obtain the projection position coordinates of each visible mark.
[0087] S503. Calculate the reprojection error of each visible mark based on its projected position coordinates and two-dimensional position coordinates.
[0088] S504. Determine the pose verification result based on the reprojection error of each visible marker.
[0089] Specifically, reprojection error refers to the pixel deviation between the actual observed coordinates of the feature points of the visible markers in the image and the predicted coordinates obtained by projection based on pose information; pose verification result refers to the result obtained by judging the validity of the navigation stick pose information based on the reprojection error.
[0090] This embodiment calculates the three-dimensional coordinates of each visible marker in the first camera coordinate system based on the final determined pose information of the navigation stick and the preset relative pose information of each visible marker relative to the navigation stick. Then, based on the imaging parameters of the first camera, the three-dimensional coordinates are projected onto the image acquired by the first camera to obtain the projected position coordinates of each visible marker feature point. The deviation between the projected position coordinates of each feature point and the actual observed two-dimensional position coordinates is calculated to obtain the reprojection error of each visible marker. The pose information of the navigation stick is validated based on the reprojection error, and the pose validation result is output.
[0091] For example, following the aforementioned scenario, in determining the pose verification result, firstly, based on the final output navigation stick pose information and combined with the preset relative pose information of each visible marker, the three-dimensional coordinates of each marker in the telephoto RGB camera coordinate system are calculated; then, based on the intrinsic parameter matrix of the telephoto RGB camera, the three-dimensional coordinates of the four corner points of each visible marker are projected onto the acquired color image to obtain the projected position coordinates of each corner point of the visible marker; the Euclidean distance between the projected position coordinates of each corner point and the actual extracted two-dimensional position coordinates in step S102 is calculated, and the average value is taken to obtain the reprojection error of the marker as 0.18 pixels; with a preset reprojection error threshold of 0.5 pixels, the pose is determined to be valid based on this error value, and the pose verification result that has passed the verification is output.
[0092] This embodiment verifies the validity of the navigation stick pose information by calculating the reprojection error. It can accurately identify abnormal pose tracking situations and provides a quantitative verification basis for the reliability of pose results. It solves the problems of existing technologies that cannot judge the validity of pose tracking in real time and are prone to outputting incorrect positioning results.
[0093] In some embodiments, the pose verification result includes: verification failed.
[0094] The method also includes: S601, if the pose verification result is a failure, re-acquire the image captured by the first camera within a preset imaging distance range, so as to redetermine the pose information of the navigation stick.
[0095] Specifically, a verification failure refers to a verification result where the reprojection error exceeds a preset threshold, rendering the current pose tracking result invalid; reinitialization refers to re-executing the entire process of image acquisition, marker recognition, and pose calculation to restore the navigation stick pose tracking process.
[0096] In this embodiment, after receiving the pose verification result that failed, the output of the current invalid pose information is first terminated, the image acquisition process of the first camera is retried, a new image within the preset imaging distance range is acquired, and the tracking process of the navigation stick pose information is re-executed based on the new image to complete the re-initialization of pose tracking.
[0097] For example, following the aforementioned scenario, when the navigation stick deflects significantly and only a local area of a single marker is visible, the calculated reprojection error is 0.8 pixels, exceeding the preset threshold of 0.5 pixels, and the pose verification result is a failure. By terminating the output of the current invalid pose, the telephoto RGB camera is triggered to re-acquire a color image. Based on the newly acquired image, the marker recognition, pose calculation and verification process is re-executed until the navigation stick pose information that has passed verification is obtained, and stable tracking is restored.
[0098] This embodiment can quickly recover from the pose tracking failure state by re-initializing the process when the verification fails. This avoids the output of erroneous pose information and ensures the continuity and reliability of pose tracking during surgical navigation. It solves the problem in the prior art that pose tracking failure cannot be automatically recovered and is prone to navigation interruption.
[0099] The pose tracking system provided by the present invention is described below. The pose tracking system described below can be referred to in correspondence with the pose tracking method described above.
[0100] Figure 3 This is one of the structural schematic diagrams of the pose tracking system provided by the present invention. For example... Figure 3 As shown, the pose tracking system includes: The navigation stick 101 has a head that includes multiple non-coplanar outer surfaces, each of which is provided with visible markings.
[0101] The first camera 102 is used to acquire images within a preset imaging distance range.
[0102] The second camera 103 is used to acquire depth information of the target scene, and the depth information is used to construct a reference coordinate system for pose tracking.
[0103] The processing unit 104 is used to implement any of the above pose tracking methods.
[0104] Specifically, the pose tracking system refers to the integrated hardware and software system used to realize real-time six-degree-of-freedom pose tracking of the navigation stick 101; the processing unit 104 refers to the computing and processing module used to perform image acquisition control, marker recognition, pose calculation and data output.
[0105] In this embodiment, the head of the navigation stick 101 is provided with multiple non-coplanar outer surfaces, and visible marks are fixedly arranged on each outer surface; the first camera 102 and the second camera 103 are fixedly installed by a rigid structure. The first camera 102 is used to acquire images containing the navigation stick within a preset imaging distance range, and the second camera 103 is used to acquire depth information of the target scene and construct a reference coordinate system for pose tracking; the processing unit 104 is communicatively connected to the first camera 102 and the second camera 103 respectively, and is used to execute any of the above pose tracking methods.
[0106] For example, in a neurosurgical navigation application scenario, the pose tracking system includes a neurosurgical navigation probe, a telephoto RGB camera, a depth camera, and an embedded processing unit. The head of the navigation probe has a regular square polyhedron structure, with a unique AprilTag visible mark affixed to each of its four non-coplanar sides. The telephoto RGB camera and the depth camera are rigidly fixed by an aluminum alloy bracket and installed on the head end of the operating table in the operating room, at a distance of 1.8 meters from the patient's surgical area, within a preset imaging distance range of 1.5-2 meters. The telephoto RGB camera is used to acquire color images including the navigation probe, and the depth camera is used to acquire depth information of the surgical area, with its coordinate system serving as the reference coordinate system for navigation. The embedded processing unit uses an NVIDIA Jetson Xavier NX module, which is connected to the telephoto RGB camera and the depth camera via gigabit Ethernet communication, respectively, to execute the aforementioned pose tracking method and complete the real-time pose tracking and navigation output of the navigation probe.
[0107] In this embodiment, the pose tracking system achieves functional decoupling of long-distance marker imaging and navigation reference establishment through the coordinated setting of the first camera and the second camera. The tracking robustness is improved by the multi-faceted, multi-marker navigation rod structure. It does not rely on expensive infrared optical positioning equipment. The system has a simple structure, is easy to deploy, and has controllable cost, solving the problems of high cost, complex deployment, and poor long-distance tracking stability of existing surgical navigation systems.
[0108] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840. The processor 810, communications interface 820, and memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a pose tracking method, which includes a pose tracking method.
[0109] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the pose tracking method provided by the above methods, the method including: pose tracking method.
[0111] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the pose tracking method provided by the above methods, the method comprising: a pose tracking method.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pose tracking method, characterized in that, include: Acquire an image captured by a first camera within a preset imaging distance range. The image contains a navigation stick located within the target scene. Visible marks are distributed on multiple non-coplanar outer surfaces of the navigation stick. Identify at least one visible marker in the image and obtain the two-dimensional position coordinates of each visible marker in the image; Based on the two-dimensional position coordinates of each visible marker and the imaging parameters of the first camera, determine the first pose information of each visible marker in the coordinate system corresponding to the first camera; Based on the spatial transformation relationship between the coordinate system corresponding to the first camera and the coordinate system corresponding to the second camera, each first pose information is converted into second pose information in the coordinate system corresponding to the second camera, wherein the second camera is used to establish a reference coordinate system for pose tracking. Based on the preset relative pose information of each visible marker relative to the navigation stick, and combined with the second pose information corresponding to each visible marker, at least one candidate pose information of the navigation stick is determined; The pose information of the navigation stick is determined based on at least one candidate pose information of the navigation stick.
2. The method according to claim 1, characterized in that, Each candidate pose information includes: position information and attitude information; When there are multiple candidate pose information, determining the pose information of the navigation stick based on at least one candidate pose information includes: Based on the position information and the attitude information, calculate the translation consistency error and rotation consistency error between each candidate pose information; Based on the translation consistency error and the rotation consistency error, target candidate pose information that meets the preset consistency condition is determined from the plurality of candidate pose information; The target candidate pose information is fused to obtain the pose information of the navigation stick.
3. The method according to claim 2, characterized in that, The target candidate poses are fused to obtain the pose information of the navigation stick, including: Robust statistical fusion is performed on the position information of the target candidate pose information to obtain fused position information; Quaternion pose fusion is performed on the pose information of the target candidate pose information to obtain fused pose information; Based on the fused location information and the fused attitude information, the target pose information of the navigation stick is determined.
4. The method as described in claim 1, characterized in that, Determining the pose information of the navigation stick based on at least one candidate pose information of the navigation stick includes: Based on the at least one candidate pose information, determine the initial pose information of the navigation stick; Obtain the pose information of the navigation stick from the previous moment as reference pose information; Based on the reference pose information, the position information in the initial pose information is subjected to temporal domain filtering, and the rotation information in the initial pose information is subjected to temporal domain interpolation to obtain smoothed position information and rotation information. The pose information of the navigation stick is obtained based on the smoothed position and rotation information.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: Based on the pose information of the navigation stick and the preset relative pose information of each visible mark relative to the navigation stick, calculate the three-dimensional coordinates of each visible mark in the coordinate system corresponding to the first camera; Based on the imaging parameters of the first camera, the three-dimensional coordinates of each visible mark are projected onto the image to obtain the projection position coordinates of each visible mark; Calculate the reprojection error of each visible mark based on the projected position coordinates and the two-dimensional position coordinates of each visible mark; The pose verification result is determined based on the reprojection error of each visible marker.
6. The method according to claim 5, characterized in that, The pose verification result includes: verification failed; The method further includes: If the pose verification result is a failure, the image captured by the first camera within the preset imaging distance range is reacquired to redetermine the pose information of the navigation stick.
7. A pose tracking system, characterized by include: The navigation stick has a head comprising multiple non-coplanar outer surfaces, each of which is provided with visible markings. The first camera is used to acquire images within a preset imaging distance range; The second camera is used to acquire depth information of the target scene, and the depth information is used to construct a reference coordinate system for pose tracking; A processing unit is configured to implement the pose tracking method as described in any one of claims 1-6.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the pose tracking method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the pose tracking method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the pose tracking method as described in any one of claims 1 to 6.