Surgical instrument tracking method and system based on composite structure positioner

By employing a surgical instrument tracking method based on a composite structure locator and utilizing binocular vision and consistency verification technology, the problem of decreased accuracy in surgical instrument pose calculation under complex operating room environments was solved, achieving high-precision and high-reliability surgical instrument tracking.

CN121867948APending Publication Date: 2026-04-17SHANGHAI QUANSHI INTELLIGENT SENSE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI QUANSHI INTELLIGENT SENSE TECHNOLOGY CO LTD
Filing Date
2026-02-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing visual tracking solutions struggle to distinguish between real feature points and false observation points caused by metallic reflections or bloodstains in the highly complex environment of the operating room, leading to a decrease in the accuracy of surgical instrument pose calculation.

Method used

A surgical instrument tracking method based on a composite structure locator is adopted. Distortion correction and noise reduction are performed through a binocular vision acquisition device, the positioning coding structure is detected, sub-pixel feature point extraction and 3D reconstruction are performed by combining the circular feature structure, outliers are removed, and the pose reliability index is calculated by using consistency verification to ensure the reliability of the output pose data.

Benefits of technology

It significantly improves the tracking accuracy and reliability of surgical instruments in complex lighting environments, ensures high confidence of output data, avoids misleading pose calculations, and enhances the safety and real-time response capability of medical navigation systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121867948A_ABST
    Figure CN121867948A_ABST
Patent Text Reader

Abstract

The invention provides a surgical instrument tracking method and system based on a composite structure positioner, and the method comprises the steps: the composite structure positioner is disposed on a surgical instrument, and comprises a positioning coding structure and a circular ring feature structure, and the positioning coding structure and the circular ring feature structure have a predetermined relative geometrical relationship; the method comprises the following steps: acquiring a binocular image containing a composite structure positioner through a binocular vision acquisition device, and performing correction and denoising processing to generate a standardized binocular image; and coarse positioning is completed based on the positioning coding structure to obtain an initial pose estimated value and determine a prediction search area of the circular ring feature structure, observation three-dimensional point information of circular ring features is reconstructed by using left and right eye binocular images, the observation three-dimensional point information is compared with preset three-dimensional point information, abnormal points with large errors are eliminated, and a surgical instrument fine positioning pose result is calculated. And when the final pose credibility does not meet the condition, skipping the current frame and processing the next frame of image. Error points are eliminated through binocular three-dimensional point cloud comparison, and the surgical instrument tracking precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and medical device navigation technology, and in particular to a surgical instrument tracking method and system based on a composite structure locator. Background Technology

[0002] In the fields of computer-assisted surgery and medical robot navigation, vision-based surgical instrument tracking technology is widely used due to its non-contact nature. However, the extreme complexity of real-world surgical scenarios presents significant challenges to existing vision tracking solutions. First, the operating room environment is typically characterized by strong illumination from shadowless lamps, and surgical instruments, mostly made of metal, are highly prone to specular reflection. Furthermore, unavoidable blood splatter or tissue obstruction during surgery can lead to noise or missing local features in visual markers during imaging. Existing tracking algorithms often lack effective mechanisms for eliminating physical anomalies, making it difficult to distinguish between genuine feature points and false observations caused by reflections or contaminants, resulting in decreased pose calculation accuracy. Therefore, we propose a surgical instrument tracking method and system based on a composite structure locator.

[0003] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by providing a surgical instrument tracking method and system based on a composite structure locator, thereby solving the technical problems mentioned in the background section.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A surgical instrument tracking method based on a composite structure locator includes the following steps: S1. Synchronously acquire binocular images containing a composite structure locator using a binocular vision acquisition device, and perform distortion correction and noise reduction processing on the acquired binocular images to obtain standardized binocular images. S2. Detect the localization coding structure and determine the decoding region in the standardized binocular image. Calculate the initial pose estimate of the composite structure locator based on the localization coding corner information. Based on this, determine the prediction search region of the circular feature structure in the standardized binocular image. S3. Extract sub-pixel feature points from the circular ring feature structure within the prediction search area, and perform three-dimensional reconstruction based on the extracted sub-pixel feature points of the circular ring and the calibration parameters of the binocular camera to obtain the observation three-dimensional point information of the circular ring feature structure. S4. Compare the observed three-dimensional point information with the preset three-dimensional point information of the composite structure locator, remove abnormal points whose spatial position error exceeds the preset threshold, and use the effective three-dimensional point set after removal to calculate the fine positioning pose result. S5. Perform integrity monitoring processing on the fine positioning pose results. The integrity monitoring processing includes at least the following indicators for consistency verification: binocular reprojection error, circular feature geometric fitting residual, and calculate the pose credibility index based on the consistency verification results. S6. When the pose confidence index meets the preset confidence threshold condition, the precise positioning pose result is output as the final tracking pose of the surgical instrument; when the pose confidence index does not meet the preset confidence threshold condition, the current pose result is not output, the current frame is skipped, the next pair of binocular images is read, and steps S1 to S6 are re-executed.

[0006] S1 specifically involves: synchronously acquiring the left and right eye images, including the composite structure locator, through a binocular vision acquisition device, and adding the same timestamp to each pair of images to obtain the original binocular image pair; Based on preset camera intrinsic parameters, distortion correction and epipolar correction are performed on the original binocular image pairs to obtain corrected binocular image pairs, so as to ensure that the correspondence between the left and right eyes can be used for subsequent pose calculation. Denoising is performed on the calibrated binocular image pairs to obtain preprocessed binocular image pairs, thereby suppressing random noise from the image sensor; The preprocessed stereo image pairs are defined as normalized stereo images, and the normalized stereo images are used as input for the coarse localization step of the subsequent localization coding structure.

[0007] S2 specifically involves: detecting candidate regions for localization coding structures in standardized binocular images and determining the effective region range that needs to be decoded; Decode and validate the effective region, and extract the corner pixel coordinates of the localization coding structure in the image. Based on the corner pixel coordinate information and the known geometric parameters of the positioning coding structure, the initial pose estimate of the composite structure locator is calculated. By utilizing the initial pose estimate and the predetermined relative geometric relationship between the localization coding structure and the circular feature structure, the predicted search region of the circular feature structure is determined in the standardized binocular image, and the predicted search region is used as the input for the fine localization step of the circular feature structure.

[0008] S3 specifically involves: performing corner detection and filtering on the circular feature structure within the predicted search area, extracting checkerboard feature corners that conform to the geometric distribution of the circular ring, and performing sub-pixel refinement to obtain the sub-pixel circular feature corner point set in the left and right eye images; using the principle of binocular stereo vision triangulation, based on the sub-pixel circular feature corner point set in the left and right eye images and the binocular camera calibration parameters, reconstructing the observation three-dimensional point set of the circular feature structure in the current camera coordinate system.

[0009] S4 specifically involves: using the observed three-dimensional points calculated in step S3, performing three-dimensional circle fitting based on a preset radius, marking points with fitting errors greater than a preset threshold as outliers and removing them, retaining the remaining points as a valid three-dimensional point set, establishing a point-to-point registration model using the valid three-dimensional point set and the preset three-dimensional model points of the composite structure locator, directly solving the pose of the composite structure locator in the current camera coordinate system, and obtaining the fine positioning pose result.

[0010] S5 specifically involves: calculating a set of indicators required for consistency verification of the fine positioning pose results, the set of indicators including: binocular reprojection error based on the fine positioning pose results, and geometric fitting residual of the circular feature; comparing the set of indicators with preset consistency conditions, performing consistency verification, and obtaining the verification result of each type of indicator; generating a pose reliability index based on the verification results of each consistency verification to characterize the reliability of the fine positioning pose results.

[0011] S6 specifically determines whether the pose confidence index meets the preset confidence threshold condition; if it does, the fine positioning pose result is determined as the final tracking pose and output; if it does not meet the condition, the current frame data is determined to be unreliable, the current calculation result is discarded, the downgraded pose output is not executed, the system state is directly reset and the acquisition and processing process of the next frame image is triggered.

[0012] A surgical instrument tracking system based on a composite structure locator includes: The system includes a binocular vision acquisition device, a composite structure locator, an image preprocessing module, a coarse positioning module with a positioning coding structure, a fine positioning module with a circular ring feature structure, a 3D point cloud cleaning and pose calculation module, an integrity monitoring and consistency verification module, and a pose output control module.

[0013] The beneficial effects of this invention are as follows: This invention employs a "binocular 3D point cloud comparison and removal" strategy. Compared to traditional temporal filtering or multi-frame optimization, this method directly utilizes the geometric constraints of binocular vision to identify and remove 3D reconstruction defects caused by metallic reflections, bloodstains, or uneven local lighting within a single frame. This "point cloud cleaning" mechanism, based on physical spatial scale, effectively eliminates the impact of observational gross errors on pose calculation, significantly improving the tracking accuracy of surgical instruments in complex lighting environments.

[0014] This invention constructs a hierarchical solution architecture consisting of a localization coding structure and a circular feature structure. The coarse localization stage quickly locks the ROI, while the fine localization stage utilizes the high signal-to-noise ratio of the circular features for independent solution, avoiding the propagation of coarse localization errors and achieving sub-pixel-level high-precision tracking.

[0015] This invention establishes a rigorous integrity monitoring and frame loss re-acquisition mechanism. Unlike degraded prediction outputs that carry the risk of misleading results, this system immediately discards the current unreliable frame and quickly triggers the acquisition of the next frame when it detects reprojection errors or excessive fitting residuals. This mechanism ensures that every frame of pose data output to the surgical navigation system is high-confidence data that has undergone rigorous verification, greatly improving the safety and real-time response capability of the medical navigation system. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of a surgical instrument tracking method based on a composite structure locator according to the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1: As Figure 1 As shown, this embodiment provides a surgical instrument tracking method based on a composite structure locator, including the following steps: S1. Synchronously acquire binocular images containing a composite structure locator using a binocular vision acquisition device, and perform distortion correction and noise reduction processing on the acquired binocular images to obtain standardized binocular images. S2. Detect the localization coding structure and determine the decoding region in the standardized binocular image. Calculate the initial pose estimate of the composite structure locator based on the localization coding corner information. Based on this, determine the prediction search region of the circular feature structure in the standardized binocular image. S3. Extract sub-pixel feature points from the circular ring feature structure within the prediction search area, and perform three-dimensional reconstruction based on the extracted sub-pixel feature points of the circular ring and the calibration parameters of the binocular camera to obtain the observation three-dimensional point information of the circular ring feature structure. S4. Compare the observed three-dimensional point information with the preset three-dimensional point information of the composite structure locator, remove abnormal points whose spatial position error exceeds the preset threshold, and use the effective three-dimensional point set after removal to calculate the fine positioning pose result. S5. Perform integrity monitoring processing on the fine positioning pose results. The integrity monitoring processing includes at least the following indicators for consistency verification: binocular reprojection error, circular feature geometric fitting residual, and calculate the pose credibility index based on the consistency verification results. S6. When the pose confidence index meets the preset confidence threshold condition, the precise positioning pose result is output as the final tracking pose of the surgical instrument; when the pose confidence index does not meet the preset confidence threshold condition, the current pose result is not output, the current frame is skipped, the next pair of binocular images is read, and steps S1 to S6 are re-executed.

[0019] S1 specifically includes the following sub-steps: S110. Synchronous Acquisition: Synchronous acquisition of the scene containing the composite structure locator is performed using a binocular vision acquisition device, obtaining the original left-eye image and the original right-eye image, and recording timestamps for both. The left-eye timestamp is recorded as follows: The right eye timestamp is Calculate synchronization deviation Its definition is: ; in, The degree of inconsistency in acquisition time between the left and right eyes. Synchronization threshold with preset In comparison, when When the left and right original images are combined, it is determined that they constitute a valid original binocular image pair; when If either binocular image is missing, the original binocular image pair is discarded, and a discard flag is generated for statistical input in subsequent integrity monitoring processing. At the same time, synchronous acquisition is retried to obtain the next pair of original binocular images.

[0020] S120, Geometric Correction: Perform distortion correction and epipolar correction on the valid original binocular image pairs to obtain corrected binocular image pairs. Distortion correction is performed based on preset camera intrinsics, which include at least the intrinsic matrix and distortion parameters of the left and right cameras. These preset camera intrinsics are obtained from factory calibration or the calibration process after installation and stored in the system configuration. Distortion correction includes at least radial distortion correction, which corrects pixel coordinates as follows: , ; in, For normalized image plane coordinates, , To preset the distortion coefficient, The corrected normalized image plane coordinates, The radial distance from the pixel to the center of the image is given. Subsequently, epipolar correction is performed based on the binocular geometry to ensure that corresponding points of the same spatial point in the left and right eye images are aligned with the epipolar lines (i.e., located in the same row or approximately in the same row), thereby obtaining a corrected binocular image pair for subsequent pose calculation and binocular 3D reconstruction.

[0021] S130. Denoising Processing: Denoising is performed on the calibrated binocular image pair to obtain a preprocessed binocular image pair. Since the operating room environment may contain electromagnetic interference or low-light noise, a filtering algorithm is needed to suppress random noise generated by the image sensor to improve the stability of subsequent sub-pixel edge extraction. Gaussian filtering is preferred as the denoising method, and its kernel function is defined as: ; in The preset standard deviation is used to control the smoothing level. In actual processing, a 3×3 or 5×5 convolution kernel is used to perform convolution operations on the image. Alternatively, if salt-and-pepper noise exists in the image, median filtering can be used for processing. After denoising, the jaggedness of the image edges is reduced, and the signal-to-noise ratio in flat areas is improved. This represents the local coordinate distance within the convolution kernel.

[0022] S140. Standardized Output: The preprocessed stereo image pair is defined as a standardized stereo image, and image quality gating is performed before output to ensure the stable implementation of the subsequent coarse localization step of the localization coding structure. The standardized stereo image includes at least: a standardized left-eye image, a standardized right-eye image, corresponding timestamps, and identification information associated with preset camera intrinsic parameters. Image quality gating is achieved by calculating the sharpness index Q, which is defined as: ; in, Let p be the grayscale value. Let p be the gradient at pixel p. The total number of pixels in the image. Indicates amplitude. Compare Q with a preset sharpness threshold. Comparison: When When, the standardized binocular image is output as input for the subsequent coarse localization step of the localization coding structure; when When the current image is determined to be blurry (e.g., due to focus failure or motion blur), the image pair corresponding to the standardized stereo image is discarded and a quality insufficiency flag is generated. At the same time, the synchronization acquisition is retried to obtain a new valid original stereo image pair.

[0023] S2 specifically includes the following sub-steps: S210, Candidate Region Detection and Determination: Detection of the localization coding structure is performed in the standardized binocular image to determine the effective region range to be decoded. Specifically, adaptive thresholding is performed on the standardized binocular image to obtain a binary image; connected components or contours are extracted from the binary image; polygon approximation is performed on each contour to obtain a set of polygon vertices, and approximately quadrilateral regions are selected as candidate regions for localization coding.

[0024] The candidate regions for localization coding are further filtered based on preset geometric screening conditions, which include at least a candidate region area threshold, a candidate region aspect ratio threshold, and a convexity condition. The regions retained after screening are the set of candidate regions for localization coding. These regions are identified as regions of interest (ROIs) containing potential coding information and will proceed to the subsequent decoding process.

[0025] S220. Decoding and Corner Extraction: Decode and validate each candidate region in the set of candidate regions for localization coding, and extract the corner pixel coordinate information. Specifically, extract the coordinates of the four vertices of each candidate region and calculate the homography matrix. Correct the perspective transformation of the candidate region to a frontal view code image. Sample and binarize the code elements on the code image according to a preset grid to obtain a binary code element sequence.

[0026] Perform validity checks on the binary symbol sequence (such as Cyclic Redundancy Check (CRC) or Hamming distance check). If the check passes, the region is determined to be a valid localization coding region, and the set of corner pixel coordinates of the localization coding structure in the image coordinate system is extracted. (Usually the four outer corners of the positioning code or specific internal feature points), which are used as key observation inputs for subsequent initial pose calculations.

[0027] S230, Initial pose estimation: Based on the set of corner pixel coordinates output in step S220. Calculate the initial pose estimate of the composite structure locator. Specifically, predefine the set of three-dimensional coordinates of the corresponding corner points of the positioning encoding structure in the coordinate system of the composite structure locator. .

[0028] Based on the preset camera intrinsic parameter matrix K, the rotation matrix is ​​solved using the PnP (Perspective-n-Point) algorithm. With translation vector This minimizes the reprojection error. ; in, This is the projection function from 3D to the pixel plane. The distance is the Euclidean distance. The calculated distance will be... This is output as the initial pose estimate. This step also calculates the initial reprojection error. This is used to evaluate whether the coarse localization has converged. If the value is too large, a retest will be required.

[0029] S240. Predicting the search region: using the initial pose estimate. Furthermore, a pre-determined relative geometric relationship between the localization coding structure and the annular feature structure is established to determine the prediction search region of the annular feature structure in the standardized binocular image. Specifically, a set of three-dimensional sampling points of the annular feature structure is pre-defined in the coordinate system of the composite structure locator. (For example, the center of a ring and discrete points on the circumference).

[0030] It should be noted that the relative geometric relationship and the set of three-dimensional points can be guaranteed by high-precision CNC machining, or pre-calibrated by an offline optical calibration instrument and stored in the system configuration file.

[0031] Based on the initial pose estimate, the set of 3D sampling points is projected onto the image plane to obtain the set of projected points. : ; For the set of projection points Calculate the minimum / maximum pixel coordinates in the horizontal and vertical directions. Construct a bounding rectangle containing the annular feature. Then, expand the bounding rectangle outwards by a predefined search margin. (Unit: pixels), forming the final predicted search area: ; The predicted search region is used as input for the fine localization step (S3) of the circular feature structure. Through this step, the system can eliminate background interference and search for circular features only within a very small image range, thereby significantly improving the speed and reliability of subsequent sub-pixel extraction.

[0032] S3 specifically includes the following sub-steps: S310. Corner point detection and screening of checkerboard circular features: Within the prediction search area (ROI) determined in step S2, corner point feature extraction is performed on the left and right standardized images respectively.

[0033] X-shaped corner detection: Addressing the texture of a checkerboard pattern with alternating black and white squares, this method abandons contour extraction methods affected by jagged edges and employs corner detection algorithms (such as Harris corner detection or FAST) to search for feature points with drastic pixel grayscale changes within the ROI region. The focus is on extracting X-shaped corners (X-junctions) at the boundaries of the black and white squares to form an initial corner set.

[0034] Subpixel corner refinement: To overcome image resolution limitations, subpixel refinement (such as the cornerSubPix iterative algorithm) is performed on the extracted initial corner set. By calculating the gray-level gradient and vector product in the neighborhood of the corner, the corner coordinate accuracy is converged to the subpixel level (within 0.1 pixels) to meet the requirements of high-precision positioning.

[0035] Noise removal based on geometric distribution constraints: Cleaning sub-pixel corner point sets using geometric prior knowledge of checkerboard annulus.

[0036] Constructing a circle / ellipse fitting model: Building a geometric fitting function based on the extracted corner coordinates (e.g., ... Or the general equation of an ellipse).

[0037] Residual filtering: Calculate the geometric distance from each corner point to the fitted curve and set a distance threshold. Remove isolated noise points or background interference points that deviate too much from the fitted curve, and retain the effective checkerboard corner point set that conforms to the circular distribution pattern as input for subsequent 3D reconstruction.

[0038] S320. Stereo matching based on epipolar constraints and topological order: Establishing sub-pixel corner point sets for the left eye image. Sub-pixel corner point set of the right eye image The point-to-point correspondence between them.

[0039] Epipolar correction is achieved by ensuring that corresponding physical points in the left and right eye images are located on the same image line based on the epipolar correction completed in step S1.

[0040] Matching strategy: Polar line constraint: Traverse every point in the left eye corner point set Search within the right corner point set for points that satisfy the row alignment constraint. ( Candidate points (with a tolerance threshold, such as 1.0 pixels).

[0041] Topological order consistency constraint: Utilizing the geometric closure of the checkerboard annulus, calculate the polar angles of the left and right corner point sets relative to their geometric centers, and sort the corner points in angular order (e.g., counter-clockwise). Perform one-to-one matching based on the sorted indices, eliminating outliers that do not satisfy the order consistency requirement, and generating a high-confidence set of matching point pairs. .

[0042] S330. Triangulation and reconstruction of observational 3D point sets: Using the principle of binocular stereo vision, the matched 2D sub-pixel corner point pairs are converted into 3D spatial coordinates, completing the mapping from the image domain to the spatial domain.

[0043] Projection model: Based on preset binocular camera calibration parameters, including the left camera projection matrix. With right camera projection matrix : ; ; Least squares solution: For each pair of matching corner points, solve the above system of equations to construct a linear system of equations, and use SVD (Singular Value Decomposition) to solve for the three-dimensional coordinate vector. .

[0044] Output: Traverse all matching point pairs to reconstruct the 3D point set of the circular feature structure in the current camera coordinate system. This point set consists of high-precision 3D coordinates of the checkerboard corner points, serving as input data for subsequent step S4 for geometric cleaning and pose calculation.

[0045] S4 specifically includes the following sub-steps: S410. Constrained 3D circle fitting based on a preset radius: using the observation 3D point set reconstructed in step S3. Based on the known geometric parameters (preset physical radius) of the annular feature structure in the composite structure locator ), perform constrained three-dimensional circle fitting.

[0046] Fitting objective: To find an optimal torus model in three-dimensional space, which is determined by the coordinates of the center C and the normal vector of the plane containing the torus. Sure.

[0047] Algorithm selection: The least squares method with radius constraints or the RANSAC (random sample consensus) algorithm is preferred.

[0048] Specific process: Plane fitting: First, calculate the point set using Principal Component Analysis (PCA) or Singular Value Decomposition (SVD). The covariance matrix is ​​obtained by taking the eigenvector corresponding to the smallest eigenvalue as the plane normal vector. The initial estimate is that the centroid of the point set is taken as the center of the plane.

[0049] Center optimization: Project all observation points onto the fitting plane. Construct the objective function in a two-dimensional planar coordinate system: ; in Let be the projection point. Let be the two-dimensional coordinates of the center of the circle in the plane. The radius is a pre-defined known radius (serving as a fixed constraint and not involved in optimization). The center position of the circle in the plane is solved through nonlinear optimization.

[0050] Parameter recovery: Transform the coordinates of the center of the circle in the plane back to three-dimensional space to obtain the fitted center of the circle. and final normal vector .

[0051] S420. Outlier detection and removal based on fitting error: Using the optimal spatial circle model obtained in step S410, the original observation data is quality assessed and cleaned to remove outliers caused by metal reflection, bloodstain obscuration, or matching errors.

[0052] Error definition: Define each observation point Spatial location error (fit residual) This error consists of two parts: Perpendicular distance: The perpendicular distance from a point to the fitted plane ; Radial distance: The difference between the distance from the projection of a point onto the center of a circle on a plane and the preset radius. .

[0053] Overall error: (i.e., the shortest Euclidean distance from a point to a three-dimensional circular loop).

[0054] Elimination logic: Set a preset error threshold (For example, set it to 0.2mm~0.5mm, the specific value to be determined based on the accuracy specifications of the surgical navigation system).

[0055] If a certain point error If the point is determined to be an outlier, it is removed from the point set. like The point was determined to be a valid observation point and was retained.

[0056] Output: The remaining points after traversing all observation points constitute a valid 3D point set. .

[0057] S430. Establish a 3D-3D point pair registration model: using the cleaned effective 3D point set. With the preset 3D model point set of the composite structure locator Establish point set registration relationships.

[0058] Correspondence Determination: Due to the rotational symmetry of the ring structure, a one-to-one correspondence between observation points and model points needs to be determined. Using the "initial pose estimate" calculated in step S2 as prior information, the nearest neighbor search or projection matching method is used to determine the correspondence between each valid observation point. Associated with the nearest corresponding point on the preset 3D model (CAD model) .

[0059] Model Construction: Constructing the rigid body transformation equations: ; Where R is a 3×3 rotation matrix and t is a 3×1 translation vector. To observe noise. The goal is to find the optimal... Make the registration error function minimize.

[0060] S440. Direct calculation of precise positioning pose: Based on the established point-to-point registration model, the precise pose of the composite structure positioner in the current camera coordinate system is directly solved to obtain the precise positioning pose result.

[0061] Solution method: The Absolute Orientation Problem is solved using the Singular Value Decomposition (SVD) method.

[0062] Decentralization: Calculate the centroids of the effective observation set and the corresponding model point set separately. and And calculate the coordinates after removing the centroid: , .

[0063] Covariance matrix calculation: Calculate the cross covariance matrix .

[0064] SVD decomposition: Singular value decomposition of matrix H .

[0065] Solving by rotation and translation: Calculate the rotation matrix (like Then, modify the last column sign of V to ensure it is a rotation matrix.

[0066] Calculate the translation vector: .

[0067] Output result: The solution obtained This output serves as the precise positioning and pose result. The result is directly derived from a high signal-to-noise ratio effective 3D point set, without introducing the cumulative error of the S2 coarse positioning, and has passed the S420 geometric cleaning, exhibiting extremely high accuracy.

[0068] S5 specifically includes the following sub-steps: S510, Integrity monitoring index calculation: Based on the fine positioning pose results output in step S4. Calculate the set of key indicators used for integrity monitoring.

[0069] Binocular reprojection error Using the precise positioning and pose results, the preset 3D model point set of the composite structure locator (including coded corner points and circular feature points) is projected onto the left and right image planes of the current frame, respectively. All valid projection points are calculated. Compared with the pixel observation points actually extracted in steps S2 and S3 The Euclidean distance between them is taken as the root mean square (RMS) value as the binocular reprojection error: ; This metric reflects the self-consistency of the final solved pose in the 2D image domain.

[0070] Circular feature geometric fitting residual : Obtain the valid 3D point set generated in step S420 The average three-dimensional geometric fitting residual is calculated as follows: This is the average distance from all retained valid points to the optimal spatial circular model fitted in step S410. This metric directly reflects the degree of spatial geometric agreement between the observed three-dimensional structure and the theoretical physical model.

[0071] S520. Consistency Threshold Verification: The calculated index set is compared with the system's preset consistency threshold conditions to eliminate obviously erroneous calculations.

[0072] Set reprojection error threshold (e.g., 1.0 pixel) and the fitted residual threshold (e.g., 0.2 mm).

[0073] Perform the following logical judgment: judge Is it valid? judge Whether it is valid or not.

[0074] Only when both of the above conditions are met is the current solution result considered to be consistent in terms of geometric constraints; if either index exceeds the limit, the current frame solution is directly considered abnormal (which may be due to camera calibration parameter drift, incomplete image distortion correction, or incorrect feature matching).

[0075] S530, Pose Reliability Index Generation: Based on the consistency verification results, a quantified pose reliability index is calculated. This is used to characterize the reliability level of the current output pose.

[0076] A normalized weighted scoring mechanism is adopted: Each error index is normalized into a score (range of values). ): ; ; Calculate the overall credibility based on preset weights: ; in, And usually set (For example , This is because binocular reprojection error can better reflect the accuracy of 3D pose on 2D images globally.

[0077] S540. Monitoring Result Judgment: Based on the pose reliability index. Compared with the preset confidence threshold (e.g., 0.8) Generate the final decision state.

[0078] Trusted state: when If all the hard threshold conditions in step S520 are met, a "credible" judgment result is generated, allowing the pose to enter the output process of step S6.

[0079] Untrusted state: when If any hard threshold condition is not met, an "unreliable" result is generated. At the same time, the specific reason for the unreliability is recorded (such as "reprojection error exceeds the limit" or "fitting residual is too large"), which is used as a status code for system debugging.

[0080] S6 specifically includes the following sub-steps: S610, Reliable Output Judgment: Determine the pose reliability index generated in step S5. Does it meet the preset confidence threshold condition (i.e.) And all hard consistency checks passed.

[0081] When the conditions are met: the fine positioning pose result obtained in step S4 is used. The final tracking pose of the surgical instruments at the current moment is determined. The system sends this final tracking pose, the corresponding timestamp, and the confidence status to the surgical navigation host computer or robot control system via the communication interface, completing the effective output of the current frame.

[0082] S620, Abnormal Frame Drop Handling: When the pose reliability index If the preset trust threshold condition is not met, or if any hard consistency check fails, the system determines that the calculation result of the current frame is unreliable.

[0083] At this point, to ensure the safety of medical navigation, the system implements a strict frame dropping strategy: Output Denied: Directly intercepts the calculation results of the current frame and does not output any pose data (nor the maintained or predicted pose of the previous frame), preventing outdated or incorrect pose data from misleading surgical procedures.

[0084] Status flag: Sends an "invalid" or "lost" status flag to the external system, indicating that the location cannot be found at the moment and prompting the doctor or control algorithm to wait for the next frame of data.

[0085] S630, Fast Reset and Resampling Trigger: After performing frame loss processing, the system immediately terminates unnecessary subsequent logical operations of the current frame and resets the temporary state buffer inside the algorithm (such as clearing the current abnormal feature point set).

[0086] The acquisition command of the binocular vision acquisition device is triggered synchronously, forcing the system to skip the remaining waiting time of the current frame and immediately read the next pair of latest binocular image data. This mechanism ensures that the system can attempt to resume tracking as quickly as possible when encountering interference (such as brief occlusion or reflection), rather than wasting computational resources on erroneous data.

[0087] S640, Closed-loop continuous operation: Update the system running status and record the processing result of the current frame ("successful output" or "frame loss skipped") into the system log. Then, the system pointer automatically jumps back to step S1 (image acquisition and preprocessing), and uses step S630 to trigger the acquisition of a new frame of image data to start a new round of tracking calculation.

[0088] Through the logic of S610 to S640 above, the system forms a real-time and rigorous closed-loop control flow: only outputs poses with high confidence, and immediately discards and refreshes when an anomaly is encountered, thereby maximizing the real-time response capability of the system while ensuring accuracy.

[0089] Example 2: This example provides a surgical instrument tracking system based on a composite structure locator, including: A binocular vision acquisition device is used to synchronously acquire binocular images including a composite structure locator. This device typically consists of two rigorously calibrated industrial-grade image sensors, equipped with hardware synchronization triggering circuitry to ensure that the time synchronization of left and right eye image acquisition meets microsecond-level requirements. A composite structure locator, set on surgical instruments (such as surgical probes, bone drills, etc.), includes a positioning coding structure and a circular feature structure, with a predetermined relative geometric relationship between the two; Positioning coding structure: It is preferred to use visual markers with ID recognition capabilities such as ArUco code, AprilTag code or DM QR code to provide device identification information and rough initial pose; Circular feature structure: A high-contrast checkerboard circular pattern (such as a black and white checkerboard ring) surrounding the positioning coding structure, used to provide sub-pixel level high-precision corner features; The image preprocessing module is used to perform distortion correction and noise reduction on the binocular images acquired by the binocular vision acquisition device. Specifically, it includes loading preset camera intrinsic parameters to perform distortion correction and epipolar correction, and using algorithms such as Gaussian filtering to remove image sensor noise, thereby obtaining standardized, high-quality standardized images for the left and right eyes. The localization coding structure coarse localization module is used to detect localization coding structures and determine decoding regions in standardized binocular images. This module is responsible for: The localization coding structure is identified and decoded to extract corner information; the initial pose estimate of the composite structure locator is calculated based on the corner information; the circular feature is projected back onto the image plane using the initial pose estimate and the geometric parameters of the locator, thereby accurately delineating the prediction search area (ROI) of the circular feature structure in the standardized binocular image. The circular feature structure fine localization module is used to extract sub-pixel feature points and reconstruct 3D structures of the circular feature structure within the predicted search area. Specifically, this module performs the following: Corner detection and filtering: Perform corner detection (e.g., for X-shaped corners) within the ROI region, extract checkerboard feature corners that conform to the circular geometric distribution, and perform sub-pixel refinement to obtain the sub-pixel circular feature corner set in the left and right eye images; Stereo matching and reconstruction: The correspondence between the left and right eye corner points is established by using epipolar constraints and topological order consistency strategies. Based on the principle of binocular stereo vision triangulation and combined with the binocular camera calibration parameters, the matched corner points are transformed into observation 3D point information (observation 3D point set) in the current camera coordinate system.

[0090] The 3D point cloud cleaning and pose calculation module is used to clean the observation data at the physical geometry level and perform final pose calculation. This module specifically performs the following: Geometric cleaning: Perform constrained 3D circle / ellipse fitting (based on preset physical radius) using observed 3D point information, calculate the geometric residual from each observation point to the fitted model; remove outliers (such as reflective noise points) whose residuals exceed a preset threshold, and obtain a high-confidence effective 3D point set. Pose calculation: A 3D-3D point pair registration model is established using the effective 3D point set and the preset 3D model points of the composite structure locator. The pose of the composite structure locator in the current camera coordinate system is directly solved by algorithms such as SVD (Singular Value Decomposition), and the fine positioning pose result is output. The integrity monitoring and consistency verification module is used to perform integrity monitoring on the fine positioning pose results. This module calculates a set of indicators, including at least the binocular reprojection error and the geometric fitting residual of the circular feature, compares the indicator set with a preset consistency threshold, and calculates a quantified pose reliability index based on the verification results to characterize the reliability of the current solution result; The pose output control module is used to control the final data flow based on the credibility index.

[0091] When the pose confidence index meets the preset confidence threshold condition, the precise positioning pose result is output as the final tracking pose of the surgical instrument. If the conditions are not met, the current frame output is intercepted, a frame dropping operation is performed, and the binocular vision acquisition device is immediately triggered to read the next pair of binocular images and enter the re-acquisition process.

[0092] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0093] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions according to the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.

[0094] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0095] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0096] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0097] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0098] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0099] If a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0100] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0101] In conclusion, the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A surgical instrument tracking method based on a composite structure locator, characterized in that, Includes the following steps: S1. Synchronously acquire binocular images containing a composite structure locator using a binocular vision acquisition device, and perform distortion correction and noise reduction processing on the acquired binocular images to obtain standardized binocular images. S2. Detect the localization coding structure and determine the decoding region in the standardized binocular image. Calculate the initial pose estimate of the composite structure locator based on the localization coding corner information. Based on this, determine the prediction search region of the circular feature structure in the standardized binocular image. S3. Extract sub-pixel feature points from the circular ring feature structure within the prediction search area, and perform three-dimensional reconstruction based on the extracted sub-pixel feature points of the circular ring and the calibration parameters of the binocular camera to obtain the observation three-dimensional point information of the circular ring feature structure. S4. Compare the observed three-dimensional point information with the preset three-dimensional point information of the composite structure locator, remove abnormal points whose spatial position error exceeds the preset threshold, and use the effective three-dimensional point set after removal to calculate the fine positioning pose result. S5. Perform integrity monitoring processing on the fine positioning pose results. The integrity monitoring processing includes at least the following indicators for consistency verification: binocular reprojection error, circular feature geometric fitting residual, and calculate the pose credibility index based on the consistency verification results.

2. The surgical instrument tracking method based on a composite structure locator according to claim 1, characterized in that, It also includes S6, where when the pose confidence index meets the preset confidence threshold condition, the precise positioning pose result is output as the final tracking pose of the surgical instrument; when the pose confidence index does not meet the preset confidence threshold condition, the current pose result is not output, the current frame is skipped, the next pair of binocular images is read, and steps S1 to S6 are re-executed.

3. The surgical instrument tracking method based on a composite structure locator according to claim 1, characterized in that, S1 specifically refers to: The left and right eye images, including the composite structure locator, are simultaneously acquired by a binocular vision acquisition device, and the same timestamp is added to each pair of images to obtain the original binocular image pairs. Based on preset camera intrinsic parameters, distortion correction and epipolar correction are performed on the original binocular image pairs to obtain corrected binocular image pairs, so as to ensure that the correspondence between the left and right eyes can be used for subsequent pose calculation. Denoising is performed on the calibrated binocular image pairs to obtain preprocessed binocular image pairs, thereby suppressing random noise from the image sensor; The preprocessed stereo image pairs are defined as normalized stereo images, and the normalized stereo images are used as input for the coarse localization step of the subsequent localization coding structure.

4. The surgical instrument tracking method based on a composite structure locator according to claim 1, characterized in that, S2 specifically refers to: In standardized binocular images, candidate regions for localization coding structures are detected to determine the effective region range that needs to be decoded. Decode and validate the effective region, and extract the corner pixel coordinates of the localization coding structure in the image. Based on the corner pixel coordinate information and the known geometric parameters of the positioning coding structure, the initial pose estimate of the composite structure locator is calculated. By utilizing the initial pose estimate and the predetermined relative geometric relationship between the localization coding structure and the circular feature structure, the predicted search region of the circular feature structure is determined in the standardized binocular image, and the predicted search region is used as the input for the fine localization step of the circular feature structure.

5. The surgical instrument tracking method based on a composite structure locator according to claim 1, characterized in that, S3 specifically refers to: Within the predicted search area, corner point detection and filtering are performed on the circular feature structure to extract checkerboard feature corner points that conform to the geometric distribution of the circular ring. Subpixel refinement is then performed to obtain the subpixel circular feature corner point set in the left and right eye images. Using the principle of binocular stereo vision triangulation, based on the subpixel circular feature corner point set in the left and right eye images and the binocular camera calibration parameters, the observation three-dimensional point set of the circular feature structure in the current camera coordinate system is reconstructed.

6. The surgical instrument tracking method based on a composite structure locator according to claim 1, characterized in that, S4 specifically refers to: Using the observed three-dimensional points calculated in step S3, a three-dimensional circle is fitted according to a preset radius. Points with fitting errors greater than a preset threshold are marked as outliers and removed. The remaining points are retained as a valid three-dimensional point set. A point-to-point registration model is established using the valid three-dimensional point set and the preset three-dimensional model points of the composite structure locator. The pose of the composite structure locator in the current camera coordinate system is directly solved to obtain the fine positioning pose result.

7. The surgical instrument tracking method based on a composite structure locator according to claim 1, characterized in that, S5 specifically refers to: A set of indicators is calculated for consistency verification of the fine positioning pose results. The indicator set includes: binocular reprojection error based on the fine positioning pose results and geometric fitting residual of the circular feature. The indicator set is compared with preset consistency conditions, and consistency verification is performed to obtain the verification result of each type of indicator. Based on the verification results of each consistency verification, a pose reliability index is generated to characterize the reliability of the fine positioning pose results.

8. A surgical instrument tracking method based on a composite structure locator according to claim 2, characterized in that, S6 specifically refers to: Determine whether the pose confidence index meets the preset confidence threshold condition; if it does, determine the fine positioning pose result as the final tracking pose and output it; if it does not meet the condition, determine that the current frame data is unreliable, discard the current calculation result, do not perform downgrade pose output, directly reset the system state and trigger the acquisition and processing process of the next frame image.

9. A surgical instrument tracking system based on a composite structure locator, employing a surgical instrument tracking method based on a composite structure locator according to any one of claims 1-8, characterized in that, include: The system includes a binocular vision acquisition device, a composite structure locator, an image preprocessing module, a coarse positioning module with a positioning coding structure, a fine positioning module with a circular ring feature structure, a 3D point cloud cleaning and pose calculation module, an integrity monitoring and consistency verification module, and a pose output control module.