A multimodal feature joint intraoperative medical instrument tracking method and system
Patent Information
- Application Number
- CN202610259177.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-04
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-03-04
AI Technical Summary
[0003]针对现有技术中的上述不足,本发明提供的一种多模态特征联合的术中医疗器械追踪方法及系统解决了现有医疗器械追踪方法难以同时满足临床对“零侵入、免训练、即拆即用”的要求的问题
[0015]本发明的有益效果为:本发明通过构建基于多模态观测集合与形状先验的联合概率模型,实现医疗器械在复杂术中环境下无标志点的定位,规避了传统定位方式依靠人工标志物带来的生物相容性与脱落风险,同时无需训练数据,满足医疗器械“即拆即用”的临床合规需求,实现高精度、低延迟、零侵入的术中医疗器械识别与追踪导航。
Smart Images

Figure CN122163321B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer-aided surgical navigation technology, specifically to a method and system for intraoperative medical device tracking based on multimodal feature integration. Background Technology
[0002] With the rapid development of microsurgery and minimally invasive surgery, the demand for real-time instrument positioning during surgery has increased to sub-millimeter accuracy and millisecond latency. However, existing technologies still largely rely on artificial markers or data-driven models, making it difficult to simultaneously meet clinical requirements for "zero invasiveness, no training required, and immediate usability." Traditional reflective balls, QR codes, or LED target point solutions not only increase the burden of preoperative preparation and disinfection but also pose biocompatibility risks; once the target point is obscured by blood or gauze, the system immediately fails. End-to-end pose regression methods based on deep learning require a large amount of labeled data and a lengthy training period; replacing new instruments necessitates re-collection, severely violating surgical compliance. Furthermore, single-modal vision systems are prone to feature loss in dark fields, highly reflective surfaces, or on transparent instrument surfaces, and common intraoperative occlusions further amplify mismatches and pose drift problems. Therefore, there is an urgent need for a clinical instrument tracking method that abandons artificial markers and training data, providing a feasible and robust solution for high-precision, low-latency, and zero-invasive instrument navigation in complex surgical fields. Summary of the Invention
[0003] To address the aforementioned shortcomings in existing technologies, this invention provides a multimodal feature-based intraoperative medical device tracking method and system that solves the problem that existing medical device tracking methods cannot simultaneously meet the clinical requirements of "zero invasiveness, no training required, and immediate use".
[0004] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows: A method for intraoperative medical device tracking using multimodal feature combination is provided, comprising the following steps: Register the spatial coordinate systems of the depth camera and the binocular microscope; Extracting multimodal features of a 3D model of a medical device before surgery; these multimodal features include feature point modal features, edge modal features, and depth modal features; Real-time images of the surgical area are captured using a depth camera, and depth modal features are extracted from them. The surgical field images of medical instruments performing surgical operations are acquired using a binocular microscope, and the feature point modal features and edge modal features are extracted from them; A joint objective function is constructed by establishing one-to-one and / or one-to-many matching between the multimodal features from images captured by depth cameras and binocular microscopes and the multimodal features of the 3D model of medical devices before surgery; Using the 6-DOF pose parameters of the medical device as optimization variables, the joint objective function is solved to obtain the pose of the medical device during surgery and complete the intraoperative tracking of the medical device.
[0005] Furthermore, specific methods for registering the spatial coordinate systems of the depth camera and the binocular microscope include: By placing the stereo calibration plate simultaneously within the common field of view of the depth camera and the binocular microscope, and acquiring multi-view images, the internal and external parameters of the depth camera and the binocular microscope are jointly calibrated. This yields a rigid transformation matrix from the pixel coordinate system of the depth camera and the binocular microscope to a unified coordinate system, enabling the observation data from the depth camera and the binocular microscope to be mapped to a unified spatial coordinate system, thus providing a basis for spatial consistency.
[0006] Furthermore, the specific method for extracting multimodal features of a 3D model of a medical device before surgery includes the following steps: In the 3D CAD file of medical devices, the extreme points of Gaussian curvature and geometric discontinuities are extracted by parsing the CAD triangular mesh and used as the first feature points. Operators with scale invariance and rotation invariance are used to calculate the descriptor of the first feature points. The first feature points that are stable in more than 80% of the simulated viewpoints are selected to construct a spatial index structure, thus completing the feature point modal feature extraction of the 3D model of the medical device before surgery. In the 3D CAD file of medical devices, the crease edges and contour edges that meet the preset angle threshold are parameterized into straight lines / spline curves and discretized according to the preset density. Tangential / normal covariance coding uncertainty is added to extract the edge modal features of the 3D model of the medical device before surgery. In the 3D CAD file of the medical device, depth maps and normal maps from several representative viewpoints are rendered offline to construct a viewpoint-related depth template library. Based on the initial pose value of the medical device, the nearest neighbor templates of a preset number are called for interpolation to synthesize the predicted depth map and the corresponding surface normal in real time. Combined with the 3D observation point cloud obtained by the depth camera in real time during surgery and back-projected, and the local tangent plane of the model determined by the predicted depth map and the corresponding surface normal obtained by interpolation synthesis, the residual from the midpoint of the 3D observation point cloud to the local tangent plane of the model is calculated to obtain the depth modal features of the 3D model of the medical device before surgery. The initial pose value of the medical device is obtained through global feature matching in the first frame of tracking startup, which is acquired before surgery.
[0007] Furthermore, the method for extracting depth modal features from real-time images captured by the depth camera is as follows: The depth observation value in the real-time image captured by the depth camera is directly measured by infrared, structured light or Time-of-Flight principle, and the depth observation value is used as the depth modal feature in the real-time image captured by the depth camera.
[0008] Furthermore, the method for extracting the modal features of feature points in the surgical field image acquired by binocular microscope includes the following steps: Extract candidate target regions from the acquired surgical field images; Within the candidate target region, the gradient magnitude of the pixels is calculated based on the image grayscale information. Points that satisfy the condition that the gradient magnitude of the neighborhood is greater than a preset threshold and that satisfy the preset tracking stability condition in a set number of consecutive frames are taken as the second feature points. A directional consistency analysis is performed on the second feature point to obtain a set of three-dimensional feature points containing positional information and local geometric attributes, thus obtaining the modal features of the feature points in the acquired surgical field.
[0009] Furthermore, the method for extracting edge modal features from the surgical field image acquired by binocular microscopy includes the following steps: Extract candidate target regions from the acquired surgical field images; Continuous boundary information is extracted within the candidate target area, and parametric modeling is performed based on geometric consistency to generate three-dimensional linear or curved edge observations, thereby obtaining edge modal features in the acquired surgical field images.
[0010] Furthermore, the method for constructing a joint objective function by establishing one-to-one and / or one-to-many matching between the multimodal features from images acquired by depth cameras and binocular microscopes and the multimodal features of the medical device 3D model before surgery includes the following steps: The depth consistency error between the depth modal features extracted from the images captured by the depth camera and the depth modal features of the 3D model of the medical device before surgery is calculated, and then the depth modal observation residual is constructed. The modal features of feature points extracted from images captured by a binocular microscope are compared with the spatial index structure using a nearest neighbor search. Feature points whose descriptor feature distance is less than [a certain value] are selected. And the distance ratio is less than The corresponding first feature points are associated to establish a relationship, and then point-level observation residuals are constructed; whereby and Preset; The edge modal features extracted from images captured by a binocular microscope are compared with the edge modal features of the medical device's 3D model before surgery using a bidirectional nearest neighbor search. The search aims to find the nearest neighbor with a directional error less than [a certain value]. And the distance is less than The edge modal features correspond one-to-one, thus constructing a linear residual; where and Preset; Constructing a joint objective function Its expression is:
[0011] Where ln represents the natural logarithm; These are the 6-DOF pose parameters for medical devices; This represents the likelihood of the overall observation under the condition that each modal characteristic is independent. , , and These represent the sub-probability likelihood terms for feature point modal features, edge modal features, and depth modal features under conditional independence, respectively. Adaptive weights for modal features of feature points. , The basic weights of the modal features of feature points. For natural index; The sensitivity coefficient for the modal features of feature points; The residuals are point-level observations. Adaptive weights for edge modality features , The basic weights for edge modality features, The sensitivity coefficient for edge modality features; For linear residuals; Adaptive weights for deep modal features , The basic weights for deep modal features, The sensitivity coefficient for deep modal features; The depth modal observation residuals are represented by: i, j, and k, which represent the modal feature of the i-th feature point, the j-th edge modal feature, and the k-th depth modal feature, respectively; and F, E, and D, which represent the set of modal features of feature points, the set of edge modal features, and the set of depth modal features, respectively. This represents a constant term that is independent of the pose parameters.
[0012] Furthermore, when solving the joint objective function, if the adaptive weight of any modal feature is lower than the preset lower limit, the basic weight of that modal feature in the next frame optimization is reduced to avoid error amplification caused by occlusion, motion blur or depth loss. Methods for solving the joint objective function include Newton's method and the Levenberg-Marquardt method; the condition for outputting the current solution value and using it as the pose of the medical device during surgery is that any of the following conditions are met: Condition 1: The relative decrease in the joint objective function value between two adjacent iterations is less than the first threshold; Condition 2: The six-degree-of-freedom pose increment is less than the second threshold; Condition 3: The gradient norm of the joint objective function with respect to the pose parameters is less than the third threshold; Condition 4: The current iteration count has reached the preset limit.
[0013] Furthermore, after obtaining the position of the medical instruments during surgery, the following operations are performed: The position of the medical device during surgery is mapped to the patient tracking coordinate system in real time through a preoperative calibration matrix, and a semi-transparent virtual model with a 1:1 size matching the medical device is drawn on the display screen; when the medical device moves or rotates, the semi-transparent virtual model is updated with the same 6 degrees of freedom, realizing zero-marker, zero-invasive augmented reality guidance.
[0014] A system for intraoperative medical device tracking based on multimodal feature integration is provided, comprising: Depth camera, used to capture real-time images of the surgical area; A binocular microscope is used to capture images of the surgical field when medical instruments are used for surgical procedures. The registration module is used to register the spatial coordinate system of the depth camera with that of the binocular microscope; The first modal feature extraction module is used to extract the multimodal features of the 3D model of the medical device before surgery; The second modal feature extraction module is used to extract the depth modal features from the real-time images of the surgical area captured by the depth camera; The third modal feature extraction module is used to extract feature point modal features and edge modal features from the surgical field image of medical devices performing surgical operations acquired by a binocular microscope; The pose calculation module is used to establish one-to-one and / or one-to-many matching between the multimodal features from the images captured by the depth camera and binocular microscope and the multimodal features of the 3D model of the medical device before surgery, and construct a joint objective function; and solve the joint objective function with the 6-DOF pose parameters of the medical device as optimization variables to obtain the pose of the medical device during surgery and complete the intraoperative medical device tracking. The augmented reality guidance module is used to map the position of medical devices during surgery to the patient tracking coordinate system in real time through a preoperative calibration matrix, and draw a semi-transparent virtual model on the display screen that matches the size of the medical device at a 1:1 scale. When the medical device moves or rotates, the semi-transparent virtual model is updated with the same 6 degrees of freedom, realizing zero-marker, zero-intrusion augmented reality guidance.
[0015] The beneficial effects of this invention are as follows: By constructing a joint probability model based on multimodal observation sets and shape priors, this invention enables the positioning of medical devices in complex intraoperative environments without markers, avoiding the biocompatibility and detachment risks associated with traditional positioning methods that rely on artificial markers. At the same time, no training data is required, meeting the clinical compliance requirements of "ready to use" medical devices and achieving high-precision, low-latency, and non-invasive intraoperative medical device identification and tracking navigation. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the method. Detailed Implementation
[0017] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0018] like Figure 1 As shown, this multimodal feature-based intraoperative medical device tracking method includes the following steps: S1. Register the spatial coordinate system of the depth camera and the binocular microscope; S2. Extract multimodal features of the 3D model of the medical device before surgery; the multimodal features include feature point modal features, edge modal features, and depth modal features; S3. Acquire real-time images of the surgical area using a depth camera and extract the depth modal features from them; S4. Acquire surgical field images of medical instruments performing surgical operations using a binocular microscope, and extract feature point modal features and edge modal features from them; S5. Establish one-to-one and / or one-to-many matching between the multimodal features from the images captured by the depth camera and binocular microscope and the multimodal features of the medical device 3D model before surgery, and construct a joint objective function; S6. Using the 6-DOF pose parameters of the medical device as optimization variables, solve the joint objective function to obtain the pose of the medical device during surgery and complete the intraoperative medical device tracking. S7. The position of the medical device during surgery is mapped to the patient tracking coordinate system in real time through the preoperative calibration matrix, and a semi-transparent virtual model with a 1:1 size matching the medical device is drawn on the display screen; when the medical device moves or rotates, the semi-transparent virtual model is updated with the same 6 degrees of freedom, realizing zero-marker, zero-invasive augmented reality guidance.
[0019] Based on this method, the intraoperative medical device tracking system combining multimodal features includes: A depth camera is used to capture real-time images of the surgical area, specifically 640×480 RGB-D frames; A binocular microscope is used to acquire surgical field images of medical instruments performing surgical procedures, specifically 2×1920×1080 high-resolution RGB images; The registration module is used to register the spatial coordinate system of the depth camera with that of the binocular microscope; The first modal feature extraction module is used to extract the multimodal features of the 3D model of the medical device before surgery; The second modal feature extraction module is used to extract the depth modal features from the real-time images of the surgical area captured by the depth camera; The third modal feature extraction module is used to extract feature point modal features and edge modal features from the surgical field image of medical devices performing surgical operations acquired by a binocular microscope; The pose calculation module is used to establish one-to-one and / or one-to-many matching between the multimodal features from the images captured by the depth camera and binocular microscope and the multimodal features of the 3D model of the medical device before surgery, and construct a joint objective function; and solve the joint objective function with the 6-DOF pose parameters of the medical device as optimization variables to obtain the pose of the medical device during surgery and complete the intraoperative medical device tracking. The augmented reality guidance module is used to map the position of medical devices during surgery to the patient tracking coordinate system in real time through a preoperative calibration matrix, and draw a semi-transparent virtual model on the display screen that matches the size of the medical device at a 1:1 scale. When the medical device moves or rotates, the semi-transparent virtual model is updated with the same 6 degrees of freedom, realizing zero-marker, zero-intrusion augmented reality guidance.
[0020] In some embodiments, a specific method for registering the spatial coordinate systems of a depth camera and a binocular microscope includes: By simultaneously placing a stereo calibration board within the common field of view of both the depth camera and the binocular microscope, and acquiring multi-view images, the intrinsic and extrinsic parameters of both the depth camera and the binocular microscope are jointly calibrated. This yields a rigid transformation matrix from the pixel coordinate systems of the depth camera and the binocular microscope to a unified coordinate system, mapping the observation data from both systems to a unified spatial coordinate system and providing a basis for spatial consistency. The rigid transformation matrix can directly map the RGB-D point cloud and the left view of the binocular microscope to the same world coordinate system, achieving a one-to-one registration between pixels and point clouds. At this point, any pixel can simultaneously acquire color, parallax, and 3D coordinates, providing a spatially consistent data foundation for subsequent multimodal feature extraction.
[0021] In some embodiments, in order to construct a multimodal feature prior of a medical device, a three-dimensional CAD model of the medical device to be tracked can be imported into the system in the preoperative stage, and geometric analysis can be performed on the three-dimensional CAD model to extract geometric primitives and feature information related to the structural stability of the medical device, including but not limited to key point locations, linear or axisymmetric structures, and surface depth distribution characteristics.
[0022] For example, a specific method for extracting multimodal features of a 3D model of a medical device before surgery includes the following steps: A1. In the 3D CAD file of medical devices, Gaussian curvature extrema points and geometric discontinuities are extracted by parsing the CAD triangular mesh and used as the first feature points. Operators with scale invariance and rotation invariance (such as ORB or SIFT) are used to calculate the descriptor of the first feature points. The first feature points that are stable in more than 80% of the simulated viewpoints are selected to construct a spatial index structure, thus completing the feature point modal feature extraction of the 3D model of the medical device before surgery. For example, the 3D CAD model of the medical device is first rendered into a sequence of 2D images from multiple simulated perspectives. The 3D first feature points extracted based on Gaussian curvature are projected onto these 2D images. Within the local image neighborhood of the projection location, a feature vector (i.e., the first feature point descriptor) with rotation and scale invariance is calculated using the ORB (based on gray-scale centroid and binary test) or SIFT (based on gradient histogram) algorithm. If the ORB operator is used, specific point pairs are selected in the neighborhood for pixel brightness comparison (BRIEF algorithm) to generate a binary string (e.g., 256 bits) as the descriptor for that perspective. If the SIFT operator is used, the neighborhood is divided into 4×4 sub-regions, generating a 128-dimensional floating-point vector, which is then normalized to adapt to changes in illumination (scale invariance) to obtain the first feature point descriptor.
[0023] A2. In the 3D CAD file of medical devices, the crease edges and contour edges that meet the preset angle threshold (e.g., dihedral angle > 15°) are parameterized into straight lines / spline curves and discretized according to the preset density (e.g., 3-5 points per millimeter). Tangential / normal covariance coding uncertainty is added to extract the edge modal features of the 3D model of the medical device before surgery. A3. In the 3D CAD file of the medical device, depth maps and normal maps of several (e.g., 50) representative viewpoints are rendered offline to construct a viewpoint-related depth template library. Based on the initial pose value of the medical device, a preset number (e.g., 3) of nearest neighbor templates are called for interpolation to synthesize the predicted depth map and the corresponding surface normal in real time. Combined with the 3D observation point cloud obtained by the depth camera in real time during the operation and back-projected, and the predicted depth map and the corresponding surface normal (surface normal map) obtained by interpolation, the local tangent plane of the model is determined. The residual from the midpoint of the 3D observation point cloud to the local tangent plane of the model is calculated to obtain the depth modal features of the 3D model of the medical device before the operation. The initial pose value of the medical device is obtained through global feature matching (e.g., FPFH) in the first frame of tracking start, which is obtained before the operation.
[0024] It should be noted that the depth modal features of the medical device 3D model before surgery used the first frame image data from the depth camera at the start of medical device tracking. In subsequent tracking phases, the pose calculation results of the previous frame were used as the initial prediction values for the current frame. The preoperative phase mainly focuses on building an offline depth template library; while the processes of "interpolation" and "real-time synthesis" occur during the intraoperative real-time solution phase, used to generate the desired observations under the current pose.
[0025] In practical implementation, extracting depth modal features during surgery is to obtain the true depth information of the medical device surface. For example, the method for extracting depth modal features from real-time images captured by a depth camera is as follows: This method involves directly measuring the depth observation value in real-time images captured by a depth camera using infrared, structured light, or Time-of-Flight (ToF) principles, and then using this depth observation value as the depth modal feature of the real-time images captured by the depth camera. This is a relatively conventional method and will not be elaborated further.
[0026] In practical implementation, extracting modal features of feature points mainly involves extracting the features of three-dimensional feature points with stable geometric distribution. For example, the method for extracting modal features of feature points in the surgical field image acquired by a binocular microscope includes the following steps: B1. Extract candidate target regions from the acquired surgical field images; B2. Within the candidate target region, based on image grayscale information, calculate the gradient magnitude of pixels. Points whose neighborhood gradient magnitude is greater than a preset threshold and who satisfy a preset tracking stability condition in a set number of consecutive frames (e.g., 5 frames) are designated as second feature points. The tracking stability condition can be set to a matching success rate greater than 60%. B3. Perform directional consistency analysis on the second feature point to obtain a set of three-dimensional feature points containing positional information and local geometric attributes, thereby obtaining the modal features of the feature points in the acquired surgical field.
[0027] In practical implementation, edge modal feature extraction mainly involves extracting continuous contours or linear structures and constructing corresponding geometric descriptions to form features. For example, the method for extracting edge modal features from surgical field images acquired by a binocular microscope includes the following steps: C1. Extract candidate target regions from the acquired surgical field images; C2. Extract continuous boundary information within the candidate target area and perform parametric modeling based on geometric consistency to generate three-dimensional linear or curved edge observations, thereby obtaining edge modal features in the acquired surgical field images.
[0028] In some embodiments, medical devices (such as surgical forceps, scissors, etc.) often have symmetrical structures. A certain feature point or local edge extracted during surgery may have multiple similar candidate matching points on the three-dimensional model. In this case, it is necessary to establish a preliminary association through a 'one-to-many' relationship, and then use multimodal constraints (such as depth information and motion priors) in the joint probability model for screening.
[0029] For example, the specific method for constructing a joint objective function by establishing one-to-one and / or one-to-many matching between multimodal features from images acquired by depth cameras and binocular microscopes and multimodal features of a 3D model of a medical device before surgery includes the following steps: D1. Calculate the depth consistency error between the depth modal features extracted from the images captured by the depth camera and the depth modal features of the 3D model of the medical device before surgery, and then construct the depth modal observation residual. D2. Perform a nearest neighbor search on the modal features of the feature points extracted from the images captured by the binocular microscope and the spatial index structure, selecting feature points whose Euclidean distance is less than 1. And the distance ratio is less than The corresponding first feature points are associated to establish a relationship, and then point-level observation residuals are constructed; whereby and Preset; D3. Perform a bidirectional nearest neighbor search between the edge modal features extracted from the images captured by the binocular microscope and the edge modal features of the medical device 3D model before surgery, ensuring that the directional error is less than [a certain value]. And the distance is less than The edge modal features correspond one-to-one, thus constructing a linear residual; where and Preset; When constructing line-level residuals, an edge observation in the image often corresponds to multiple discrete sampling points on the model. This is a typical 'one-to-many' association, used to calculate the shortest distance from the observed edge to the model surface. In the presence of occlusion or noise, retaining the 'one-to-many' matching assumption can avoid pose jumps caused by mismatches, thereby ensuring the continuity and stability of tracking.
[0030] D4. Construct a joint objective function (based on a joint probability model of multimodal observation set and shape prior). Its expression is:
[0031] Where ln represents the natural logarithm; These are the 6-DOF pose parameters for medical devices; This represents the likelihood of the overall observation under the condition that each modal characteristic is independent. , , and These represent the sub-probability likelihood terms for feature point modal features, edge modal features, and depth modal features under conditional independence, respectively. Adaptive weights for modal features of feature points. , The basic weights of the modal features of feature points. For natural index; The sensitivity coefficient for the modal features of feature points; The residuals are point-level observations. Adaptive weights for edge modality features , The basic weights for edge modality features, The sensitivity coefficient for edge modality features; For linear residuals; Adaptive weights for deep modal features , The basic weights for deep modal features, The sensitivity coefficient for deep modal features; The depth modal observation residuals are represented by: i, j, and k, which represent the modal feature of the i-th feature point, the j-th edge modal feature, and the k-th depth modal feature, respectively; and F, E, and D, which represent the set of modal features of feature points, the set of edge modal features, and the set of depth modal features, respectively. This represents a constant term that is independent of the pose parameters.
[0032] In the adaptive weights of modal features, the exponential part of the exponent decays exponentially as the residual increases, thereby reducing the weight value of the modal feature. This enables dynamic calculation of the correspondence and weight allocation of different modal features during the operation: features with stable observations, high confidence, and low residuals are given high weights to dominate the optimization process; features that are occluded, have weak textures, blurred contours, or high depth noise are automatically weighted to suppress interference from abnormal information.
[0033] When solving the joint objective function, if the adaptive weight of any modal feature is lower than the preset lower limit, the basic weight of that modal feature in the next frame optimization is reduced to avoid error amplification caused by occlusion, motion blur or depth loss, while ensuring that the mathematical solvability of the model is maintained even when some markers are covered by blood or gauze.
[0034] Methods for solving the joint objective function include Newton's method and the Levenberg-Marquardt method; the condition for outputting the current solution value and using it as the pose of the medical device during surgery is that any of the following conditions are met: Condition 1: The relative decrease in the joint objective function value between two adjacent iterations is less than the first threshold; Condition 2: The six-degree-of-freedom pose increment is less than the second threshold; Condition 3: The gradient norm of the joint objective function with respect to the pose parameters is less than the third threshold; Condition 4: The current iteration count has reached the preset limit.
[0035] In summary, this invention constructs a joint probability model based on multimodal observation sets and shape priors to achieve landmark-free positioning of medical devices in complex intraoperative environments. This avoids the biocompatibility and detachment risks associated with traditional positioning methods that rely on artificial markers. Furthermore, it requires no training data, meeting the clinical compliance requirements for "ready-to-use" medical devices and achieving high-precision, low-latency, and non-invasive intraoperative medical device identification and tracking navigation.
Claims
1. A system for intraoperative medical device tracking based on multimodal feature combination, characterized in that, include: Depth camera, used to capture real-time images of the surgical area; A binocular microscope is used to capture images of the surgical field when medical instruments are used for surgical procedures. The registration module is used to register the spatial coordinate system of the depth camera with that of the binocular microscope; The first modal feature extraction module is used to extract the multimodal features of the 3D model of the medical device before surgery; The second modal feature extraction module is used to extract the depth modal features from the real-time images of the surgical area captured by the depth camera; The third modal feature extraction module is used to extract feature point modal features and edge modal features from the surgical field image of medical devices performing surgical operations acquired by a binocular microscope; The pose calculation module is used to establish one-to-one and / or one-to-many matching between the multimodal features from the images captured by the depth camera and binocular microscope and the multimodal features of the 3D model of the medical device before surgery, and construct a joint objective function; and solve the joint objective function with the 6-DOF pose parameters of the medical device as optimization variables to obtain the pose of the medical device during surgery and complete the intraoperative medical device tracking. The augmented reality guidance module is used to map the position of medical devices during surgery to the patient tracking coordinate system in real time through a preoperative calibration matrix, and draw a semi-transparent virtual model that matches the size of the medical device at a 1:1 scale on the display screen; when the medical device moves or rotates, the semi-transparent virtual model is updated with the same 6 degrees of freedom, realizing zero-marker, zero-intrusion augmented reality guidance. The intraoperative medical device tracking method based on multimodal feature integration includes the following steps: Register the spatial coordinate system of the depth camera with that of the binocular microscope; Extracting multimodal features of a 3D model of a medical device before surgery; these multimodal features include feature point modal features, edge modal features, and depth modal features; Real-time images of the surgical area are captured using a depth camera, and depth modal features are extracted from them. The surgical field images of medical instruments performing surgical operations are acquired using a binocular microscope, and the feature point modal features and edge modal features are extracted from them; A joint objective function is constructed by establishing one-to-one and / or one-to-many matching between the multimodal features from images captured by depth cameras and binocular microscopes and the multimodal features of the 3D model of medical devices before surgery; Using the 6-DOF pose parameters of the medical device as optimization variables, the joint objective function is solved to obtain the pose of the medical device during surgery and complete the intraoperative tracking of the medical device. The specific method for extracting multimodal features of a 3D model of a medical device before surgery includes the following steps: In the 3D CAD file of medical devices, the extreme points of Gaussian curvature and geometric discontinuities are extracted by parsing the CAD triangular mesh and used as the first feature points. Operators with scale invariance and rotation invariance are used to calculate the descriptor of the first feature points. The first feature points that are stable in more than 80% of the simulated viewpoints are selected to construct a spatial index structure, thus completing the feature point modal feature extraction of the 3D model of the medical device before surgery. In the 3D CAD file of medical devices, the crease edges and contour edges that meet the preset angle threshold are parameterized into straight lines / spline curves and discretized according to the preset density. Tangential / normal covariance coding uncertainty is added to extract the edge modal features of the 3D model of the medical device before surgery. In the 3D CAD file of medical devices, depth maps and normal maps from several representative viewpoints are rendered offline to construct a viewpoint-related depth template library. Based on the initial pose of the medical device, a preset number of nearest neighbor templates are called for interpolation to synthesize a predicted depth map and the corresponding surface normal in real time. Combining the 3D observation point cloud obtained by the depth camera in real time during surgery and back-projected with the predicted depth map and the corresponding surface normal determined by the model's local tangent plane, the residual from the midpoint of the 3D observation point cloud to the local tangent plane of the model is calculated to obtain the depth modal features of the medical device's 3D model before surgery. The initial pose of the medical device is obtained through global feature matching in the first frame of tracking startup, which is acquired before surgery. The specific method for constructing a joint objective function by establishing one-to-one and / or one-to-many matching between the multimodal features from images acquired by depth cameras and binocular microscopes and the multimodal features of the medical device 3D model before surgery includes the following steps: The depth consistency error between the depth modal features extracted from the images captured by the depth camera and the depth modal features of the 3D model of the medical device before surgery is calculated, and then the depth modal observation residual is constructed. The modal features of feature points extracted from images captured by a binocular microscope are compared with the spatial index structure using a nearest neighbor search. Feature points whose descriptor feature distance is less than [a certain value] are selected. And the distance ratio is less than The corresponding first feature points are associated to establish a relationship, and then point-level observation residuals are constructed; whereby and Preset; The edge modal features extracted from images captured by a binocular microscope are compared with the edge modal features of the medical device's 3D model before surgery using a bidirectional nearest neighbor search. This search aims to find the nearest neighbor with a directional error less than [a certain value]. And the distance is less than The edge modal features correspond one-to-one, thus constructing a linear residual; where and Preset; Constructing a joint objective function Its expression is: Where ln represents the natural logarithm; These are the 6-DOF pose parameters for medical devices; This represents the likelihood of the overall observation under the condition that each modal characteristic is independent. , , and These represent the sub-probability likelihood terms for feature point modal features, edge modal features, and depth modal features under conditional independence, respectively. Adaptive weights for modal features of feature points. , The basic weights for the modal features of feature points. For natural index; The sensitivity coefficient for the modal features of feature points; The residuals are point-level observations. Adaptive weights for edge modality features , The basic weights for edge modality features, The sensitivity coefficient for edge modality features; For linear residuals; Adaptive weights for deep modal features , The basic weights for deep modal features, The sensitivity coefficient for deep modal features; The depth modal observation residuals are represented by: i, j, and k, which represent the modal feature of the i-th feature point, the j-th edge modal feature, and the k-th depth modal feature, respectively; and F, E, and D, which represent the set of modal features of feature points, the set of edge modal features, and the set of depth modal features, respectively. This represents a constant term that is independent of the pose parameters.
2. The system according to claim 1, characterized in that, Specific methods for registering the spatial coordinate systems of a depth camera and a binocular microscope include: By placing the stereo calibration plate simultaneously within the common field of view of the depth camera and the binocular microscope, and acquiring multi-view images, the internal and external parameters of the depth camera and the binocular microscope are jointly calibrated. This yields a rigid transformation matrix from the pixel coordinate system of the depth camera and the binocular microscope to a unified coordinate system, enabling the observation data from the depth camera and the binocular microscope to be mapped to a unified spatial coordinate system, thus providing a basis for spatial consistency.
3. The system according to claim 1, characterized in that, The method for extracting depth modal features from real-time images captured by a depth camera is as follows: The depth observation value in the real-time image captured by the depth camera is directly measured by infrared, structured light or Time-of-Flight principle, and the depth observation value is used as the depth modal feature in the real-time image captured by the depth camera.
4. The system according to claim 1, characterized in that, The method for extracting modal features of feature points in surgical field images acquired by binocular microscopes includes the following steps: Extract candidate target regions from the acquired surgical field images; Within the candidate target region, the gradient magnitude of the pixels is calculated based on the image grayscale information. Points that satisfy the condition that the gradient magnitude of the neighborhood is greater than a preset threshold and that satisfy the preset tracking stability condition in a set number of consecutive frames are taken as the second feature points. A directional consistency analysis is performed on the second feature point to obtain a set of three-dimensional feature points containing positional information and local geometric attributes, thus obtaining the modal features of the feature points in the acquired surgical field.
5. The system according to claim 1, characterized in that, The method for extracting edge modal features from surgical field images acquired by binocular microscopes includes the following steps: Extract candidate target regions from the acquired surgical field images; Continuous boundary information is extracted within the candidate target area, and parametric modeling is performed based on geometric consistency to generate three-dimensional linear or curved edge observations, thereby obtaining edge modal features in the acquired surgical field images.
6. The system according to claim 1, characterized in that, When solving the joint objective function, if the adaptive weight of any modal feature is lower than the preset lower limit, the basic weight of that modal feature in the next frame optimization is reduced to avoid error amplification caused by occlusion, motion blur or depth loss. Methods for solving the joint objective function include Newton's method and the Levenberg-Marquardt method; the condition for outputting the current solution value and using it as the pose of the medical device during surgery is that any of the following conditions are met: Condition 1: The relative decrease in the joint objective function value between two adjacent iterations is less than the first threshold; Condition 2: The six-degree-of-freedom pose increment is less than the second threshold; Condition 3: The gradient norm of the joint objective function with respect to the pose parameters is less than the third threshold; Condition 4: The current iteration count has reached the preset limit.
7. The system according to any one of claims 1 to 6, characterized in that, After obtaining the position of the medical instruments during surgery, the following procedures are performed: The position of the medical device during surgery is mapped to the patient tracking coordinate system in real time through a preoperative calibration matrix, and a semi-transparent virtual model with a 1:1 size matching the medical device is drawn on the display screen; when the medical device moves or rotates, the semi-transparent virtual model is updated with the same 6 degrees of freedom, realizing zero-marker, zero-invasive augmented reality guidance.
Citation Information
Patent Citations
Attitude estimation method based on multi-sensor multi-feature modular fusion
CN115690550A
Surgical navigation image real-time registration method and system based on multi-mode body surface mark tracking
CN120599006A