Surgical instrument recognition system and method based on multiple visual angles

Through the multi-view surgical instrument recognition system, combined with multi-view image data and geometric consistency, the problems of insufficient accuracy and robustness of single-view recognition methods in complex surgical scenarios are solved, and higher recognition accuracy and stability are achieved.

CN120725972AInactive Publication Date: 2025-09-30DONGYANG PEOPLES HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510771665.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing single-view surgical instrument recognition methods suffer from insufficient recognition accuracy and robustness in complex surgical scenarios due to occlusion, appearance similarity, and environmental changes. In particular, it is difficult to accurately determine the presence and category of instruments when multiple instruments are operated simultaneously or when the instruments and tissues are severely occluded.

Method used

A multi-view recognition system is adopted to acquire multi-view image data, combine cross-view geometric consistency and single-view visual recognition information, integrate multi-view image information, use geometric consistency constraints and visual features, detect candidate areas and perform cross-view association, and fuse recognition scores to improve recognition accuracy and robustness.

Benefits of technology

It significantly improves the recognition accuracy and robustness of surgical instruments in complex scenarios, effectively filters out single-view false detections, improves recognition accuracy in cases of occlusion and similar appearance, and enhances the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725972A_ABST
    Figure CN120725972A_ABST
Patent Text Reader

Abstract

The invention provides a multi-view-angle-based surgical instrument recognition system and method, relates to the technical field of medical management, and is used for solving the technical problem that in the related technology, a single-view-angle-based surgical instrument recognition method is insufficient in recognition precision and robustness in the surgical process. The acquisition module is used for acquiring multi-view image data and camera parameters; the preprocessing module is used for detecting a candidate area from each view angle image and extracting initial identification information; the cross-view-angle analysis module is used for determining a spatial relationship between view angles and acquiring geometric consistency between cross-view-angle candidate regions based on the spatial relationship; the multi-view integration module is used for associating the grouping candidate areas based on the geometric consistency, fusing the geometric consistency and the initial identification information to obtain combined identification scores, and aggregating the combined identification scores in the grouping to obtain aggregated identification scores; the judgment module is used for judging the ID and the recognition confidence of the surgical instrument according to the aggregated recognition score; and an output module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of medical management technology, and in particular to a surgical instrument recognition system and method based on multiple perspectives. Background Art

[0002] In minimally invasive surgery (MIS), accurate and real-time identification of surgical instruments is crucial for improving surgical efficiency, ensuring patient safety, and enabling surgical automation. This instrument identification information can be used for a variety of applications, including surgical robot control, intraoperative navigation, and postoperative process analysis.

[0003] Currently, technical solutions for surgical instrument recognition primarily rely on analyzing and processing images from a single viewpoint. These methods typically involve image preprocessing, extracting visual features of the instrument (such as shape, texture, and color), and classifying or detecting the surgical instrument using pattern recognition or machine learning algorithms (such as support vector machines and deep learning networks). For example, some methods use convolutional neural networks for object detection in endoscopic images to select and identify surgical instruments within the image.

[0004] However, there are some inherent limitations in single-view surgical instrument recognition methods. Frequent occlusions between instruments and between instruments and tissues during surgery are common phenomena. Single-view information cannot effectively handle instrument recognition problems under severe occlusion, which can easily lead to missed detections or false detections. Many surgical instruments are similar in appearance, and it is difficult to accurately distinguish them based solely on the visual features of a single view. In addition, the surgical environment is complex and changeable. Factors such as lighting changes, reflections, and smoke will affect image quality, further reducing the robustness of single-view recognition. Although some methods attempt to enhance recognition through the temporal information of sequential images, this cannot fundamentally solve the problem of information loss caused by physical occlusion or viewing angle limitations.

[0005] Therefore, the accuracy and robustness of existing surgical instrument recognition systems and methods based on a single viewpoint still need to be improved in complex surgical scenarios. This is especially true when multiple instruments are being operated simultaneously or when instruments are severely occluded by each other or by tissue. It is difficult to accurately determine the presence and type of an instrument based solely on local information from a single viewpoint. Summary of the Invention

[0006] The embodiments of the present application provide a multi-perspective surgical instrument recognition system and method, which are used to improve the technical problems of insufficient recognition accuracy and robustness caused by occlusion, appearance similarity, and complex environmental changes faced by surgical instrument recognition methods based on a single perspective during surgery.

[0007] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions: In the first aspect, the present application provides a surgical instrument recognition system based on multiple perspectives, including: an acquisition module for acquiring multi-perspective image data and camera parameters; a preprocessing module for detecting candidate areas from each perspective image and extracting initial recognition information; a cross-perspective analysis module for determining the spatial relationship between perspectives and obtaining the geometric consistency between cross-perspective candidate areas based on the spatial relationship; a multi-perspective integration module for associating and grouping the candidate areas based on the geometric consistency, fusing the geometric consistency with the initial recognition information to obtain a combined recognition score, and aggregating the combined recognition scores within the group to obtain an aggregated recognition score; a determination module for determining the surgical instrument ID and recognition confidence based on the aggregated recognition score; and an output module.

[0008] The multi-view surgical instrument recognition system provided by this application can significantly improve the recognition accuracy and robustness of surgical instruments in complex surgical scenarios by effectively integrating image information from multiple perspectives, particularly by combining geometric consistency constraints across perspectives with visual recognition information from a single perspective. By detecting candidate regions from multiple perspectives and analyzing their cross-perspective geometric relationships, false detections from a single perspective can be effectively filtered out, and cross-perspective candidate regions belonging to the same physical instrument can be associated for collaborative recognition, enabling the system to improve recognition accuracy when instruments are occluded and have high appearance similarity.

[0009] In one possible implementation of the first aspect, the determination module is further configured to obtain a historical recognition frequency of the surgical instrument ID and to modify the aggregate recognition score based on the historical recognition frequency before making a determination. By incorporating the historical recognition frequency as prior information to modify the recognition result of the current frame, the system can leverage contextual information from the time series to improve the stability of the recognition results, reduce transient misjudgments or recognition jitter, and further enhance the robustness of the system.

[0010] In a possible implementation of the first aspect, the cross-perspective analysis module obtains the geometric consistency based on the spatial distance between the source perspective candidate area position information and the target perspective candidate area position information after the source perspective candidate area position information is projected to the target perspective using the spatial relationship, and the geometric consistency is negatively correlated with the spatial distance.

[0011] In this way, the method of obtaining geometric consistency based on the projection space distance can quantify the degree of spatial correspondence of candidate regions under different perspectives, providing a basis for subsequent cross-perspective association and information fusion.

[0012] In a possible implementation of the first aspect, the multi-perspective integration module obtains a cross-perspective consistency score based on the geometric consistency, wherein the cross-perspective consistency score is obtained for any candidate area by traversing other perspectives different from the current perspective and determining the maximum geometric consistency of other perspective candidate areas that have the highest geometric consistency with the candidate area, and summing the maximum geometric consistency.

[0013] By aggregating the geometric consistency between the candidate region and the best matching regions in all other viewpoints, the cross-view consistency score characterizes the comprehensive geometric support strength obtained by the candidate region from all other viewpoints, and more comprehensively reflects the possibility of its real existence in three-dimensional space.

[0014] In a possible implementation of the first aspect, the multi-view integration module obtains the set of combined recognition scores of the candidate area for each device ID by multiplying the cross-view consistency score by the initial likelihood value in the initial recognition information.

[0015] This allows the single-view visual recognition results to be weighted by geometric consistency across views. A candidate region's combined score is high only when it visually resembles a device in a single view and has strong geometric support in other views. This effectively suppresses false alarms in a single view and enhances recognition reliability.

[0016] In one possible implementation of the first aspect, the multi-view integration module constructs an undirected graph with the candidate regions as nodes. If the geometric consistency between the cross-view candidate regions exceeds a preset threshold, edges are added between the corresponding nodes, and connected components of the undirected graph are searched for grouping. By constructing a graph based on geometric consistency and performing connected component analysis, the system can automatically group candidate regions from different viewpoints that correspond to the same physical surgical instrument instance, providing a basis for subsequent information aggregation within the group.

[0017] In one possible implementation of the first aspect, the multi-view integration module sums, for each device ID, the combined recognition scores of all candidate regions in the group for that device ID to obtain the aggregated recognition score set for each device ID in the group. Summing and aggregating the combined recognition scores of all members within a recognition group aggregates evidence for the same physical device instance from different perspectives. Even if the device is partially obscured or difficult to identify from one perspective, information from other perspectives can compensate, thereby improving the accuracy and confidence of the overall recognition of the physical device.

[0018] In the second aspect, the present application provides a surgical instrument identification method based on multiple perspectives, including: obtaining multi-perspective image data and camera parameters; preprocessing the image data to detect candidate areas from each perspective image and extract initial identification information; determining the spatial relationship between perspectives and obtaining the geometric consistency between cross-perspective candidate areas based on the spatial relationship; grouping the candidate areas based on the geometric consistency; fusing the geometric consistency with the initial identification information to obtain a combined identification score; aggregating the combined identification scores within the group to obtain an aggregated identification score; determining the surgical instrument ID and identification confidence based on the aggregated identification score; and outputting the surgical instrument ID and the identification confidence.

[0019] In a possible implementation of the second aspect, the determination step further includes maintaining a historical recognition frequency of the surgical instrument ID, and correcting the aggregate recognition score according to the historical recognition frequency before performing the determination.

[0020] In a possible implementation of the second aspect, the step of obtaining the geometric consistency includes: obtaining the geometric consistency based on the spatial distance between the source perspective candidate area position information and the target perspective candidate area position information after the source perspective candidate area position information is projected to the target perspective using the spatial relationship, and the geometric consistency is negatively correlated with the spatial distance.

[0021] In a possible implementation of the second aspect, the step of obtaining a cross-perspective consistency score based on the geometric consistency includes: for any candidate area, traversing other perspectives different from the current perspective and determining the maximum geometric consistency of other perspective candidate areas that have the highest geometric consistency with the candidate area, and summing the maximum geometric consistencies.

[0022] In a possible implementation of the second aspect, the step of fusing the geometric consistency with the initial identification information includes: multiplying the cross-view consistency score with the initial likelihood value in the initial identification information to obtain the combined identification score set of the candidate area for each device ID.

[0023] In a possible implementation of the second aspect, the step of associating and grouping the candidate areas includes: constructing an undirected graph with the candidate areas as nodes, adding edges between corresponding nodes if the geometric consistency between the cross-view candidate areas is greater than a preset threshold, and searching for connected components of the undirected graph.

[0024] In a possible implementation of the second aspect, the step of aggregating the combined recognition scores within the group includes: for each device ID, summing the combined recognition scores of all candidate areas in the group for the device ID to obtain the aggregated recognition score set of the group for each device ID. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic diagram of the structure of an identification system provided in some embodiments of the present application; Figure 2 A flowchart of an identification method provided for some embodiments of the present application. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.

[0027] Hereinafter, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature identified with "first," "second," etc., may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0028] In addition, in this application, directional terms such as "up", "down", "left", and "right" may be defined including but not limited to the orientation relative to the schematic placement of the components in the drawings. It should be understood that these directional terms may be relative concepts. They are used for relative descriptions and clarifications, and they may change accordingly according to changes in the orientation of the components in the drawings.

[0029] In this application, unless otherwise specified or limited, the term "connection" should be understood broadly. For example, "connection" can mean fixed connection, detachable connection, or integration; it can mean direct connection or indirect connection through an intermediate medium. In addition, the term "electrical connection" can refer to the method of electrical connection that enables signal transmission.

[0030] As used herein, “about,” “substantially,” or “approximately” includes the stated value and reference values ​​that are within an acceptable range of deviation from the particular value, as determined by one of ordinary skill in the art taking into account the measurements in question and the errors associated with the measurement of the particular quantity (i.e., the limitations of the measurement method).

[0031] The present invention provides a multi-view based surgical instrument recognition system and method, which aims to use data from multiple viewpoints to improve the recognition accuracy and robustness of surgical instruments by combining geometric consistency and visual features, especially in cases of partial occlusion or similar appearance.

[0032] Figure 1 FIG1 shows a system structure diagram of the surgical instrument recognition system based on multiple perspectives of the present invention. Figure 1 As shown, the system may include an acquisition module 101 , a pre-processing module 102 , a cross-perspective analysis module 103 , a multi-perspective integration module 104 , a determination module 105 and an output module 106 .

[0033] Acquisition module 101 is configured to acquire a multi-view image sequence of the surgical area from at least two perspectives, along with camera parameters corresponding to the at least two perspectives. The multi-view image sequence comprises multiple frames for real-time or near-real-time processing. The camera parameters include intrinsic and extrinsic parameters for each perspective. The intrinsic parameters describe the camera's internal optical and geometric properties, while the extrinsic parameters describe the camera's position and posture in three-dimensional space.

[0034] For example, in minimally invasive surgery, two or more stereo endoscopes or cameras mounted on a surgical robot can be used to capture color image sequences of the surgical field from different angles. Camera parameters can be pre-calibrated using standard camera calibration methods (e.g., using a checkerboard grid) or dynamically estimated during surgery. The acquisition module transmits these image frames and the corresponding camera parameters to subsequent processing modules.

[0035] The preprocessing module 102 is configured to perform a preprocessing operation on each view image of any current frame in the multi-view image sequence. The preprocessing operation includes detecting multiple candidate regions in the view image and extracting regional features and an initial device ID likelihood set for each candidate region.

[0036] The goal of detecting candidate regions is to preliminarily identify areas in an image that may contain surgical instruments. This can be achieved using existing object detection techniques. For example, a pretrained convolutional neural network (CNN) model can be used as a detector. This model, trained on a dataset of images containing various surgical instruments, can identify potential instrument instances in an image and generate bounding boxes for them. These bounding boxes are the candidate regions.

[0037] The goal of extracting region features is to obtain visual descriptions of each candidate region. Region features can include the candidate region's location in the image (such as the pixel coordinates and center point coordinates of its bounding box) as well as visual features reflecting its appearance. Visual features can be deep feature vectors extracted from the aforementioned detection model or other feature extraction networks (such as ResNet and VGG), or traditional image features (such as SIFT and HOG). The initial instrument ID likelihood set reflects the preprocessing module's preliminary judgment of the class of known surgical instruments to which each candidate region may belong. This is typically a confidence score output by the detection model, indicating the likelihood that the candidate region belongs to each class in the preset instrument ID set (such as scissors, forceps, and needle holder). The initial instrument ID likelihood set can be represented as a vector with a dimension equal to the number of preset instrument classes, where each element in the vector corresponds to the initial likelihood value for an instrument ID.

[0038] For example, the preprocessing module runs a pretrained YOLOv7 model on a single-view image in the current frame. This model outputs multiple bounding boxes and their corresponding class confidence vectors. A detected bounding box is a candidate region. The pixel coordinates of the bounding box (e.g., the coordinates of the top left and bottom right corners) constitute the location information. The feature map extracted by the model during detection is pooled to obtain a feature vector, which serves as the visual feature. The class confidence vector corresponding to the bounding box is the initial instrument ID likelihood set, for example, [0.1 (scissors), 0.8 (pliers), 0.05 (needle holder), ...].

[0039] The cross-view analysis module 103 is configured to determine the spatial transformation relationship between the at least two views and obtain geometric consistency between the cross-view candidate regions.

[0040] Determining the spatial relationship between perspectives is based on the camera parameters provided by the acquisition module. For stereo vision systems, the camera intrinsic and extrinsic parameters can be used to construct a basic matrix or an essential matrix, or the camera matrix can be used directly to represent the projection relationship from a three-dimensional space point to a two-dimensional image point. For more complex non-rigid scenes, feature matching combined with motion recovery structure or simultaneous localization and mapping techniques can be used to dynamically estimate the relative pose between perspectives. The spatial transformation relationship can be expressed as a transformation from the image coordinate system of one perspective to the image coordinate system of another perspective, such as a homography or a more general nonlinear mapping.

[0041] Acquiring geometric consistency is to assess which candidate regions from one view spatially match those in another view after being projected onto the other view. For any source view in the multiple views and any target view different from the source view, for any candidate region from the source view, the position information of the candidate region from the source view is projected onto the target view using the spatial transformation relationship between the source and target views. This position information can be simplified to the pixel coordinates of the candidate region's center point. After projection, a projected position in the target view's image coordinate system is obtained. For each candidate region in the target view, its position information (e.g., center point coordinates) is determined. Based on the spatial distance between the projected position of the candidate region from the source view and the position information of each candidate region in the target view, a pairwise geometric consistency score is obtained between the candidate region from the source view and the candidate region in the target view. The pairwise geometric consistency score represents the degree of spatial match between the candidate region from the source view and the candidate region in the target view after being projected onto the target view, and is negatively correlated with the spatial distance. A closer distance indicates a higher consistency score. The pairwise geometric consistency scores may be filled into corresponding positions of the pairwise geometric consistency matrix between the source view and the target view.

[0042] For example, consider a candidate region R1 in source view V1, whose center pixel coordinates are (u1, v1). Using the spatial transformation T(V1->V2) from V1 to target view V2, determined by the camera parameters, (u1, v1) is projected onto V2, yielding the projected point (u1', v1'). In V2, there are multiple candidate regions R2_a, R2_b, R2_c, etc., whose center coordinates are (u2_a, v2_a), (u2_b, v2_b), and (u2_c, v2_c), respectively. The Euclidean distance between the projected point (u1', v1') and the center point of each candidate region in V2 is calculated as follows: d_a = ||(u1', v1') - (u2_a, v2_a)||_2, d_b = ||(u1', v1') - (u2_b, v2_b)||_2, and so on. Pairwise geometric consistency scores can be obtained based on these distances. For example, a maximum allowed distance threshold D_max can be set. If the distance d>D_max, the consistency score is 0. If d<= D_max, the consistency score can be obtained as Score = exp(-alpha * d^2) or Score = max(0, 1 - d / D_max), where alpha is a decay factor. For example, if R1 is closest to R2_b after being projected onto V2, the distance d_b is 5 pixels, and D_max is 10 pixels, the consistency score is max(0, 1 - 5 / 10) = 0.5. These pairwise scores form a matrix, with rows representing candidate regions from the source view and columns representing candidate regions from the target view.

[0043] The multi-view integration module 104 is used to comprehensively utilize the geometric consistency between different views and the initial recognition information from a single view. Its main functions include obtaining a cross-view consistency score for each candidate region, combining the cross-view consistency score with the initial recognition information, and associating and grouping corresponding candidate regions from different views based on geometric consistency and aggregating their recognition scores.

[0044] For each of the candidate regions in the at least two perspectives, obtain its cross-perspective consistency score. The cross-perspective consistency score characterizes the degree of geometric support obtained by the candidate region from other perspectives. For any candidate region of any current perspective, traverse all other perspectives different from the current perspective. For each of the other perspectives, determine the maximum pairwise geometric consistency score of the candidate region in the other perspective that has the highest pairwise geometric consistency score with the candidate region of the current perspective. Sum the maximum pairwise geometric consistency scores of the candidate region of the current perspective and all other perspectives to obtain the cross-perspective consistency score of the candidate region of the current perspective. This is equivalent to the extent to which a candidate region can find its geometrically corresponding matching region in other perspectives.

[0045] For example, candidate region R1 is located in view V1. The other view angles are V2 and V3. The maximum value of the pairwise geometric consistency score between R1 and all candidate regions in V2 is Score(R1, V2_max), and the maximum value of the pairwise geometric consistency score between R1 and all candidate regions in V3 is Score(R1, V3_max). Then the cross-view consistency score of R1 is CrossView_Score(R1) = Score(R1, V2_max) + Score(R1, V3_max). If there are N view angles, then the N-1 maximum pairwise consistency scores are summed.

[0046] For each of the candidate areas in the at least two perspectives, the cross-perspective consistency score is combined with the initial device ID likelihood set to obtain a combined identification score set for the candidate area. The combination method may be to multiply the cross-perspective consistency score with the initial likelihood. For each device ID in the initial device ID likelihood set and its corresponding initial likelihood value, the cross-perspective consistency score of the candidate area is multiplied with the initial likelihood value corresponding to the device ID to obtain a combined identification score of the candidate area for the device ID. The combined identification scores of the candidate area for all device IDs in the initial device ID likelihood set constitute a combined identification score set for the candidate area.

[0047] For example, the cross-view consistency score of candidate region R1 is CrossView_Score(R1) = 1.2. Its initial instrument ID likelihood set is [0.1 (scissors), 0.8 (pliers), 0.05 (needle holder), ...]. The combined recognition score set of R1 is: Scissors: 0.1 * 1.2 = 0.12 Pliers: 0.8 * 1.2 = 0.96 Needle holder: 0.05 * 1.2 = 0.06, etc.

[0048] This shows that even if a candidate region looks very similar to a device in a single view (high initial likelihood), if it has no geometric correspondence in other views (low cross-view consistency), its combination score will be reduced; conversely, if it has strong geometric support in other views, its combination score will be improved due to high cross-view consistency.

[0049] A candidate region association graph is constructed based on the pairwise geometric consistency between the candidate regions in the at least two perspectives, and a connected component analysis is performed on the candidate region association graph to obtain a plurality of candidate region connected components. Each candidate region connected component represents a potential actual surgical instrument instance in the current frame. An undirected graph is constructed, and the nodes of the undirected graph are all the candidate regions in the at least two perspectives. For any source perspective in the at least two perspectives and any target perspective different from the source perspective, for any candidate region of the source perspective and any candidate region of the target perspective, if the pairwise geometric consistency score between the candidate region of the source perspective and the candidate region of the target perspective is greater than a preset threshold (for example, 0.3), an edge is added between the node corresponding to the candidate region of the source perspective and the node corresponding to the candidate region of the target perspective. All connected components are searched in the undirected graph.

[0050] For example, V1 has a candidate region R1, V2 has a candidate region R2, and V3 has a candidate region R3. If the pairwise geometric consistency score between R1 and R2 is greater than the threshold, and the pairwise geometric consistency score between R2 and R3 is greater than the threshold, then an edge is connected between nodes R1 and R2 in the graph, and an edge is connected between R2 and R3. If the consistency score between R1 and R3 is less than the threshold, there is no direct edge between them. Connected component analysis will find that R1, R2, and R3 belong to the same connected component, indicating that these three candidate regions are likely to correspond to the same actual surgical instrument. If V1 also has a candidate region R4, and its consistency score with any candidate region of V2 or V3 is less than the threshold, then R4 will constitute a connected component alone.

[0051] For the current frame, for each candidate region connected component, aggregate the combined identification score set of all candidate regions belonging to the connected component to obtain the aggregated identification score set of the candidate region connected component. Aggregation is intended to integrate the identification information of the candidate regions corresponding to the same actual instrument instance under different perspectives. For each instrument ID of the at least two preset instrument IDs, traverse all the candidate regions in the candidate region connected component. Sum the combined identification scores of all the candidate regions in the candidate region connected component for the current instrument ID to obtain the aggregated identification score of the candidate region connected component for the current instrument ID. The aggregated identification scores of the candidate region connected component for the at least two preset instrument IDs constitute the aggregated identification score set of the candidate region connected component.

[0052] For example, a connected component contains R1 from V1, R2 from V2, and R3 from V3. Assume that their combined recognition score sets are: R1: [0.12 (scissors), 0.96 (pliers), 0.06 (needle holder)] R2: [0.15 (scissors), 0.90 (pliers), 0.10 (needle holder)] R3: [0.08 (scissors), 0.95 (pliers), 0.07 (needle holder)] Then the aggregated recognition score set of the connected component is: Scissors: 0.12 + 0.15 + 0.08 = 0.35 Pliers: 0.96 + 0.90 + 0.95 = 2.81 Needle holder: 0.06 + 0.10 + 0.07 = 0.23 The aggregate score vector is [0.35, 2.81, 0.23].

[0053] The determination module 105 is configured to determine the surgical instrument ID and recognition confidence corresponding to each candidate region connected component based on the aggregated recognition score set of each candidate region connected component. For each connected component, the instrument ID with the highest score in its aggregated recognition score set is determined to be the surgical instrument ID corresponding to that connected component. The highest score can be used as the recognition confidence.

[0054] Optionally, the determination module is also used to maintain the historical recognition frequency of each surgical instrument ID that has been identified during the current operation. The historical recognition frequency can be a statistical count or weighted average of the instrument IDs identified in the past several frames or a period of time. For example, maintain a window to record the instrument IDs that were finally recognized in the past 100 frames and their number of times. Then, for each of the candidate area connected components and each instrument ID of the at least two preset instrument IDs, the aggregate recognition score of the candidate area connected component for the instrument ID is corrected according to the historical recognition frequency of the instrument ID and the preset frequency weight factor to obtain a corrected aggregate recognition score. This can introduce a context-based prior knowledge, that is, which instruments are more likely to appear in the current stage of the operation. For example, the correction formula can be used: Corrected aggregate recognition score_i = aggregate recognition score_i * (1 + frequency weight factor * historical frequency_i) Where aggregate recognition score_i is the aggregate recognition score of the connected component for device ID i, historical frequency_i is the historical recognition frequency of device ID i (for example, normalized to [0, 1]), and the frequency weight factor is an adjustable parameter used to control the influence of historical frequency. This enhances the stability of multi-view recognition results for a single frame by leveraging cross-frame temporal information, thus avoiding recognition errors in certain frames.

[0055] For example, for the above aggregated score vector [0.35 (scissors), 2.81 (pliers), 0.23 (needle holder)], the highest score is 2.81, corresponding to "pliers." Therefore, the instrument ID corresponding to this connected component is determined to be "pliers," with a confidence level of 2.81. If historical frequency correction is introduced, assuming that the frequency of "pliers" in the historical recognition frequency is higher (e.g., 0.9) and the frequency of "scissors" is lower (e.g., 0.2), and the frequency weight factor is set to 1.0, the corrected score may be: Scissors: 0.35 * (1 + 1.0 * 0.2) = 0.42 Pliers: 2.81 * (1 + 1.0 * 0.9) = 5.339 After the correction, "pliers" still has the highest score, with a confidence level of 5.339. In another example, if the initial aggregate scores for "scissors" and "pliers" are very close, but the historical frequency of "pliers" is much higher than that of "scissors," the correction will significantly improve the score for "pliers," helping to stabilize the recognition results.

[0056] Output module 106 is configured to output the surgical instrument ID identified in the current frame and its identification confidence. The output may include the instrument ID of each identified instrument instance, its corresponding confidence score, and the location of that instance in each viewport image (e.g., the bounding boxes of each candidate region in the corresponding connected component). This information can be used for visual annotation in the image or transmitted to a surgical robot for further navigation or manipulation.

[0057] This invention effectively leverages the redundancy and complementarity of multi-view data by combining geometric consistency information from multiple viewpoints with visual recognition information from a single viewpoint. As can be understood, geometric consistency information effectively filters out false detections and alarms caused by occlusion or similar appearance under a single viewpoint. It also associates candidate regions belonging to the same physical device from different viewpoints, thereby improving recognition robustness and accuracy through information aggregation.

[0058] like Figure 2 As shown, the present application also provides a surgical instrument recognition method based on multiple perspectives, including: S100, obtaining multi-view image data and camera parameters; S200, pre-processing the image data to detect candidate regions from each perspective image and extract initial recognition information; S300, determining a spatial relationship between perspectives and obtaining geometric consistency between cross-perspective candidate regions based on the spatial relationship; S400, grouping the candidate regions based on the geometric consistency; S500, fusing the geometric consistency with the initial recognition information to obtain a combined recognition score; S600, aggregating the combined recognition scores within the group to obtain an aggregate recognition score; determining the surgical instrument ID and recognition confidence based on the aggregate recognition score; and S700: Output the surgical instrument ID and the recognition confidence.

[0059] In some embodiments, the determination step further includes maintaining a historical recognition frequency of the surgical instrument ID, and correcting the aggregate recognition score according to the historical recognition frequency before making a determination.

[0060] In some embodiments, the step of obtaining the geometric consistency includes: obtaining the geometric consistency based on the spatial distance between the source perspective candidate area position information and the target perspective candidate area position information after projecting the source perspective candidate area position information to the target perspective using the spatial relationship, and the geometric consistency is negatively correlated with the spatial distance.

[0061] In some embodiments, the step of obtaining a cross-perspective consistency score based on the geometric consistency includes: for any candidate area, traversing other perspectives different from the current perspective and determining the maximum geometric consistency of other perspective candidate areas that have the highest geometric consistency with the candidate area, and summing the maximum geometric consistency.

[0062] In some embodiments, the step of fusing the geometric consistency with the initial identification information includes: multiplying the cross-view consistency score with the initial likelihood value in the initial identification information to obtain the combined identification score set of the candidate area for each device ID.

[0063] In some embodiments, the step of associating and grouping the candidate regions includes: constructing an undirected graph with the candidate regions as nodes, adding edges between corresponding nodes if the geometric consistency between the cross-view candidate regions is greater than a preset threshold, and searching for connected components of the undirected graph.

[0064] In some embodiments, the step of aggregating the combined recognition scores within the group includes: for each device ID, summing the combined recognition scores of all candidate regions in the group for the device ID to obtain the aggregated recognition score set for each device ID in the group. Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0065] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0066] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units shown in this embodiment may be selected according to actual needs to achieve the purpose of this embodiment.

[0067] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware.

[0068] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A surgical instrument recognition system based on multiple perspectives, characterized in that: include: Acquisition module, used to obtain multi-view image data and camera parameters; A preprocessing module is used to detect candidate regions from images of each viewpoint and extract initial recognition information; A cross-view analysis module, configured to determine the spatial relationship between views and obtain geometric consistency between cross-view candidate regions based on the spatial relationship; a multi-view integration module, configured to group the candidate regions based on the geometric consistency, fuse the geometric consistency with the initial recognition information to obtain a combined recognition score, and aggregate the combined recognition scores within the group to obtain an aggregated recognition score; a determination module, configured to determine the surgical instrument ID and recognition confidence according to the aggregated recognition score; as well as An output module is used to output the surgical instrument ID and the recognition confidence.

2. The system according to claim 1, wherein: The determination module is further configured to obtain a historical recognition frequency of the surgical instrument ID and to modify the aggregate recognition score according to the historical recognition frequency.

3. The system according to claim 1, wherein: The cross-view analysis module obtains the geometric consistency based on the spatial distance between the source view candidate area position information and the target view candidate area position information after the source view candidate area position information is projected to the target view using the spatial relationship.

4. The system according to claim 3, characterized in that The multi-perspective integration module obtains a cross-perspective consistency score based on the geometric consistency. For the cross-perspective consistency score of any candidate area, the maximum value of the geometric consistency between any candidate area and the candidate areas in the other perspectives is determined for each other perspective different from the current perspective, and the cross-perspective consistency score is obtained by summing up multiple maximum values.

5. The system according to claim 4, characterized in that The multi-view integration module combines the cross-view consistency score with the initial likelihood in the initial recognition information to obtain the combined recognition score set of the candidate area for each device ID.

6. The system according to claim 1, wherein: The multi-view integration module constructs an undirected graph with the candidate regions as nodes. If the geometric consistency between the cross-view candidate regions is greater than a preset threshold, edges are added between the corresponding nodes, and the connected components of the undirected graph are searched for grouping.

7. The system according to claim 1, wherein: The multi-view integration module obtains the aggregated recognition score set for each instrument ID of the group by summing the combined recognition scores of all candidate regions in the group for the instrument ID.

8. A surgical instrument recognition method based on multiple perspectives, characterized in that: include: Obtain multi-view image data and camera parameters; preprocessing the image data to detect candidate regions from each view image and extract initial recognition information; Determining the spatial relationship between views and obtaining geometric consistency between cross-view candidate regions based on the spatial relationship; Associating and grouping the candidate regions based on the geometric consistency; fusing the geometric consistency with the initial recognition information to obtain a combined recognition score; Aggregate the combined recognition scores within the group to obtain the aggregate recognition score; Determining the surgical instrument ID and recognition confidence based on the aggregated recognition score; and The surgical instrument ID and the recognition confidence level are output.

9. The method according to claim 8, characterized in that The determination step further includes obtaining a historical recognition frequency of the surgical instrument ID and correcting the aggregate recognition score according to the historical recognition frequency.

10. The method according to claim 8, characterized in that The step of obtaining the geometric consistency includes: obtaining the geometric consistency based on the spatial distance between the source perspective candidate area position information and the target perspective candidate area position information after projecting the source perspective candidate area position information to the target perspective using the spatial relationship, and the geometric consistency is negatively correlated with the spatial distance.

Citation Information

Cited By

  • Bone nail automatic identification and dosage confirmation method based on multi-feature collaborative reasoning

    CN120976584A