A method and system for tracking instruments and automatically moving the scope in the laparoscopic field of view

Through object detection, matching and tracking algorithms, combined with KCF algorithm, the problem of instrument tracking and automatic mirror operation in multi-insert scenarios in laparoscopic minimally invasive surgery is solved, and the stable tracking and reasonable display of the instrument is achieved, improving the accuracy and stability of the automatic mirror operation.

CN118750170BActive Publication Date: 2025-09-02SHANGHAI DROIDSURG MEDICAL CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410868852.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2025-09-02
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle real-time tracking and automatic mirroring of instruments in multi-instrument scenarios in laparoscopic minimally invasive surgery, especially when the instrument overlaps and obstructs the processing logic, and the depth estimation errors lead to difficulty in automatic mirroring.

Method used

Through the target detection, target matching and target tracking algorithms, combined with the kernel-related filtering KCF algorithm, an automatic mirror operation strategy is generated, and the main tracking target is specified and the mirror operation strategy of the mirror-holding robot is generated. The device tracking and mirror operation are performed based on the position and size information in the image.

Benefits of technology

The instrument tracking stability and tracking accuracy in multi-instrument scenarios are improved, the stable display and reasonable size of the front end of the instrument in the image field of view are achieved, and the accuracy and sustainability of the automatic operation mirror are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118750170B_ABST
    Figure CN118750170B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of laparoscopic surgical robots, and provides a method for tracking instruments and automatically moving the camera in a laparoscopic field of view, comprising: S1: selecting a tracking instrument as a primary tracking target on an initial video frame image, outlining positioning frame information surrounding the front end of the primary tracking target, and setting an ideal length of the front end of the primary tracking target; S2: acquiring a video stream image in the laparoscopic field of view at a fixed frequency, and acquiring the positioning frame information of the front end of the tracking instrument in the current image frame based on the acquired current frame image through target detection, target matching, target tracking, positioning frame synthesis and update, and post-processing of the tracked target; S3: generating an automatic camera movement strategy based on the position and size of the positioning frame information of the primary tracking target in the current image frame obtained in the video stream image. The method is used to solve the problem of real-time tracking of instruments in a multi-instrument scenario in the laparoscopic field of view, and to improve the accuracy, stability, and sustainability of the automatic camera movement of the camera-holding robot during surgery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of laparoscopic surgical robots, and in particular to a method and system for tracking instruments and automatically moving the lens in a laparoscopic field of view. Background Art

[0002] Minimally invasive laparoscopic surgery is performed by making a hole in the abdomen to enter the abdominal cavity, using minimally invasive surgical instruments and a laparoscopic imaging system. This surgical method has the advantages of minimal trauma, minimal pain, and a quick recovery, and has gradually become a commonly used surgical method in clinical practice.

[0003] Minimally invasive laparoscopic surgery requires visual tracking of surgical instruments. Instrument visual tracking uses image processing and computer vision technologies to automatically identify and accurately track surgical instruments. This technology typically relies on advanced image processing and computer vision algorithms. These algorithms analyze laparoscopic images in real time to identify information such as the shape, size, and position of surgical instruments. Furthermore, by processing and comparing consecutive images, they can accurately calculate the trajectory and posture of surgical instruments.

[0004] Existing laparoscopic minimally invasive surgery also utilizes an active scope-holding robot. This robot receives instructions from the surgeon through a human-machine interface, adjusting the position and posture of the laparoscope to provide the surgeon with a clear surgical field of view. The development trend of this human-machine interface is towards active guidance based on laparoscopic images, with visual tracking of surgical instruments being a core technology.

[0005] In the prior art, there are already some technical solutions for performing instrument vision on minimally invasive surgical instruments during laparoscopic minimally invasive surgery.

[0006] For example, patent application CN113538522A - A method for visual tracking of instruments for laparoscopic minimally invasive surgery discloses a solution: the laparoscopic surgery video stream is processed frame by frame, the target area of ​​the surgical instrument is segmented, and the position of the tracking point on the surgical instrument joint is determined based on the segmentation result through image processing method, the depth information of the image is obtained using a monocular vision depth estimation network, and the three-dimensional spatial coordinates of the tracking point in the world coordinate system are obtained based on the principle of camera calibration to guide the dynamic control of the robot.

[0007] However, patent application CN113538522A, a method for visually tracking instruments for minimally invasive laparoscopic surgery, still suffers from the following flaws: The method, which uses segmentation of the target area of ​​surgical instruments to achieve instrument positioning and tracking, is difficult to handle when multiple instruments are present in a video image, especially when overlapping occlusions occur at the front ends of multiple instruments, causing subsequent processing logic to fail. Furthermore, the monocular visual depth estimation network is trained through deep learning on a dataset of image pairs based on real image-depth distance maps. The model parameters are highly coupled to the parameters of the laparoscopic imaging camera. Changing the laparoscope or imaging parameters can lead to depth estimation errors, resulting in errors in the converted world coordinates of the tracking points.

[0008] For another example, patent application CN114663474A - A method for visual tracking of multiple instruments in the laparoscopic field of view of a laparoscopic robot discloses a solution: based on a workflow similar to that of CN113538522A, it improves the network for segmenting the target area of ​​surgical instruments, making the positioning of surgical instrument joints more accurate. When multiple surgical instrument joints are visible, the tracking target point is defined as the midpoint of the line connecting the two joint points farthest apart, solving the problem of tracking multiple instruments.

[0009] However, patent application CN114663474A - a method for visual tracking of multiple instruments in the laparoscopic field of view of a mirror-holding robot - still has the following defects: when handling multi-instrument scenarios, this solution defines the tracking target point as the midpoint of the joint points of the two instruments farthest apart, and is unable to specify the main instrument in the multi-instrument scenario during laparoscopic surgery. In theory, it is more reasonable to use the main instrument as the tracking target. Summary of the Invention

[0010] To address the above-mentioned issues, the present invention aims to provide a method and system for tracking and automatically moving instruments within a laparoscopic field of view, thereby solving the problem of real-time tracking of a specific instrument in a multi-instrument scenario. Furthermore, to address the difficulty in automatically moving the laparoscopic image due to the inability to effectively perceive the depth of the instrument target, the present invention uses the detected position and size of the instrument target in the image, and then forms an effective automatic movement strategy through algorithmic processing, thereby improving the accuracy, stability, and sustainability of the automatic movement of the scope-holding robot during surgery.

[0011] The above-mentioned object of the present invention is achieved through the following technical solutions:

[0012] A method for tracking instruments and automatically moving the scope in a laparoscopic field of view comprises the following steps:

[0013] S1: Acquire an initial video frame image of the laparoscope field of view for instrument tracking and automatic scope movement, select a tracking instrument as a primary tracking target on the initial video frame image, draw a positioning frame surrounding the front end of the primary tracking target, and set an ideal length of the front end of the primary tracking target;

[0014] S2: acquiring a video stream image in the laparoscope field of view at a fixed frequency, and obtaining, based on the acquired current frame image, the positioning frame information of the front end of the tracking instrument that includes at least the primary tracking target in the current image frame through target detection, target matching, target tracking, positioning frame synthesis and updating, and tracked target post-processing processing algorithms;

[0015] S3: Generate an automatic camera movement strategy based on the acquired position and size of the positioning frame information of the main tracking target in the current image frame in the video stream image.

[0016] Furthermore, in step S2, target detection is performed on the current frame image, specifically:

[0017] Based on the pre-trained target detection model, all pending positioning frame sets Y on the current frame image are obtained. Each element y in the positioning frame set Y i Contains the positioning frame information, device category information and confidence score, wherein the positioning frame information includes the coordinates (r, c) of the upper left corner of the rectangular frame serving as the positioning frame and the width w and height h of the rectangular frame;

[0018] The elements in the positioning box set Y are deduplicated, including removing highly overlapping positioning boxes and nested positioning boxes. The specific logic is as follows:

[0019] a: The high overlapping positioning frame is removed

[0020] Traversing all elements in the positioning box set Y, calculating the object detection evaluation index IoU between the positioning box information in the current element and the positioning box information of other elements, where the object detection evaluation index IoU is the ratio of the intersection area of ​​two positioning boxes to the combined area. If the object detection evaluation index IoU is greater than a preset evaluation critical threshold, deleting the element with the smaller confidence score;

[0021]

[0022] Wherein, A is the current positioning frame, and B is the other positioning frames in the positioning frame set Y;

[0023] b: The nested positioning frame is removed

[0024] Traverse all elements in the positioning frame set Y and calculate the inner_rate of the positioning frame information of the current element and other elements with the same device category information. The inner_rate is the ratio of the intersection area of ​​the two positioning frames to the area of ​​the smaller positioning frame. If the inner_rate is greater than a preset inner_rate critical threshold, delete the element with the larger area.

[0025]

[0026] Wherein, A is the current positioning frame, and B is other positioning frames in the positioning frame set Y.

[0027] Furthermore, in step S2, the target matching is performed on the current frame image, specifically:

[0028] Taking the positioning frame information of the front end of the tracking instrument as the target to be tracked, which is drawn by the user, as input, wherein the target to be tracked includes the primary tracking target drawn on the initial video frame image, and secondary tracking targets dynamically generated based on the result of the target detection for tracking other surgical instruments in the video stream image;

[0029] Obtain the deduplicated positioning frame set Y and the positioning frame set X of the primary tracking target and the secondary tracking target after calculation and settlement of the previous image frame, and solve the weighted bipartite graph maximum matching of the positioning frame set Y and the positioning frame set X. The weight matrix is ​​G(X,Y), and the matrix element G ij =IoU(x i ,y j ) is the target detection evaluation index IoU of the i-th positioning frame information in the positioning frame set X and the j-th positioning frame information in the positioning frame set Y;

[0030] A matching combination that maximizes the weighted sum is found based on the weight matrix to find the best pairing of the objects in the positioning frame set X and the positioning frame set Y, in preparation for subsequent positioning frame synthesis updates.

[0031] Furthermore, in step S2, the target tracking is performed on the current frame image, specifically:

[0032] Traverse the primary tracking target and all the secondary tracking targets, track the targets using the kernel correlation filter (KCF) algorithm, and obtain the corresponding tracking target frames z of the primary tracking target and all the secondary tracking targets in the current frame image.

[0033] Furthermore, in step S2, the positioning frame synthesis update is performed on the current frame image, specifically:

[0034] Based on the primary tracking target and all the secondary tracking targets, and according to the calculation results of the target matching and the target tracking, updating the tracking results of all current tracking targets including the primary tracking target and the secondary tracking targets in the current frame image, including:

[0035] When the positioning frame information corresponding to the current tracking target in the positioning frame set X in the positioning frame set Y is matched in the target matching, and the target detection evaluation index IoU is greater than 0:

[0036] a: The kernel correlation filter (KCF) algorithm successfully tracks the current tracking target, and the tracking target frame is z. The positioning frame information y and the tracking target frame z are synthesized as the positioning frame information updated for the current tracking target in the current frame image. The center point of the positioning frame information is the midpoint of the line connecting the center points of the positioning frame information y and the center points of the tracking target frame z. The area is the average of the areas of the positioning frame information y and the tracking target frame z. The aspect ratio remains consistent with the positioning frame information y.

[0037] b: The kernel correlation filter (KCF) algorithm fails to track the current tracking target and cannot obtain the latest tracking target frame z, then the front end of the tracking device undergoes significant deformation or displacement, and the matched positioning frame information y is used as the updated positioning frame information of the current tracking target in the current frame image;

[0038] When the positioning frame information of the current tracking target in the positioning frame set X in the positioning frame set Y is not matched in the target matching, and the target detection evaluation index IoU is equal to 0:

[0039] a: The kernel correlation filter (KCF) algorithm successfully tracks the current tracking target, the tracking target frame is z, the tracking target frame z is used as the positioning frame information updated for the current tracking target in the current frame image, and the current tracking target is marked as suspected lost;

[0040] b: The kernel correlation filter (KCF) algorithm fails to track the current tracking target, and the current tracking target is not updated in the current frame image. The result of the previous image frame is retained, and the current tracking target is marked as suspected lost.

[0041] Furthermore, in step S2, the tracked target post-processing is performed on the current frame image, specifically:

[0042] Traverse all the secondary tracking targets. When the number of times the secondary tracking target is continuously marked as suspected of being lost is greater than the first preset number threshold, delete the secondary tracking target, and the secondary tracking target will no longer be tracked in the subsequent video stream images;

[0043] If the number of times the primary tracking target is continuously marked as suspected of being lost is greater than the second preset number threshold, prompt that the primary tracking target is lost, and the user is allowed to choose to stop or restart the instrument tracking;

[0044] Extract the positioning frame information in the set Y of the positioning frames that are not successfully paired in the target matching, and initialize to generate new secondary tracking targets.

[0045] Further, in step S3, based on the position and size of the positioning frame information of the primary tracking target in the obtained current image frame in the video stream image, generate the automatic camera movement strategy, specifically:

[0046] Calculate the distance from the center point of the positioning frame information of the primary tracking target to the center point of the video stream image. If the distance is greater than the preset distance threshold, prompt the camera-holding robot to move a fixed step length along the specified direction in the laparoscopic image plane;

[0047] Calculate the diagonal length a of the positioning frame information of the primary tracking target, and compare it with the ideal length L set at the front end of the primary tracking target in the initial video frame image. The preset ratio threshold is r (0 < r < 1);

[0048] a: a / L < r, the lens of the camera-holding robot is far from the primary tracking target, and the primary tracking target appears small in the video stream image. Prompt the camera-holding robot to move forward a fixed step length in the depth direction;

[0049] b: r ≤ a / L ≤ 1 / r, the lens of the camera-holding robot is at an appropriate distance from the primary tracking target, and the primary tracking target appears reasonable in size in the video stream image. Keep the depth of the lens of the camera-holding robot unchanged;

[0050] c: a / L > 1 / r, the lens of the camera-holding robot is close to the primary tracking target, and the primary tracking target appears large in the video stream image. Prompt the camera-holding robot to move backward a fixed step length in the depth direction.

[0051] A laparoscopic vision instrument tracking and automatic camera movement system for performing the laparoscopic vision instrument tracking and automatic camera movement method as described above, including:

[0052] A tracking algorithm initialization module is used to obtain an initial video frame image of the laparoscope field of view to be used for instrument tracking and automatic scope movement, select a tracking instrument as the primary tracking target on the initial video frame image, outline positioning frame information surrounding the front end of the primary tracking target, and set the ideal length of the front end of the primary tracking target;

[0053] a positioning frame information acquisition module, configured to acquire a video stream image in the laparoscope field of view at a fixed frequency, and obtain the positioning frame information of the front end of the tracking instrument that includes at least the primary tracking target in the current image frame through target detection, target matching, target tracking, positioning frame synthesis and update, and tracked target post-processing processing algorithms based on the acquired current frame image;

[0054] The automatic camera movement strategy generation module is used to generate an automatic camera movement strategy based on the position and size of the positioning frame information of the main tracking target in the current image frame obtained in the video stream image.

[0055] A computer device includes a memory and one or more processors, wherein the memory stores computer code, and when the computer code is executed by the one or more processors, the one or more processors execute the above method.

[0056] A computer-readable storage medium stores computer code. When the computer code is executed, the above method is performed.

[0057] Compared with the prior art, the present invention has at least one of the following beneficial effects:

[0058] (1) The present invention combines the algorithmic processing logic of target detection, target matching, and target tracking to specify the tracking target of an instrument in a multi-instrument scenario and improve the stability and tracking accuracy of the selected instrument. During the implementation of the solution, the primary tracking target object is generated based on the specified tracking target, and the secondary tracking target objects are generated based on other targets in the target detection results. During the dynamic tracking process, the target detection results of the current image frame are paired with the historical positioning information of all tracked objects through multi-target matching, effectively improving the tracking stability of the algorithm when the targets are moving rapidly or are mutually occluded.

[0059] (2) When the present invention deals with the problem of converting the tracking target positioning information into the mirror movement strategy of the mirror-holding robot, because the depth information cannot be accurately obtained, the present invention does not directly calculate the real-world physical coordinates of the tracking instrument target. Based on the center point position of the rectangular frame of the front end of the instrument located in the image, the endoscope is instructed to move in the direction from the center point of the image to the center point of the front end of the instrument in the image plane, keeping the front end of the instrument in the central area of ​​the image field of view; based on the diagonal length of the positioning frame, the size of the front end of the instrument in the lens is estimated, and the mirror-holding robot is instructed to move forward or backward in the depth direction, keeping the size of the front end of the instrument in the image within a reasonable range. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 This is an overall flow chart of the method for tracking instruments and automatically moving the scope in the laparoscopic field of view of the present invention;

[0061] Figure 2 A flowchart of the present invention for obtaining the positioning frame information of the front end of the tracking device including at least the main tracking target in the current image frame;

[0062] Figure 3 A schematic diagram of a laparoscopic surgical image with a positioning frame and instrument category marked for the present invention;

[0063] Figure 4 This is a tracking diagram when the automatic camera movement strategy is implemented in the present invention;

[0064] Figure 5 The figure is a diagram showing the overall structure of the instrument tracking and automatic mirror movement system in the laparoscope field of view of the present invention. DETAILED DESCRIPTION

[0065] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0066] Those skilled in the art will appreciate that, unless otherwise stated, the singular forms "a," "an," "said," and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0067] To address the issues raised in the background technology, the present invention trains a target detection model to obtain detection frames for the front ends of all instruments in an image, and combines this with a tracking algorithm to achieve long-term, stable tracking of the selected instrument. The tracking target can be switched at any time based on the user's intent. To provide the camera-holding robot with positional information for its movement, the present invention does not directly provide precise three-dimensional world coordinates. Instead, it provides the camera displacement direction within the image plane and in the depth direction based on the real-time position and size of the instrument target in the image, thereby achieving stable and continuous automatic camera movement and tracking of instruments within the field of view.

[0068] The present invention mainly provides a method for tracking surgical instruments in the laparoscopic field of view of a mirror-holding robot and an automatic mirror movement strategy based on tracking information. The present invention does not require preoperative marking of surgical instruments. Instead, it detects multiple detection frames of the front ends of surgical instruments from the surgical video stream based on a deep learning method, and combines a multi-target matching algorithm and a target tracking algorithm to improve the stability and tracking accuracy of the selected instrument. The direction of the mirror movement displacement in the image plane is calculated based on the position of the instrument detection frame, and the forward or backward movement of the mirror in the depth direction is determined based on the size of the detection frame. The present invention realizes the automated detection and tracking of the front ends of instruments in laparoscopic surgery, and designs a practical automatic mirror movement strategy based on tracking information.

[0069] The following is described by specific examples:

[0070] First embodiment

[0071] like Figure 1 As shown, this embodiment provides a method for tracking instruments and automatically moving the scope in a laparoscopic field of view, comprising the following steps:

[0072] S1: Acquire an initial video frame image of the laparoscope field of view for instrument tracking and automatic mirror movement, select a tracking instrument as the main tracking target on the initial video frame image, outline positioning frame information surrounding the front end of the main tracking target, and set the ideal length of the front end of the main tracking target.

[0073] Specifically, in this embodiment, before initiating the instrument tracking method of the present invention, appropriate initialization is required. Specifically, the following steps are performed: 1. Selecting a primary tracking target from among multiple surgical instruments to be tracked in the initial video frame and drawing a positioning frame around the leading end of the primary instrument. In this embodiment, the positioning frame is a rectangular frame; 2. Setting the ideal length L of the leading end of the tracking instrument. This length is defined as the diagonal length of the rectangular frame within the video stream image when the distance between the target instrument and the laparoscope is optimal. After initialization, the algorithm generates the primary tracking target object, and the drawn rectangular frame indicates the size and position of the tracking target in the initial video frame.

[0074] S2: The video stream image in the laparoscope field of view is acquired at a fixed frequency, and based on the acquired current frame image, the positioning frame information of the front end of the tracking instrument including at least the main tracking target in the current image frame is obtained through target detection, target matching, target tracking, positioning frame synthesis update, and tracked target post-processing processing algorithms.

[0075] Specifically, such as Figure 2 As shown, in this embodiment, a video stream image is acquired frame by frame at a fixed frequency (e.g., 15 fps), and the current frame image and the tracking target positioning frame information obtained by processing the previous image frame are input. Then, the positioning frame information of the front end of the tracking device including at least the primary tracking target in the current image frame is acquired through the processing algorithms of target detection, target matching, target tracking, positioning frame synthesis and update, and tracked target post-processing. The details of the principles and implementation steps of each algorithm are as follows:

[0076] (1) performing target detection on the current frame image, specifically:

[0077] Based on the pre-trained target detection model, all pending positioning frame sets Y on the current frame image are obtained. Each element y in the positioning frame set Y i Contains the positioning frame information, device category information and confidence score. The positioning frame information includes the coordinates (r, c) of the upper left corner of the rectangular frame serving as the positioning frame and the width w and height h of the rectangular frame.

[0078] To ensure a sufficiently high target detection rate, a smaller confidence threshold is set. At this time, there may be multiple positioning frames on the same instrument front end. It is necessary to perform deduplication operations on the elements in the positioning frame set Y, including removing highly overlapping positioning frames and nested positioning frames. The specific logic is as follows:

[0079] a: The high overlapping positioning frame is removed

[0080] Traverse all elements in the positioning box set Y and calculate the target detection evaluation index IoU (Intersection over Union) between the positioning box information in the current element and the positioning box information of other elements. The target detection evaluation index IoU is the ratio of the intersection area of ​​two positioning boxes to the combined area. If the target detection evaluation index IoU is greater than a preset evaluation critical threshold, delete the element with the smaller confidence score;

[0081]

[0082] Wherein, A is the current positioning frame, and B is the other positioning frames in the positioning frame set Y;

[0083] b: The nested positioning frame is removed

[0084] Traverse all elements in the positioning frame set Y, and calculate the inner_rate of the positioning frame information of the current element and the positioning frame information of other elements with the same instrument category information. The inner_rate is the ratio of the intersection area of ​​the two positioning frames to the area of ​​the smaller positioning frame. If the inner_rate is greater than a preset inner overlap critical threshold, it means that the smaller positioning frame between the two more accurately locates the front end of the instrument, while the larger positioning frame additionally includes part of the instrument shaft. Delete the element with the larger area.

[0085]

[0086] Wherein, A is the current positioning frame, and B is other positioning frames in the positioning frame set Y.

[0087] In addition, the pre-trained target detection model in this embodiment is briefly described: the target detection model in this embodiment selects the yolo detection network for training before performing surgical instrument tracking, and the training data is used as follows Figure 3 The following laparoscopic surgical images are annotated with localization boxes and instrument categories. YOLO is a classic object detection algorithm that can quickly and accurately detect multiple objects in an image and simultaneously provide their locations and categories. The training data consists of laparoscopic surgical images annotated with localization boxes and instrument categories. This means that in each image, each surgical instrument is annotated with a rectangular box and its category. This annotated data is used to train the YOLO model, enabling it to learn the appearance characteristics and location distribution of surgical instruments, enabling accurate detection of surgical instruments in the images during testing. During training, the YOLO model optimizes its parameters using a backpropagation algorithm to minimize the discrepancy between detection results and annotations. After training, the model is able to detect objects in new images and output the location and category of each detected surgical instrument. Using the trained YOLO model for object detection enables real-time tracking of surgical instruments during surgery, improving surgical efficiency and safety.

[0088] In addition, you can also choose the target detection algorithms CenterNet and Faster R-CNN instead of the YOLO model to train the target detection model. CenterNet is a target detection algorithm based on target center point detection. It can not only detect the position of the target, but also accurately detect the center point of the target. CenterNet is simple and efficient, and is suitable for real-time target detection tasks. It can simultaneously predict the target position and category in a single network, so it has lower computational complexity and faster inference speed. Faster R-CNN is a classic target detection algorithm that adopts a two-stage detection framework. It generates candidate detection boxes through the region proposal network (RPN) and uses the region classification network to classify and regress the candidate boxes. Faster R-CNN has high performance in accuracy and stability, and is particularly suitable for target detection tasks in complex scenes.

[0089] (2) performing target matching on the current frame image, specifically:

[0090] The positioning frame information of the front end of the tracking instrument to be tracked, which is drawn by the user, is used as input, where the target to be tracked includes the main tracking target drawn on the initial video frame image, and the secondary tracking targets dynamically generated based on the result of the target detection for tracking other surgical instruments in the video stream image. The target matching method is used to make the tracking positioning clearer and the tracking results more stable and accurate.

[0091] Obtain the deduplicated positioning frame set Y and the positioning frame set X of the primary tracking target and the secondary tracking target after calculation and settlement of the previous image frame, and solve the weighted bipartite graph maximum matching of the positioning frame set Y and the positioning frame set X. The weight matrix is ​​G(X,Y), and the matrix element G ij= IoU(x i ,y j ) is the target detection evaluation index IoU of the i-th positioning frame information in the positioning frame set X and the j-th positioning frame information in the positioning frame set Y;

[0092] A matching combination that maximizes the weighted sum is found based on the weight matrix to find the best pairing of the objects in the positioning frame set X and the positioning frame set Y, in preparation for subsequent positioning frame synthesis updates.

[0093] In addition, for a given set of positioning boxes X and Y, and the corresponding weight matrix (i.e., IoU value matrix), a weighted bipartite graph maximum matching algorithm is used to find the best matching combination. The following are the general steps to solve this problem:

[0094] Construct a weighted bipartite graph: Construct the positioning box set X and the positioning box set Y into a weighted bipartite graph, where each positioning box in X is connected to each positioning box in Y, and the weight of the edge is the IoU value between them.

[0095] Solve the maximum matching problem: Use an optimization algorithm such as the Hungarian algorithm to solve the maximum matching of a weighted bipartite graph. This process will find the matching combination that maximizes the weighted sum, that is, the matching solution that maximizes the sum of the IoU values ​​of all matches.

[0096] Obtain matching results: Based on the maximum matching result, determine the best pairing method for the objects in the positioning box set X and the positioning box set Y. Each match corresponds to a pair of positioning boxes, and the IoU value between them indicates their similarity.

[0097] This optimal matching method can be used to find the optimal match between the targets in set X and set Y, providing accurate matching information for subsequent processing. In practice, existing optimization algorithm libraries can also be used to implement a weighted bipartite graph maximum matching algorithm, such as the linear_sum_assignment function in the scipy library.

[0098] (3) performing target tracking on the current frame image, specifically:

[0099] Traverse the primary tracking target and all the secondary tracking targets, track the targets using the Kernel Correlation Filter (KCF) algorithm, and obtain the corresponding tracking target frames z of the primary tracking target and all the secondary tracking targets in the current frame image.

[0100] The general steps of tracking the target using the kernel correlation filter KCF algorithm are:

[0101] Obtain target information: Extract the primary tracking target and all secondary tracking targets from the detected targets, including their location and appearance characteristics.

[0102] Initialize the tracker: For each tracked target, a target tracker is initialized using the kernel correlation filter (KCF) algorithm. This tracker will use the target's appearance features and the correlation between adjacent frames to track the target.

[0103] Frame-by-frame tracking: For the current frame image, the initialized tracker is used to track the primary tracking target and all secondary tracking targets. The tracker will update the position of the target in the current frame image based on the target's appearance characteristics and correlation.

[0104] Get the tracking target frame: For each tracked target, get the target frame output by the tracker, that is, the tracking target frame of the corresponding target in the current frame image.

[0105] Store tracking results: Save the tracking target frames of all primary tracking targets and all secondary tracking targets in the current frame image for subsequent analysis and application.

[0106] Alternatively, similar discriminative methods like the kernel correlation filter (KCF) algorithm, such as MOSSE and TLD, can be used. Generative methods, such as Kalman filtering and optical flow, are also commonly used. Deep learning methods for target tracking can also be used as alternatives.

[0107] (4) performing the positioning frame synthesis update on the current frame image, specifically:

[0108] Based on the primary tracking target and all the secondary tracking targets, and according to the calculation results of the target matching and the target tracking, updating the tracking results of all current tracking targets including the primary tracking target and the secondary tracking targets in the current frame image, including:

[0109] When the positioning frame information corresponding to the current tracking target in the positioning frame set X in the positioning frame set Y is matched in the target matching, and the target detection evaluation index IoU is greater than 0:

[0110] a: The kernel correlation filter (KCF) algorithm successfully tracks the current tracking target, and the tracking target frame is z. The positioning frame information y and the tracking target frame z are synthesized as the positioning frame information updated for the current tracking target in the current frame image. The center point of the positioning frame information is the midpoint of the line connecting the center points of the positioning frame information y and the center points of the tracking target frame z. The area is the average of the areas of the positioning frame information y and the tracking target frame z. The aspect ratio remains consistent with the positioning frame information y.

[0111] b: The kernel correlation filter (KCF) algorithm fails to track the current tracking target and cannot obtain the latest tracking target frame z, then the front end of the tracking device undergoes significant deformation or displacement, and the matched positioning frame information y is used as the updated positioning frame information of the current tracking target in the current frame image;

[0112] When the positioning frame information of the current tracking target in the positioning frame set X in the positioning frame set Y is not matched in the target matching, and the target detection evaluation index IoU is equal to 0:

[0113] a: The kernel correlation filter (KCF) algorithm successfully tracks the current tracking target, the tracking target frame is z, the tracking target frame z is used as the positioning frame information updated for the current tracking target in the current frame image, and the current tracking target is marked as suspected lost;

[0114] b: The kernel correlation filter (KCF) algorithm fails to track the current tracking target, and the current tracking target is not updated in the current frame image. The result of the previous image frame is retained, and the current tracking target is marked as suspected lost.

[0115] (5) performing post-processing of the tracked target on the current frame image, specifically:

[0116] Traversing all the secondary tracking targets, when the number of times the secondary tracking target is continuously marked as suspected lost is greater than a first preset number threshold, deleting the secondary tracking target, and no longer tracking the secondary tracking target in subsequent video stream images;

[0117] If the number of times the primary tracking target is continuously marked as suspected lost exceeds a second preset number threshold, a prompt is given that the primary tracking target is lost, and the user can choose to stop or restart device tracking;

[0118] The positioning frame information in the positioning frame set Y that is not successfully matched in the target matching is extracted, and the new secondary tracking target is initialized and generated.

[0119] It should be noted that, in this embodiment, it is necessary to sample the image frames of the laparoscopic video stream at a fixed frequency, and execute step S2 frame by frame. The specific flow chart is shown in FIG. Figure 2 shown.

[0120] S3: Based on the acquired position and size of the positioning frame information of the main tracking target in the current image frame in the video stream image, an automatic camera movement strategy is generated, specifically:

[0121] Calculate the distance between the center point of the positioning frame information of the main tracking target and the center point of the video stream image. If the distance is greater than the preset distance threshold, prompt the scope holding robot to move a fixed step length in the specified direction on the laparoscope image plane (i.e., the plane perpendicular to the optical axis), such as Figure 4 As shown, the length of the dotted arrow line is the measured distance, and the direction indicated by the arrow is the specified direction.

[0122] Calculate the diagonal length a of the positioning frame information of the main tracking target ( Figure 4 The diagonal line in the rectangular box is drawn), and compared with the ideal length L of the front end of the main tracking target set on the initial video frame image, the preset ratio threshold is r(0 <r<1);

[0123] a: a / L < r, where the lens of the endoscope-holding robot is far from the main tracking target, and the main tracking target appears small in the video stream image, indicating that the endoscope-holding robot advances a fixed step length in the depth direction;

[0124] b: r ≤ a / L ≤ 1 / r, where the distance between the lens of the endoscope-holding robot and the main tracking target is appropriate, and the main tracking target appears reasonable in size in the video stream image, keeping the depth of the lens of the endoscope-holding robot unchanged;

[0125] c: a / L > 1 / r, where the lens of the endoscope-holding robot is close to the main tracking target, and the main tracking target appears large in the video stream image, indicating that the endoscope-holding robot retreats a fixed step length in the depth direction.

[0126] Second Embodiment

[0127] As Figure 5 shown, this embodiment provides a system for instrument tracking and automatic endoscope movement in a laparoscopic field of view for performing the instrument tracking and automatic endoscope movement method in the laparoscopic field of view as in the first embodiment, including:

[0128] Tracking algorithm initialization module 1, which is used to obtain the initial video frame image of the laparoscopic field of view for instrument tracking and automatic endoscope movement, select the tracking instrument as the main tracking target on the initial video frame image, draw the positioning frame information surrounding the front end of the main tracking target, and set the ideal length of the front end of the main tracking target;

[0129] Positioning frame information acquisition module 2, which is used to obtain the video stream image in the laparoscopic field of view at a fixed frequency, and obtain the positioning frame information of at least the front end of the tracking instrument including the main tracking target in the current image frame based on the obtained current frame image through the processing algorithms of target detection, target matching, target tracking, positioning frame synthesis update, and post-processing of the tracked target in sequence;

[0130] Automatic endoscope movement strategy generation module 3, which is used to generate an automatic endoscope movement strategy based on the position and size of the positioning frame information of the main tracking target in the obtained current image frame in the video stream image.

[0131] A computer-readable storage medium stores computer code. When the computer code is executed, the above-described method is performed. A person skilled in the art will appreciate that all or part of the steps in the various methods of the above-described embodiments can be performed by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0132] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

[0133] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0134] It should be noted that the above embodiments can be freely combined as needed. The above description is only a preferred embodiment of the present invention. It should be pointed out that those skilled in the art can make several improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for tracking instruments and automatically moving the scope in a laparoscopic field of view, characterized in that: The following steps are involved: S1: Acquire an initial video frame image of the laparoscope field of view for instrument tracking and automatic scope movement, select a tracking instrument as a primary tracking target on the initial video frame image, draw a positioning frame surrounding the front end of the primary tracking target, and set an ideal length of the front end of the primary tracking target; S2: acquiring a video stream image in the laparoscope field of view at a fixed frequency, and obtaining, based on the acquired current frame image, the positioning frame information of the front end of the tracking instrument that includes at least the primary tracking target in the current image frame through target detection, target matching, target tracking, positioning frame synthesis and updating, and tracked target post-processing algorithms; S3: generating an automatic camera movement strategy based on the acquired position and size of the positioning frame information of the main tracking target in the current image frame in the video stream image; In step S2, target detection is performed on the current frame image, specifically: Based on the pre-trained target detection model, all pending positioning frame sets Y on the current frame image are obtained. Each element in the positioning frame set Y Contains the positioning frame information, device category information and confidence score, wherein the positioning frame information includes the coordinates (r, c) of the upper left corner of the rectangular frame serving as the positioning frame and the width w and height h of the rectangular frame; The elements in the positioning box set Y are deduplicated, including removing highly overlapping positioning boxes and nested positioning boxes. The specific logic is as follows: a: The high overlapping positioning frame is removed Traversing all elements in the positioning box set Y, calculating the object detection evaluation index IoU between the positioning box information in the current element and the positioning box information of other elements, where the object detection evaluation index IoU is the ratio of the intersection area of ​​two positioning boxes to the combined area. If the object detection evaluation index IoU is greater than a preset evaluation critical threshold, deleting the element with the smaller confidence score; Wherein, A is the current positioning frame, and B is the other positioning frames in the positioning frame set Y; b: The nested positioning frame is removed Traverse all elements in the positioning frame set Y and calculate the inner_rate of the positioning frame information of the current element and other elements with the same device category information. The inner_rate is the ratio of the intersection area of ​​the two positioning frames to the area of ​​the smaller positioning frame. If the inner_rate is greater than a preset inner_rate critical threshold, delete the element with the larger area. Wherein, A is the current positioning frame, and B is the other positioning frames in the positioning frame set Y; In step S2, target matching is performed on the current frame image, specifically: Taking the positioning frame information of the front end of the tracking instrument as the target to be tracked, which is drawn by the user, as input, wherein the target to be tracked includes the primary tracking target drawn on the initial video frame image, and secondary tracking targets dynamically generated based on the result of the target detection for tracking other surgical instruments in the video stream image; Get the deduplicated positioning frame set Y and the positioning frame set X of the primary tracking target and the secondary tracking target after calculation and settlement of the previous image frame, and solve the weighted bipartite graph maximum matching of the positioning frame set Y and the positioning frame set X. The weight matrix is , the matrix elements The target detection evaluation index IoU of the i-th positioning frame information in the positioning frame set X and the j-th positioning frame information in the positioning frame set Y; A matching combination that maximizes the weighted sum is found based on the weight matrix to find the best pairing of the objects in the positioning frame set X and the positioning frame set Y, in preparation for subsequent positioning frame synthesis updates.

2. The method for tracking instruments and automatically moving the laparoscopic field of view according to claim 1, characterized in that: In step S2, the target tracking is performed on the current frame image, specifically: Traverse the primary tracking target and all the secondary tracking targets, track the targets using the kernel correlation filter (KCF) algorithm, and obtain the corresponding tracking target frames z of the primary tracking target and all the secondary tracking targets in the current frame image.

3. The method for tracking instruments and automatically moving the laparoscopic field of view according to claim 2, characterized in that: In step S2, the positioning frame synthesis update is performed on the current frame image, specifically: Based on the primary tracking target and all the secondary tracking targets, and according to the calculation results of the target matching and the target tracking, updating the tracking results of all current tracking targets including the primary tracking target and the secondary tracking targets in the current frame image, including: When the positioning frame information corresponding to the current tracking target in the positioning frame set X in the positioning frame set Y is matched in the target matching, and the target detection evaluation index IoU is greater than 0: a: The kernel correlation filter (KCF) algorithm successfully tracks the current tracking target, and the tracking target frame is z. The positioning frame information y and the tracking target frame z are synthesized and used as the positioning frame information updated for the current tracking target in the current frame image. The center point of the positioning frame information is the midpoint of the line connecting the center points of the positioning frame information y and the center points of the tracking target frame z. The area is the average of the areas of the positioning frame information y and the tracking target frame z. The aspect ratio remains consistent with the positioning frame information y. b: The kernel correlation filter (KCF) algorithm fails to track the current tracking target and cannot obtain the latest tracking target frame z, then the front end of the tracking device undergoes significant deformation or displacement, and the matched positioning frame information y is used as the updated positioning frame information of the current tracking target in the current frame image; When the positioning frame information of the current tracking target in the positioning frame set X in the positioning frame set Y is not matched in the target matching, and the target detection evaluation index IoU is equal to 0: a: The kernel correlation filter (KCF) algorithm successfully tracks the current tracking target, the tracking target frame is z, the tracking target frame z is used as the positioning frame information updated for the current tracking target in the current frame image, and the current tracking target is marked as suspected lost; b: The kernel correlation filter (KCF) algorithm fails to track the current tracking target, and the current tracking target is not updated in the current frame image. The result of the previous image frame is retained, and the current tracking target is marked as suspected lost.

4. The method for tracking instruments and automatically moving the laparoscopic field of view according to claim 3, characterized in that: In step S2, the tracked target is post-processed on the current frame image, specifically: Traversing all the secondary tracking targets, when the number of times the secondary tracking target is continuously marked as suspected lost is greater than a first preset number threshold, deleting the secondary tracking target, and no longer tracking the secondary tracking target in subsequent video stream images; If the number of times the primary tracking target is continuously marked as suspected lost exceeds a second preset number threshold, a prompt is given that the primary tracking target is lost, and the user can choose to stop or restart device tracking; Extract the bounding box information of the unpaired bounding box set Y in the target match, and initialize to generate a new secondary tracking target.

5. The method for tracking instruments and automatically moving the laparoscopic field of view according to claim 1, characterized in that: In step S3, based on the position and size of the bounding box information of the primary tracking target in the obtained current image frame in the video stream image, generate the automatic camera movement strategy. Specifically: Calculate the distance from the center point of the bounding box information of the primary tracking target to the center point of the video stream image. If the distance is greater than the preset distance threshold, prompt the camera-holding robot to move a fixed step length along the specified direction in the laparoscopic image plane. Calculate the diagonal length a of the bounding box information of the primary tracking target, and compare it with the ideal length L set at the front end of the primary tracking target in the initial video frame image. The preset ratio threshold is r (0 < r < 1). a: a / L < r, the lens of the camera-holding robot is far from the primary tracking target, and the primary tracking target appears small in the video stream image. Prompt the camera-holding robot to move forward a fixed step length in the depth direction. b: r ≤ a / L ≤ 1 / r, the lens of the camera-holding robot is at an appropriate distance from the primary tracking target, and the primary tracking target appears reasonable in size in the video stream image. Keep the depth of the lens of the camera-holding robot unchanged. c: a / L > 1 / r, the lens of the camera-holding robot is close to the primary tracking target, and the primary tracking target appears large in the video stream image. Prompt the camera-holding robot to move backward a fixed step length in the depth direction.

6. A laparoscopic instrument tracking and automatic mirror movement system for executing the laparoscopic instrument tracking and automatic mirror movement method according to any one of claims 1 to 5, characterized in that: It includes: A tracking algorithm initialization module, which is used to obtain the initial video frame image of the laparoscopic field of view for instrument tracking and automatic camera movement, select a tracking instrument as the primary tracking target on the initial video frame image, draw the bounding box information surrounding the front end of the primary tracking target, and set the ideal length of the front end of the primary tracking target. A bounding box information acquisition module, which is used to obtain the video stream image in the laparoscopic field of view at a fixed frequency, and based on the obtained current frame image, obtain the bounding box information of at least the front end of the tracking instrument of the primary tracking target in the current image frame through processing algorithms such as object detection, object matching, object tracking, bounding box synthesis update, and post-processing of the tracked target. An automatic camera movement strategy generation module, which is used to generate an automatic camera movement strategy based on the position and size of the bounding box information of the primary tracking target in the obtained current image frame in the video stream image.

7. A computer device, including a memory and one or more processors, wherein computer code is stored in the memory, and when the computer code is executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, wherein the computer-readable storage medium stores computer code, and when the computer code is executed, the method according to any one of claims 1 to 5 is executed.

Citation Information

Patent Citations

  • Instrument visual tracking method for laparoscopic minimally invasive surgery

    CN113538522A

  • Intelligent automatic moving laparoscope and precision adjustment system based on image feature recognition and tracking technology

    CN109171957A

  • Man-machine collaborative minimally invasive endoscope holding robot system

    CN113143461A