Tracking-based laparoscopic surgery operation area retrospective reasoning method and system
By processing surgical videos, using point tracking and deformation field calculations, binary map labels for the operating area of the surgical instrument are generated, which solves the problems of high automatic identification of surgical operation areas and strong ambiguity in the prior art, and achieves efficient and accurate identification and analysis of surgical areas.
Patent Information
- Application Number
- CN202510163493.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art automatic identification method of surgical operation areas in minimally invasive abdominal surgery is costly and has strong label ambiguity, making it difficult to effectively train deep learning models.
By processing the surgical video, the surgical operation time screen is automatically identified, point tracking and deformation field calculation are used to obtain the regional distribution of the surgical instrument tip, and combined with Gaussian fuzzy and conditional random field processing, binary map labels are generated.
It reduces the production cost of the surgical operation area data set, improves the accuracy and boundary quality of labels, enhances the dynamicity and continuity of surgical video analysis, and provides strong support for surgical teaching and evaluation.
Smart Images

Figure CN120259932A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of laparoscopic data processing, and more specifically, relates to a retrospective inference method and system for laparoscopic surgical operation areas based on tracking. Background Art
[0002] Minimally invasive abdominal surgery has the advantages of less trauma to patients and short patient recovery time, and has been widely popularized in surgeries. The automatic identification of the surgical operation area is of great significance for minimally invasive abdominal surgery and can be widely used in downstream tasks such as autonomous camera control systems, automated surgery, and surgical skill assessment. Most current research on the automatic identification of the surgical operation area is based on deep learning methods, which train the feature extraction and segmentation capabilities of the model through a large amount of labeled data. However, due to the complex and variable nature of medical anatomical structures, it is very difficult to create labels for the surgical operation area. It requires experienced surgeons to perform long-term annotation, and at the same time, the obtained labels are ambiguous, which is not conducive to the training of deep learning models.
[0003] To solve the above problems, in the paper "Using Semantic Segmentation to Identify Surgical Anatomy During Laparoscopic Cholecystectomy", multiple annotators were invited to annotate the surgical operation area, and then the intersection of the annotations of multiple annotators was taken as the final label. This method has a high cost and ignores the differences between different annotators, and does not further process the ambiguous labels. In the paper "Automated identification of critical structures in laparoscopic cholecystectomy", label relaxation processing was performed on the basis of taking the intersection of the annotators, and certain label ambiguity information was retained. However, this method still requires a high annotation cost, and label relaxation processing may not be able to effectively represent the surgical operation area. Summary of the Invention
[0004] Aiming at the above defects or improvement requirements of the prior art, the present invention provides a retrospective inference method and system for laparoscopic surgical operation areas based on tracking. By processing the surgical video, it automatically identifies the images at the surgical operation moment. This method can be applied to different types of surgical scenarios and has strong generalization ability. By tracking the tip points of the instruments in the surgical operation images, the labels of the surgical operation area are obtained, reducing the production cost of the surgical operation area dataset and providing data for the training of deep learning models. The methods of domain density and conditional random field are used to improve the boundary quality of the labels of the surgical operation area and enhance the accuracy of the labels in representing the surgical operation area.
[0005] To achieve the above object, according to one aspect of the present invention, a retrospective inference method for laparoscopic surgical operation areas based on tracking is proposed, including the following steps:
[0006] S100: Extract the surgical instrument operation images based on the surgical flow video and determine the query frame;
[0007] S200: Perform point tracking processing on the entire surgical instrument operation and calculate the deformation field from other frames to the query frame based on the results of the point tracking processing;
[0008] S300: Based on the deformation field, obtain the area distribution of the surgical instrument tip on the query frame and map this area distribution to other frames to obtain the operation area distribution of other frames;
[0009] S400: Fit the operation area distribution and convert it into a binary map label.
[0010] As a further preference, step one includes: Based on the surgical video, screen out the pictures of the surgical operation images, segment the surgical instruments in the pictures to obtain the segmentation results of the surgical instruments, fit the circumscribed rectangle pictures of the surgical instruments in continuously specified frames, and on this basis, make the surgical operation area labels. Each continuous segment needs to be processed separately, and the middle frame of each segment is selected as the query frame.
[0011] As a further preference, in step one, fit the minimum circumscribed rectangle pictures of the surgical instruments in continuously specified frames, and at the same time obtain the image coordinate positions of the surgical instrument tips as the positions of the surgical operation areas.
[0012] As a further preference, the obtaining the image coordinate positions of the surgical instrument tips as the positions of the surgical operation areas includes:
[0013] S101: Obtain the position P of the center point of the circumscribed rectangle of the surgical instrument contour box , the width W of the rectangle box , the height H box and the rotation angle θ, judge the long axis direction of the rectangle from the numerical sizes of the width W box , the height H box , and calculate the direction vector of the long axis through the rectangle rotation angle:
[0014]
[0015] In the formula, is the direction vector of the long axis;
[0016] S102: Obtain the offset of the surgical instrument contour relative to the center point of the circumscribed rectangle frame:
[0017] offset = Contours - P box
[0018] Wherein, Contours is the contour of the surgical instrument, and offset is the offset;
[0019] S103: Calculate the projection length of each point on the contour of the surgical instrument on the long axis according to the inner product of the direction vector and the offset. The point with the maximum or minimum projection length is the tip point of the surgical instrument, which is marked as P max and P min ;
[0020] S104: According to the distance matrix D from non-edge pixels to the nearest edge pixel in the picture of the surgical operation screen, query the minimum distances from P max and P min to the edge pixels. The point with the relatively larger distance is used as the tip point P tip of the instrument.
[0021] As a further preference, the calculation formula of the tip point P tip includes:
[0022]
[0023] Wherein, D is the distance matrix from non-edge pixels to the nearest edge pixel.
[0024] As a further preference, step two includes the following steps:
[0025] S201: Divide the segment one query frame of the screened surgical operation screen into two videos before and after, and both of the two videos need to include the query frame;
[0026] S202: Reverse-play the previous video so that the query frame is the starting frame of the previous video;
[0027] S203: Input the two videos before and after processed by step S202 into the point tracking model respectively, and obtain the positions of the grid points in the query frame on other frames to achieve the point tracking result of the grid points in the query frame;
[0028] S204: Obtain the deformation field from the query frame to each frame picture according to the point tracking result.
[0029] As a further preference, in step S203, the size of the grid points in the query frame during the point tracking process is H×W, indicating that there are H rows and W columns of grid points evenly distributed on the query frame. The coordinates of each grid point p are expressed as (x,y), and the tracking result of the point p on other frames is p tracking The point p is transformed to the point p' on other frames through the deformation field, and its calculation formula is as follows:
[0030] p' = p + f(p)
[0031] In the formula, f() is the deformation field to be obtained, and this deformation field represents the horizontal and vertical displacements of the grid points:
[0032] f(p) = (f x (p), f y (p))
[0033] In the formula, f x (p) is the horizontal displacement of the grid point, and f y (p) is the vertical displacement of the grid point.
[0034] As a further preference, in step S204, the following loss function is used to solve the deformation field, so that the deformation generated by the deformation field can align the prediction result:
[0035]
[0036] In the formula, N is the total number of grid points, that is, N = H×W, is the tracking loss, p i,tracking is the tracking result of the grid point p i on other frames, and p i is the position after being transformed by the current deformation field;
[0037] Through the iterative optimization of the deformation field by the above loss function, the deformation field from the query frame to other frames of the segment is obtained.
[0038] As a further preference, step S400 includes the following steps:
[0039] S401: Based on the neighborhood density, eliminate the sparse points in the distribution of the operation area;
[0040] S402: Use Gaussian blur operation to smooth the point data of the operation area distribution, and perform normalization processing after processing the point data into a heat map;
[0041] S403: Input the heat map processed in step S402 and the original image of the surgical scene in the surgical flow video into the conditional random field, use the color and spatial information of the original image to enhance the edge quality of the heat map, and then set the division threshold of the foreground and background to convert the heat map into a binary map label.
[0042] According to another aspect of the present invention, there is also provided a retrospective inference system for laparoscopic surgical operation areas based on tracking, including:
[0043] The first main control module is used to extract the operation screen of the surgical instrument based on the surgical flow video and determine the query frame;
[0044] The second main control module is used to perform point tracking processing on the entire surgical instrument operation and calculate the deformation field from other frames to the query frame based on the results of the point tracking processing;
[0045] The third main control module is used to obtain the regional distribution of the surgical instrument tip on the query frame based on the deformation field and map the regional distribution to other frames to obtain the operation regional distribution of other frames;
[0046] The fourth main control module is used to fit the operation regional distribution and convert it into a binary map label.
[0047] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention mainly has the following technical advantages:
[0048] 1. The present invention can improve the accuracy and efficiency of surgical operation area recognition. Through point tracking and deformation field calculation, the operation area of the surgical instrument tip can be accurately mapped from the query frame to other frames, and combined with processing methods such as Gaussian blur and conditional random field, the boundary recognition of the operation area is further optimized. This inference method based on continuous frames and deformation fields can capture the dynamic operation area of the surgical instrument more accurately compared with traditional static analysis methods. This method reduces the workload of manual annotation and analysis by automatically processing surgical videos. At the same time, by using point tracking and deformation field calculation, the inference and annotation of the entire surgical operation area can be completed in a short time, improving the efficiency of surgical video analysis.
[0049] 2. The present invention can enhance the dynamicity and continuity of surgical video analysis. This method can dynamically track and infer the operation area of the surgical instrument based on the surgical flow video. By dividing the surgical video into two front and back segments and performing reverse playback processing, the operation area of the surgical instrument at different time points can be analyzed more comprehensively, thus providing strong support for the dynamic analysis of the surgical process. By fitting the circumscribed rectangle of the surgical instrument in continuously specified frames and dynamically tracking the tip points, the continuity of the surgical operation area is ensured. This continuity is very important for understanding and analyzing complex operations during the surgical process and can help doctors better review and evaluate the details of surgical operations.
[0050] 3. The present invention can provide strong support for surgical teaching and evaluation. This method can generate detailed binary map labels of the surgical operation area, which can be used as an important auxiliary tool for surgical teaching. Through these labels, students can more intuitively understand the operation range and movement trajectory of the surgical instrument, thereby improving the learning effect of surgical skills. This method can provide an objective basis for surgical evaluation. By accurately identifying and analyzing the surgical operation area, the standardization and safety of surgical operations can be evaluated, helping doctors summarize experience, improve techniques, and provide quantitative indicators for the evaluation of surgical quality.
[0051] 4. The present invention processes surgical videos to automatically identify the images at the moments of surgical operations. This method can be applied to different types of surgical scenarios and has strong generalization ability. By tracking the tip points of the instruments in the surgical operation images, the labels of the surgical operation areas are obtained, reducing the production cost of the surgical operation area dataset and providing data for the training of deep learning models. The method using domain density and conditional random field improves the boundary quality of the labels of the surgical operation areas and enhances the accuracy of the labels in representing the surgical operation areas. Description of the Drawings
[0052] Figure 1 FIG. is a flowchart of a retrospective inference method for laparoscopic surgical operation areas based on tracking according to an embodiment of the present invention;
[0053] Figure 2 FIG. is a schematic diagram for determining the circumscribed rectangle of the instrument and the tip of the instrument in an embodiment of the present invention;
[0054] Figure 3 FIG. is a flowchart for obtaining a deformation field according to an embodiment of the present invention;
[0055] Figure 4 FIG. is a schematic diagram of projecting the tip of the instrument onto a query frame according to an embodiment of the present invention;
[0056] Figure 5 FIG. is a flowchart for processing points in the operation area to binary labels according to an embodiment of the present invention. Detailed Embodiments
[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0058] Embodiment 1
[0059] As Figure 1 shown, a retrospective inference method for laparoscopic surgical operation areas based on tracking provided by an embodiment of the present invention includes the following steps:
[0060] S100: Extract the surgical instrument operation images based on the surgical flow video and determine the query frame;
[0061] S200: Perform point tracking processing on the entire surgical instrument operation and calculate the deformation field from other frames to the query frame based on the results of the point tracking processing;
[0062] S300: Based on the deformation field, obtain the regional distribution of the surgical instrument tip in the query frame, and map this regional distribution to other frames to obtain the operation area distribution of other frames;
[0063] S400: Fit the operation area distribution and convert it into a binary map label.
[0064] Specifically, in step S100, based on the surgical video, filter out the pictures of the surgical operation scenes, segment the surgical instruments in the pictures to obtain the segmentation results of the surgical instruments, fit the minimum circumscribed rectangle pictures of the surgical instruments in consecutive specified frames. On this basis, make the surgical operation area labels. Each consecutive segment needs to be processed separately, and the middle frame of each segment is selected as the query frame. Among them, fit the minimum circumscribed rectangle pictures of the surgical instruments in consecutive specified frames, and at the same time obtain the image coordinate position of the surgical instrument tip as the position of the surgical operation area.
[0065] More specifically, after obtaining the surgical video, it is necessary to first filter out the pictures of the surgical operations before proceeding with the subsequent work. First, send all the pictures into the surgical instrument segmentation model to obtain the segmentation results of the surgical instruments, and fit their minimum circumscribed rectangles. As Figure 2 shown, retain the pictures of the segments with the length of the major axis of the minimum circumscribed rectangle of the surgical instrument contour in the surgical process being greater than 1 / 10 of the length of the major axis of the surgical image and lasting for more than 200 consecutive frames. On this basis, make the surgical operation area labels. Each consecutive segment needs to be processed separately, and the middle frame of each segment is selected as the query frame to provide a reference for the subsequent calculation of the deformation field.
[0066] When fitting the minimum circumscribed rectangle, it is necessary to obtain the image coordinate position of the surgical instrument tip as the position of the surgical operation area. Specifically as follows:
[0067] First, obtain the position P of the center point of the circumscribed rectangle of the surgical instrument contour box , the width W box , H box of the rectangle, and the rotation angle θ. Judge the major axis direction of the rectangle from the numerical values of the width and height, and calculate the direction vector of the major axis through the rectangle rotation angle:
[0068]
[0069] Then, obtain the offset offset of the surgical instrument contour Contours relative to the center point of the circumscribed rectangle frame:
[0070] offset = Contours - P box (0.2)
[0071] Calculate the projection length of each point on the contour on the major axis, that is, the inner product innerProb of the contour offset and the direction vector:
[0072]
[0073] The point with the largest or smallest inner product value is the tip point of the surgical instrument. First, take out these two points and mark them as P max and P min :
[0074]
[0075] Obtain the distance matrix D from non-edge pixels to the nearest edge pixel through the edge binary map of the surgical scene picture. The value on this matrix represents the Euclidean distance from the pixel at that location to the nearest edge pixel. Then, query P max and P min The minimum distances from the two points to the edge pixels, and take the point with the relatively larger distance as the tip point P of the instrument tip , and the calculation formula is:
[0076]
[0077] Step S200 includes the following steps:
[0078] S201: Divide the segment of the surgical operation screen filtered out into two videos before and after with the query frame as the boundary. Among them, both segments of the video need to contain the query frame;
[0079] S202: Reverse-play the previous video segment so that the query frame is the starting frame of the previous video segment;
[0080] S203: Input the two videos before and after processed in step S202 into the point tracking model respectively to obtain the positions of the grid points in the query frame on other frames, so as to achieve the point tracking result of the grid points in the query frame;
[0081] S204: Obtain the deformation field from the query frame to each frame picture according to the point tracking result.
[0082] Specifically, to obtain the position information of the operation area of each frame in the continuous segment in the query frame, it is necessary to obtain the deformation field from each frame to the reference frame. Among them, the acquisition process of the deformation field is as Figure 3 shown. The point tracking code usually initializes the point grid with the first frame. Therefore, it is necessary to first divide the filtered surgical segment into two videos with the query frame as the boundary. The previous video segment needs to be reverse-played to ensure that the query frame is the first frame, and the latter video segment does not need additional processing. Input the two videos before and after each segment into the point tracking network respectively to obtain the positions of the grid points in the query frame on other frames, and then calculate the deformation field based on the change in the positions of the grid points between two frames.
[0083] During the point tracking process, the size of the query frame grid points is H×W, indicating that there are H rows and W columns of grid points evenly distributed on the query frame. The coordinates of each grid point p can be expressed as (x, y), and the tracking result of point p on other frames is p tracking , and the point p is transformed to the point p' on other frames through the deformation field. The calculation formula is as follows:
[0084] p' = p + f(p) (0.6)
[0085] Among them, f() is the deformation field to be obtained. Since this method is for two-dimensional images, the deformation field is also two-dimensional, which represents the horizontal and vertical displacements of the grid points:
[0086] f(p) = (f x (p), f y (p)) (0.7)
[0087] The solution of the deformation field is an optimization problem. The optimization goal is to align the grid points in the target frame with the prediction results of the deformation field on the query frame. For this purpose, the following loss function is used in this embodiment to solve the deformation field.
[0088] Tracking loss It is used to measure the distance between the grid points in the target frame and the predicted grid points. By minimizing this loss, the deformation field is optimized so that the deformation generated by the deformation field can align the prediction results. The loss function is expressed as:
[0089]
[0090] In the formula, N is the total number of grid points, that is, N = H×W, is the tracking loss, p i,tracking is the grid point p i is the tracking result of the grid point p i at the position transformed by the current deformation field.
[0091] Through the iterative optimization of the deformation field by the above loss function, the deformation field from the query frame to other frames of the segment is obtained.
[0092] Step S300 specifically includes: based on the deformation field, obtaining the area distribution of the tip of the surgical instrument on the query frame, and mapping this area distribution to other frames to obtain the operation area distribution of other frames. More specifically, after obtaining the deformation field from the query frame to other frames, it is necessary to project the instrument tip points in other frames of the segment onto the query frame to obtain the operation area distribution on the query frame. As Figure 4 shown, after obtaining the operation area distribution points of the query frame, and then according to the deformation field matrix, map this distribution to other frames to provide the operation area distribution for all pictures.
[0093] Step S400 includes the following steps:
[0094] S401: Based on the neighborhood density, eliminate the sparse points in the operation area distribution;
[0095] S402: Use Gaussian blur operation to smooth the point data of the operation area distribution, process the point data into a heat map and then perform normalization processing;
[0096] S403: Input the heat map processed in step S402 and the original image of the surgical scene in the surgical flow video into the conditional random field, use the color and spatial information of the original image to enhance the edge quality of the heat map, and then set the division threshold between the foreground and the background to convert the heat map into a binary map label.
[0097] Specifically, the obtained operation area distribution information is a series of two-dimensional coordinate points, which need to be further converted into a binary map label for semantic segmentation before being used for network training. The processing flow is as Figure 5 shown. Initially filter the operation area points through the domain density; use Gaussian blur operation to process the operation area points into a heat map, and normalize the heat map result to the interval from 0 to 1; send the heat map and the original image of the surgical scene into the conditional random field together, use the color and spatial information of the original image to improve the boundary of the heat map and remove the noise area at the same time; set the division threshold to convert the heat map into a binary map label.
[0098] Embodiment 2
[0099] In this embodiment, a retrospective inference system for laparoscopic surgical operation areas based on tracking is provided, including:
[0100] The first main control module is used to extract the operation screen of the surgical instrument based on the surgical flow video and determine the query frame;
[0101] The second main control module is used to perform point tracking processing on the entire surgical instrument operation and calculate the deformation field from other frames to the query frame based on the point tracking processing result;
[0102] The third main control module is used to obtain the area distribution of the surgical instrument tip on the query frame based on the deformation field and map the area distribution to other frames to obtain the operation area distribution of other frames;
[0103] The fourth main control module is used to fit and convert the operation area distribution into a binary map label.
[0104] Specifically, the first main control module needs to execute the following steps:
[0105] Based on the surgical video, select the pictures of the surgical operation scenes, segment the surgical instruments in the pictures to obtain the segmentation results of the surgical instruments, fit the pictures of the minimum circumscribed rectangles of the surgical instruments in consecutive specified frames, and on this basis, make labels for the surgical operation areas. Each consecutive segment needs to be processed separately, and the middle frame of each segment is selected as the query frame. Among them, fit the pictures of the minimum circumscribed rectangles of the surgical instruments in consecutive specified frames, and at the same time obtain the image coordinate positions of the tips of the surgical instruments as the positions of the surgical operation areas.
[0106] More specifically, after obtaining the surgical video, it is necessary to first select the pictures of the surgical operations before proceeding with the subsequent work. First, send all the pictures into the surgical instrument segmentation model to obtain the segmentation results of the surgical instruments, and fit their minimum circumscribed rectangles. As Figure 2 shown, retain the pictures of the segments with a length of the major axis of the minimum circumscribed rectangle of the surgical instrument contour greater than 1 / 10 of the length of the major axis of the surgical image and more than 200 consecutive frames. On this basis, make labels for the surgical operation areas. Each consecutive segment needs to be processed separately, and the middle frame of each segment is selected as the query frame to provide a reference for the subsequent calculation of the deformation field.
[0107] When fitting the minimum circumscribed rectangle, it is necessary to obtain the image coordinate positions of the tips of the surgical instruments as the positions of the surgical operation areas. Specifically as follows:
[0108] First, obtain the position P box of the center point of the circumscribed rectangle of the surgical instrument contour, the width W box , H box of the rectangle, and the rotation angle θ. Judge the direction of the major axis from the numerical values of the width and height, and calculate the direction vector of the major axis through the rotation angle of the rectangle:
[0109]
[0110] Then, obtain the offset offset of the surgical instrument contour Contours relative to the center point of the circumscribed rectangle frame:
[0111] offset = Contours - P box (0.10)
[0112] Calculate the projection lengths of the points on the contour on the major axis, that is, the inner product innerProb of the contour offset and the direction vector:
[0113]
[0114] The points with the largest or smallest inner product values are the tip points of the surgical instruments. First, take out the two points and mark them as P max and P min :
[0115]
[0116] Obtain the distance matrix D from non-edge pixels to the nearest edge pixel through the edge binary map of the surgical scene picture. The value on this matrix represents the Euclidean distance from the pixel at that location to the nearest edge pixel, and then query P max and P min The minimum distances from two points to the edge pixel, and take the point with the relatively larger distance as the tip point P of the instrument tip , and the calculation formula is:
[0117]
[0118] The second main control module needs to perform the following steps:
[0119] S201: Divide the segment of the surgical operation screen filtered out with a query frame as the boundary into two videos, the front and the back. Among them, both videos need to contain the query frame;
[0120] S202: Reverse-play the front video so that the query frame is the starting frame of the front video;
[0121] S203: Input the two videos, the front and the back, processed in step S202 into the point tracking model respectively to obtain the positions of the grid points in the query frame on other frames, so as to achieve the point tracking results of the grid points in the query frame;
[0122] S204: Obtain the deformation field from the query frame to each frame picture according to the point tracking results.
[0123] Specifically, to obtain the position information of the operation area of each frame in the continuous segment in the query frame, it is necessary to obtain the deformation field from each frame to the reference frame. Among them, the acquisition process of the deformation field is as Figure 3 shown. The point tracking code usually initializes the point grid with the first frame. Therefore, it is necessary to first divide the filtered surgical segment into two videos with the query frame as the boundary. The front video needs to be reverse-played to ensure that the query frame is the first frame, and the back video does not need additional processing. Input the two videos, the front and the back, of each segment into the point tracking network respectively to obtain the positions of the grid points in the query frame on other frames, and then calculate the deformation field based on the change of the grid point positions between two frames.
[0124] The size of the grid points in the query frame during the point tracking process is H×W, indicating that there are H rows and W columns of grid points evenly distributed on the query frame. Among them, the coordinates of each grid point p can be expressed as (x,y), and the tracking result of point p on other frames is p tracking , and the point p is transformed to the point p' on other frames through the deformation field, and its calculation formula is as follows:
[0125] p′ = p + f(p) (0.14)
[0126] Among them, f() is the deformation field to be obtained. Since this method is for two-dimensional images, the deformation field is also two-dimensional, which represents the horizontal and vertical displacements of grid points:
[0127] f(p) = (f x (p), f y (p)) (0.15)
[0128] Solving the deformation field is an optimization problem. The optimization goal is to align the grid points in the target frame with the prediction results of the deformation field on the query frame. For this purpose, the following loss function is used in this embodiment to solve the deformation field.
[0129] Tracking loss It is used to measure the distance between the grid points in the target frame and the predicted grid points. By minimizing this loss, the deformation field is optimized so that the deformation generated by the deformation field can align the prediction results. The loss function is expressed as:
[0130]
[0131] In the formula, N is the total number of grid points, that is, N = H × W, is the tracking loss, p i,tracking is the tracking result of the grid point p i on other frames, and p i is the position after being transformed by the current deformation field.
[0132] Through the iterative optimization of the deformation field by the above loss function, the deformation fields from the query frame to other frames of the segment are obtained.
[0133] The third main control module needs to execute the following steps:
[0134] Based on the deformation field, obtain the area distribution of the surgical instrument tip on the query frame, and map this area distribution to other frames to obtain the operation area distribution of other frames. More specifically, after obtaining the deformation field from the query frame to other frames, it is necessary to project the instrument tip points in other frames of the segment onto the query frame to obtain the operation area distribution on the query frame. As Figure 4 shown, after obtaining the operation area distribution points of the query frame, according to the deformation field matrix, map this distribution to other frames to provide the operation area distribution for all pictures.
[0135] The fourth main control module needs to execute the following steps:
[0136] S401: Based on the neighborhood density, remove the sparse points in the operation area distribution;
[0137] S402: Use Gaussian blur operation to smooth the point data of the operation area distribution, and process the point data into a heat map and then perform normalization processing;
[0138] S403: Input the heat map processed in step S402 and the original surgical scene image in the surgical video stream into a conditional random field. Utilize the color and spatial information of the original image to enhance the edge quality of the heat map, and then set a threshold for foreground and background segmentation to convert the heat map into a binary map label.
[0139] Specifically, the obtained operation area distribution information is a series of two-dimensional coordinate points, which need to be further converted into a binary map label for semantic segmentation before being used in network training. The processing flow is as Figure 5 shown. Initially filter the operation area points through domain density; use Gaussian blur operation to process the operation area points into a heat map and normalize the heat map result to the range of 0 to 1; send the heat map and the original surgical scene image into a conditional random field together, utilize the color and spatial information of the original image to improve the boundary of the heat map and remove noise areas at the same time; set a segmentation threshold to convert the heat map into a binary map label.
[0140] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A retrospective reasoning method for laparoscopic surgical operation areas based on tracking, characterized in that, The following steps are involved: S100: extracting the operation screen of the surgical instrument based on the surgical flow video and determining the query frame; S200: performing point tracking processing on the entire surgical instrument operation, and calculating deformation fields from other frames to the query frame based on the point tracking processing results; S300: Based on the deformation field, obtaining the area distribution of the surgical instrument tip on the query frame, and mapping the area distribution to other frames to obtain the operation area distribution of the other frames; S400: Fit the operation area distribution and convert it into a binary image label.
2. The retrospective inference method for laparoscopic surgical operation area based on tracking according to claim 1, wherein Step one includes: based on the surgical video, filter out pictures of surgical operation scenes, segment the surgical instruments in the pictures to obtain the segmentation results of the surgical instruments, fit the circumscribed rectangular pictures of surgical instruments in continuous specified frames, and create surgical operation area labels on this basis. Each continuous segment needs to be processed separately, and the middle frame of each segment is selected as the query frame.
3. The retrospective reasoning method for laparoscopic surgical operation area based on tracking according to claim 2, characterized in that In step 1, the minimum circumscribed rectangle image of the surgical instrument in consecutive specified frames is fitted, and the image coordinate position of the tip of the surgical instrument is obtained as the position of the surgical operation area.
4. A retrospective reasoning method for laparoscopic surgical operation areas based on tracking according to claim 3, characterized in that, The step of obtaining the image coordinate position of the surgical instrument tip as the position of the surgical operation area comprises: S101: Obtain the position P of the center point of the circumscribed rectangle of the surgical instrument contour box , the width W of the rectangle box , height H box and rotation angle θ, judge the long axis direction of the rectangle from the numerical sizes of the width W box , height H box , and calculate the direction vector of the long axis through the rectangle rotation angle: In the formula, is the direction vector of the major axis; S102: Obtain the offset of the surgical instrument outline relative to the center point of the circumscribed rectangular frame: offset = Contours - P box In the formula, Contours is the contour of the surgical instrument, and offset is the offset; S103: Calculate the projection lengths of points on the contour of the surgical instrument on the long axis according to the inner product of the direction vector and the offset. The point with the maximum or minimum projection length is the tip point of the surgical instrument, which is marked as P max and P min ; S104: Query P according to the distance matrix D from non-edge pixels to the nearest edge pixels in the picture of the surgical operation screen max and P min The minimum distances from two points to the edge pixels, and take the point with the relatively larger distance as the tip point P of the instrument tip .
5. A retrospective reasoning method for laparoscopic surgical operation area based on tracking according to claim 4, characterized in that, The tip point P tip The calculation formula thereof includes: Where D is the distance matrix from non-edge pixels to the nearest edge pixels.
6. A retrospective reasoning method for laparoscopic surgical operation areas based on tracking according to any one of claims 1-5, characterized in that Step S200 includes the following steps: S201: dividing the screened surgical operation scene segment into two videos, one before and one after the query frame, wherein both videos must contain the query frame; S202: playing the previous video in reverse so that the query frame is the starting frame of the previous video; S203: inputting the two videos before and after the processing in step S202 into the point tracking model respectively, obtaining the position of the grid point in the query frame on other frames, so as to realize the point tracking result of the grid point in the query frame; S204: Obtain the deformation field from the query frame to each frame image according to the point tracking result.
7. A retrospective reasoning method for laparoscopic surgical operation area based on tracking according to claim 6, characterized in that, In step S203, during the point tracking process, the size of the grid points in the query frame is H×W, indicating that there are H rows and W columns of grid points evenly distributed on the query frame. The coordinates of each grid point p are represented as (x, y), and the tracking result of point p on other frames is p tracking , and the point p is transformed to the point p' on other frames through the deformation field. Its calculation formula is as follows: p′=p+f(p) Where f() is the required deformation field, which represents the horizontal and vertical displacement of the grid points: f(p) = (f x (p), f y (p)) where, f x (p) is the horizontal displacement of the grid point, and f y (p) is the vertical displacement of the grid point.
8. A retrospective reasoning method for laparoscopic surgical operation area based on tracking according to claim 6, characterized in that, In step S204, the following loss function is used to solve the deformation field so that the deformation generated by the deformation field can align the prediction results: where N is the total number of grid points, i.e., N = H × W, is the tracking loss, p i,tracking is the tracking result of grid point p i on other frames, and p i is the position after being transformed by the current deformation field; Through the iterative optimization of the deformation field using the above loss function, the deformation field from the query frame to other frames of the fragment is obtained.
9. A retrospective reasoning method for laparoscopic surgical operation areas based on tracking according to claim 4, characterized in that Step S400 includes the following steps: S401: based on the neighborhood density, remove sparse points in the distribution of the operation area; S402: using Gaussian blur operation to smooth the point data distributed in the operation area, processing the point data into a heat map and then performing normalization processing; S403: Input the heat map processed by step S402 and the original image of the surgical scene in the surgical flow video into the conditional random field, use the color and spatial information of the original image to enhance the edge quality of the heat map, and then set the foreground and background segmentation threshold to convert the heat map into a binary image label.
10. A retrospective reasoning system for laparoscopic surgical operation areas based on tracking, characterized in that, include: A first main control module is used to extract the operation screen of the surgical instrument based on the surgical flow video and determine the query frame; The second main control module is used to perform point tracking processing on the entire surgical instrument operation and calculate the deformation field from other frames to the query frame based on the point tracking processing results; The third main control module is used to obtain the regional distribution of the surgical instrument tip on the query frame based on the deformation field and map the regional distribution to other frames to obtain the operation regional distribution of other frames; The fourth main control module is used to fit the operation regional distribution and convert it into a binary map label.
Citation Information
Cited By
Laser operation coordinate positioning method and system based on robot and program product
CN121120775A