A real-time detection method for the pose of surgical instruments based on images
Through the improved image detection network, combined with the CBS module, GAM attention mechanism, SIoU loss function and PSPNet segmentation network, the problems of occlusion, jitter, reflection and insufficient light in surgical instrument detection are solved, and high-precision and high-real-time posture detection of surgical instruments are achieved, improving surgical safety.
Patent Information
- Application Number
- CN202310220537.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-03-09
AI Technical Summary
The existing surgical instrument detection methods are susceptible to interference from the instrument being blocked, shaking, incomplete display, reflection and insufficient light in real surgery, resulting in insufficient detection accuracy and real-time performance, which is difficult to meet the high safety needs of surgical robots.
Build an improved image detection network, use CBS module to replace the CBL module, introduce the GAM attention mechanism, use SIoU loss function and increase the PSPNet semantic segmentation network, combine the detection and segmentation network, and process the surgical instrument image data set through data enhancement to improve the robustness and adaptability of the network.
High-precision, high-real-time and high-rootability posture detection of surgical instruments is achieved in complex surgical environments, which can accurately identify the type and posture of the instrument, reduce missed examinations, and improve the safety and reliability of the operation.
Smart Images

Figure CN116110545B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of medical robot target detection and image segmentation, and particularly relates to a method for real-time detection of the pose of a surgical instrument based on images. Background Art
[0002] With the continuous development of robot technology and medical technology, robot-assisted minimally invasive surgery has gradually played an important role in the medical field and has become the preferred type of surgery for treating various diseases. Approximately more than 4,000 robots participate in surgical operations every day. Compared with traditional surgery, it has the characteristics of less trauma, less pain, and rapid postoperative recovery. During the operation, there is no need for a doctor to hold the surgical instrument for a long time, thus avoiding mistakes caused by human factors such as hand tremors, reducing the doctor's fatigue and improving the safety of the operation. It is precisely because minimally invasive surgery has the above advantages in various aspects that medical surgical robots have gradually attracted the attention of many researchers and become a research hotspot.
[0003] The detection and segmentation methods for surgical instruments can be mainly divided into methods based on hardware devices, methods based on traditional image algorithms, and methods based on deep learning. As early as the 1980s, Roberts et al. from Stanford University in the United States effectively combined microscopes and CT images and successfully navigated clinical surgeries by using ultrasonic waves to locate surgical instruments. However, this hardware device-dependent method not only requires considering the cost of the instrument but is also rather cumbersome in actual operation.
[0004] With the rapid development of robotics, endoscopic imaging has become clearer, enabling doctors to clearly observe the surgical area from the imaging device. Therefore, image-based detection methods have gradually gained attention. Different from hardware-based methods, they do not require expensive equipment and are more convenient to deploy. They can directly present the position, category, and posture of surgical instruments in the image. In 1994, Lee C et al. first extracted surgical instruments from laparoscopic surgery videos using color features. In 2003, Krupa A et al. installed light-emitting diodes on the tips of surgical instruments and used laser pointers and optical markers to locate and detect surgical instruments, and successfully conducted experiments on live animals. However, adding auxiliary tools to the tips of surgical instruments may affect the safety of the surgery. In 2004, Doignon et al. located and detected surgical instruments through color features and gray-scale region detection. In 2013, Sznitman R et al. combined the active testing paradigm and Bayesian filters to detect and track surgical instruments in retinal surgery. In 2013, Allan et al. proposed using color space analysis to segment surgical instruments, achieving high precision but being easily affected by the lighting conditions in the environment. In 2015, Bouget et al. proposed a two-stage surgical instrument detection method. In the first stage, each pixel of the local appearance is classified, and in the second stage, the global shape is matched. In 2016, Rieke et al. proposed a detection method based on random forests. By shrinking the candidate region to a rectangular region around the tip of the surgical instrument, the position of the instrument tip is inferred based on the histogram of oriented gradients features within the bounding box.
[0005] Deep learning-based methods generally rely on convolutional neural networks. In 2015, Sarikaya et al. proposed a multi-modal two-stream convolutional network by improving the Faster R-CNN algorithm, which can fuse image and temporal motion cues for joint detection of minimally invasive surgical instruments. In 2016, Garcia et al. proposed a method for surgical instrument segmentation that combines a Fully Convolutional Network (FCN) with optical flow tracking. In 2016, Sahu et al. combined a convolutional neural network with a random forest to achieve the detection of surgical instruments. In 2017, Attia et al. combined a convolutional neural network and a recurrent neural network to propose a new hybrid model for segmenting surgical instruments. In the same year, Chen Zhaorui et al. proposed a method that combines a convolutional neural network with a Line Segment Detector (LSD) to detect the tip of surgical instruments and uses Spatio-Temporal Context (STC) learning to automatically track the tip frame by frame. In 2018, Jin et al. based on the Faster R-CNN detection and tracking model, used a complete laparoscopic surgery video to achieve the classification, detection, and tracking of surgical instruments. In 2020, Tamer et al. proposed a method that combines a convolutional neural network and a long short-term memory model to detect surgical instruments in laparoscopic images. In 2021, Zhou et al. proposed a hierarchical feature fusion strategy, which enhanced the network's ability to fuse multi-scale feature maps and improved the accuracy of surgical instrument segmentation.
[0006] In the medical field, current surgical robots are increasingly emphasizing integration with computer vision, especially deep learning algorithms, to provide doctors with precise surgical instrument positioning and pose detection, etc., which can reduce the burden on doctors to search for surgical instruments in the imaging system when operating surgical robots. The essence of surgical instrument pose detection is to analyze the image or video through algorithms, extract the features of surgical instruments in the image, and thus determine information such as the type, real-time position, and spatial pose of surgical instruments, providing navigation for the operation of the surgery and further improving the safety of the surgery. Therefore, developing a pose detection algorithm that meets real-time, accuracy, and robustness requirements is crucial for the development of surgical robots. Summary of the Invention
[0007] The purpose of the present invention is to provide an image-based real-time surgical instrument pose detection method that solves technical problems such as surgical instruments being blocked by tissues and organs, smoke, instrument jitter, incomplete display, instrument reflection, and insufficient ambient light.
[0008] To this end, the technical solution of the present invention is as follows:
[0009] A real-time detection method for the pose of surgical instruments based on images, the steps are as follows:
[0010] Step 1: Construct an image detection network suitable for real-time pose recognition of surgical instruments, and its steps are as follows:
[0011] Step 1.1: In the backbone network and the neck network, replace the CBL module with the CBS module; the CBS module consists of a convolutional layer, a BN layer, and a SiLU activation function connected in sequence;
[0012] Step 1.2: Introduce the GAM attention mechanism into the backbone network, that is, add a GAM module behind the SPPF module, connect the output end of the SPPF module to the input end of the GAM module, and connect the output end of the GAM module to the input end of the first CBS module of the neck network;
[0013] Step 1.3: In the prediction network, replace the CIoU loss function with the SIoU loss function;
[0014] Step 1.4: Add the semantic segmentation network PSPNet, whose input end is respectively connected to the output ends of the first CBS module of the backbone network, the second CBS module of the backbone network, and the second CSP2_1 module of the neck network, and its output end is respectively connected to the output ends of the first Conv module, the second Conv module, and the third Conv module of the prediction network;
[0015] Step 2: Construct a surgical instrument image dataset for training the image detection network, which consists of a detection dataset and a segmentation dataset; among them, the detection dataset is obtained by annotating the tip of the surgical instrument with the smallest rectangular box based on the images containing different types of surgical instruments, and each annotated image generates image detection information, including the surgical instrument category, the four vertex coordinates of the rectangular box, the length and width of the rectangular box; the segmentation dataset is obtained by annotating the contour of the tip of the surgical instrument with a closed contour box based on the images containing different types of surgical instruments, and each annotated image generates image segmentation information, including the surgical instrument category and the multi-point coordinates accurately depicting the contour box;
[0016] Step 3: Perform data augmentation processing on the surgical instrument image dataset obtained in Step 2 to expand the data volume of the surgical instrument image dataset;
[0017] Step 4: Use the surgical instrument image dataset obtained in Step 3 to train the image detection network constructed in Step 1, including:
[0018] Step 4.1: Randomly divide the surgical instrument image dataset into a training set and a test set;
[0019] Step 4.2: Input the training set into the image detection network. First, use the labeled images in the detection dataset as the input and the image detection information as the output to train and validate the image detection network. Then, use the labeled images in the segmentation dataset as the input and the image segmentation information as the output to train and validate the image detection network;
[0020] Step 5: Input the surgically instrument images collected in real time into the image detection network trained in Step 4, and then the labeled images with the accurately detected postures of the surgical instruments can be output in real time.
[0021] Furthermore, in Step S1, the structure of the image detection network includes:
[0022] (1) Image input end;
[0023] (2) Backbone network, which consists of a first CBS module, a second CBS module, a first CSP1_1 module, a third CBS module, a CSP1_2 module, a fourth CBS module, a CSP1_3 module, a fifth CBS module, a second CSP1_1 module, an SPPF module, and a GAM module connected in sequence;
[0024] (3) Neck network, which consists of a first CBS module, a first upsampling module, a first Concat module, a first CSP2_1 module, a second CBS module, a second upsampling module, a second Concat module, a second CSP2_1 module, a third CBS module, a third Concat module, a third CSP2_1 module, a fourth CBS module, a fourth Concat module, and a fourth CSP2_1 module connected in sequence; Among them, the input end of the first Concat module of the neck network is also connected to the output end of the CSP1_3 module of the backbone network, the input end of the second Concat module of the neck network is also connected to the output end of the first CSP1_2 module of the backbone network, the input end of the third Concat module of the neck network is also connected to the output end of the second CBS module of the neck network, and the input end of the fourth Concat module of the neck network is also connected to the output end of the first CBS module of the neck network;
[0025] (4) Semantic segmentation network, whose input ends are respectively connected to the output ends of the first CBS module of the backbone network, the second CBS module of the backbone network, and the second CSP2_1 module of the neck network, and whose output ends are respectively connected to the output ends of the first Conv module, the second Conv module, and the third Conv module of the prediction network;
[0026] (5) Prediction network, which consists of a first Conv module, a second Conv module, and a third Conv module; The loss function in the prediction network uses the SIoU loss function.
[0027] Further, the specific implementation steps of step 2 are as follows:
[0028] Step 2.1: Collect images of different surgical instruments during the operation to obtain an original image set; among them, the number of images in the original image set should meet the requirements of network training, and the proportion of the number of images of different surgical instruments in the total number of images is the same or approximately the same;
[0029] Step 2.2: Use the image annotation software labelme to manually annotate each image obtained in step 1.1 in turn to obtain a surgical instrument image data set; the surgical instrument image data set consists of a detection data set and a segmentation data set, where,
[0030] (1) The construction method of the detection data set is as follows: Based on the original image set obtained in step 2.1, each image is marked as follows: ① If more than half of the tip part of the surgical instrument in the image is shown, then circle the visible part in the image in the form of a minimum rectangle; ② If less than half of the tip part of the surgical instrument in the image is shown, then do not mark; Generate image detection information corresponding to the images marked with the rectangle, including the surgical instrument category, the four vertex coordinates of the rectangle, the length and width of the rectangle; Images that are not marked do not generate image detection information; Finally, the marked images marked based on the original image set and the image detection information corresponding to the marked images one by one constitute the detection data set;
[0031] (2) The construction method of the segmentation data set is as follows: Based on the original image set obtained in step 2.1, each image is marked as follows: ① If more than half of the tip part of the surgical instrument in the image is shown, then draw a closed contour box along the contour of the visible part in the image; ② If less than half of the tip part of the surgical instrument in the image is shown, then do not mark; Furthermore, a segmentation data set composed of marked images is obtained; Generate image segmentation information corresponding to the images marked with the rectangle, including the surgical instrument category and the multi-point coordinates that can accurately depict the contour box; Images that are not marked do not generate image segmentation information; Finally, the marked images marked based on the original image set and the image detection information corresponding to the marked images one by one constitute the detection data set.
[0032] Further, the specific implementation of step 3 is as follows:
[0033] Step 3.1: Use the Mosaic data augmentation method to perform data augmentation on the surgical instrument image data set obtained in step 3;
[0034] Step 3.2: Use the Albumentations image enhancement library to perform data enhancement on the surgical instrument image dataset obtained in Step 3 by using at least one of the methods of random color, contrast transformation, brightness transformation, noise perturbation, random cropping, scaling, translation, and rotation in sequence.
[0035] Compared with the prior art, the real-time detection method for the pose of surgical instruments based on images solves the problems of occlusion of instruments, instrument jitter, incomplete display, insufficient ambient light or instrument reflection, etc. in real surgical scenarios. When affected by common interferences, it can still correctly detect surgical instruments, showing high adaptability in complex and changeable surgical environments. It is a method with high precision, high real-time performance, and high robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of the real-time detection method for the pose of surgical instruments based on images of the present invention;
[0037] Figure 2 is a schematic structural diagram of the image detection network in the real-time detection method for the pose of surgical instruments based on images of the present invention;
[0038] Figure 3 is a schematic structural diagram of the GAM module of the image detection network in the real-time detection method for the pose of surgical instruments based on images of the present invention;
[0039] Figure 4 is a schematic diagram of the present invention's real-time detection method for the pose of surgical instruments based on images, which respectively includes seven surgical instrument images;
[0040] Figure 5 is a schematic diagram of the labeled image of the detection dataset in the real-time detection method for the pose of surgical instruments based on images of the present invention;
[0041] Figure 6 is a schematic diagram of the labeled image of the segmentation dataset in the real-time detection method for the pose of surgical instruments based on images of the present invention;
[0042] Figure 7 is an example diagram before and after processing the original labeled image by using the Mosaic data enhancement method in the real-time detection method for the pose of surgical instruments based on images of the present invention;
[0043] Figure 8 is an example diagram before and after processing the original labeled image by using the Albumentations image enhancement library in the real-time detection method for the pose of surgical instruments based on images of the present invention;
[0044] Figure 9It is the PSPNet network structure diagram in the real-time detection method of surgical instrument posture based on images of the present invention;
[0045] Figure 10 It is a schematic diagram of the confusion matrix for binary classification;
[0046] Figure 11 It is a schematic diagram of the P-R curve drawn by different types of surgical instruments through precision and recall in the embodiment of the present invention;
[0047] Figure 12 It is a schematic diagram of the intersection over union (IoU), an evaluation index in semantic segmentation, in the real-time detection method of surgical instrument posture based on images of the present invention;
[0048] Figure 13(a) is a comparison diagram of detection results where there are missed detections in the YOLOv5 network compared with the present application in the embodiment of the present invention;
[0049] Figure 13(b) is a comparison diagram of detection results where there are mis-detections in the YOLOv5 network compared with the present application in the embodiment of the present invention;
[0050] Figure 14 It is a schematic diagram of the confusion matrix of various types of surgical instruments in the real-time detection method of surgical instrument posture based on images of the present invention;
[0051] Figure 15(a) is a precision curve graph of seven types of surgical instruments in the embodiment of the present invention;
[0052] Figure 15(b) is a confidence curve graph of seven types of surgical instruments in the embodiment of the present invention;
[0053] Figure 16(a) is an F1 curve graph of seven types of surgical instruments in the embodiment of the present invention;
[0054] Figure 16(b) is the P-R curve of seven types of surgical instruments in the embodiment of the present invention;
[0055] Figure 17 It is a change curve graph of the loss function of the training set in the real-time detection method of surgical instrument posture based on images in the embodiment of the present invention;
[0056] Figure 18 It is a change curve graph of the loss function of the validation set in the real-time detection method of surgical instrument posture based on images in the embodiment of the present invention;
[0057] Figure 19(a) is a detection result graph of the image with tissue and organ occlusion in the embodiment of the present invention;
[0058] Figure 19(b) is a detection result graph of the image with smoke occlusion in the embodiment of the present invention;
[0059] Figure 19(c) is the detection result diagram of the incomplete display of instruments in the image in the embodiment of the present invention;
[0060] Figure 19(d) is the detection result diagram of the instrument jitter in the image in the embodiment of the present invention;
[0061] Figure 19(e) is the detection result diagram of insufficient light in the image in the embodiment of the present invention;
[0062] Figure 19(f) is the detection result diagram of reflection in the image in the embodiment of the present invention;
[0063] Figure 20 is the visualization effect diagram of the surgical instrument pose detection result obtained by using the method of the present application in the embodiment of the present invention. Detailed implementation manners
[0064] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the following embodiments are by no means any limitation to the present invention.
[0065] Refer to Figure 1 , and the specific implementation steps of the image-based surgical instrument pose real-time detection method are as follows:
[0066] Step 1: Construct an image detection network suitable for real-time pose recognition of surgical instruments:
[0067] This image detection network is based on the YOLOv5 network and is constructed by improving the module connection structures in the original backbone network, neck network, and prediction network in sequence and adding a semantic segmentation network.
[0068] Specifically, the construction steps of this image detection network are as follows:
[0069] Step 1.1: Replace the CBL module with the CBS module in the backbone network and the neck network; among them, the original CBL module consists of a convolutional layer, a BN layer, and a LeakyReLU activation function connected in sequence. In the present application, the LeakyReLU activation function is replaced with a SiLU activation function, that is, a new CBS module is formed. That is, the CBS module consists of a convolutional layer, a BN layer, and a SiLU activation function connected in sequence. Refer to Figure 3 ; in this step, replacing the CBL module with the CBS module can effectively avoid gradient explosion and alleviate the situation of unstable training;
[0070] Step 1.2: Introduce the GAM attention mechanism into the backbone network, that is, add a GAM module after the SPPF module, connect the output end of the SPPF module to the input end of the GAM module, and connect the output end of the GAM module to the input end of the first CBS module of the neck network; in this step, by introducing the GAM attention mechanism into the backbone network, the network can pay more attention to important information by allocating weights, thereby improving the network's feature extraction ability.
[0071] Step 1.3: In the prediction network, replace the CIoU loss function with the SIoU loss function; the reason for modifying the loss function in this step is that the CIoU loss function is used in the existing YOLOv5 network, and the parameter v introduced by it reflects the relative value of the aspect ratio difference, which has a certain ambiguity and may sometimes hinder the optimization of the model; while in this application, replacing the CIoU loss function with the SIoU loss function can reduce the degree of freedom of regression and accelerate network convergence by introducing the vector angle relationship between the ground truth box and the predicted box, further improving the regression accuracy.
[0072] Step 1.4: Add the semantic segmentation network PSPNet, see Figure 9 , its input end is respectively connected to the output ends of the first CBS module of the backbone network, the second CBS module of the backbone network, and the second CSP2_1 module of the neck network, and its output end is respectively connected to the output ends of the first Conv module, the second Conv module, and the third Conv module of the prediction network; in this step, the addition of the PSPNet semantic segmentation head enables parallel detection and segmentation, achieving the purpose of detecting the pose of surgical instruments based on the image.
[0073] In summary, the image detection network of this application is constructed by combining object detection and semantic segmentation, in order to be able to detect and identify the type and pose of surgical instruments based on the image simultaneously, see Figure 9 .
[0074] See Figure 2 , based on the above construction steps, the specific structure of the image detection network of this application suitable for real-time pose recognition of surgical instruments is:
[0075] (1) Image input end, with an image input size of 640×640×3;
[0076] (2) The backbone network, which consists of a first CBS module, a second CBS module, a first CSP1_1 module, a third CBS module, a CSP1_2 module, a fourth CBS module, a CSP1_3 module, a fifth CBS module, a second CSP1_1 module, an SPPF module, and a GAM module connected in sequence; among them, each CBS module has the same structure and is composed of a convolutional layer, a BN layer, and a SiLU activation function connected in sequence;
[0077] (3) The neck network, which consists of a first CBS module, a first upsampling module, a first Concat module, a first CSP2_1 module, a second CBS module, a second upsampling module, a second Concat module, a second CSP2_1 module, a third CBS module, a third Concat module, a third CSP2_1 module, a fourth CBS module, a fourth Concat module, and a fourth CSP2_1 module connected in sequence; among them, the input end of the first Concat module of the neck network is also connected to the output end of the CSP1_3 module of the backbone network, the input end of the second Concat module of the neck network is also connected to the output end of the first CSP1_2 module of the backbone network, the input end of the third Concat module of the neck network is also connected to the output end of the second CBS module of the neck network, and the input end of the fourth Concat module of the neck network is also connected to the output end of the first CBS module of the neck network;
[0078] (4) The semantic segmentation network, which specifically uses PSPNet, and its input ends are respectively connected to the output ends of the first CBS module of the backbone network, the second CBS module of the backbone network, and the second CSP2_1 module of the neck network, and its output ends are respectively connected to the output ends of the first Conv module, the second Conv module, and the third Conv module of the prediction network;
[0079] (5) The prediction network, which consists of a first Conv module, a second Conv module, and a third Conv module; among them, the input end of the first Conv module is connected to the output end of the second CSP2_1 module to output a feature map of 80×80×255; the input end of the second Conv module is connected to the output end of the third CSP2_1 module to output a feature map of 40×40×255; the input end of the third Conv module is connected to the output end of the fourth CSP2_1 module to output a feature map of 20×20×255; at the same time, the CIoU loss function in the prediction network is replaced with the SIoU loss function; finally, the output result of the prediction network is the result map after fusing three scales;
[0080] Step 2: Construct a surgical instrument image dataset for training the image detection network:
[0081] Specifically, the implementation steps of step 2 are as follows:
[0082] Step 2.1: Collect images of different surgical instruments during the operation to obtain the original image set. Among them, the number of images in the original image set should meet the requirements of network training, and the proportion of the number of images of different surgical instruments in the total number of images is the same or approximately the same;
[0083] In this embodiment, the images of this step directly adopt the m2cai16 - tool - locations dataset; see Figure 4 , the surgical instruments in this image dataset include seven types, namely, grasper, surgical scissors, surgical forceps, surgical hook, irrigator, bipolar clamp, and specimen bag, with a total of 2532 frames.
[0084] Step 2.2: Use the image annotation software labelme to manually annotate each image obtained in step 1.1 in sequence to obtain the surgical instrument image dataset. Among them, the surgical instrument image dataset is composed of a detection dataset and a segmentation dataset, and the specific construction methods of the two are as follows:
[0085] (1) Construct the detection dataset:
[0086] Based on the original image set obtained in step 2.1, each image is marked as follows: ① If more than half of the tip part of the surgical instrument in the image is shown, then circle the visible part of the image in the form of a minimum rectangle, see Figure 5 ; ② If less than half of the tip part of the surgical instrument in the image is shown, then do not mark;
[0087] Generate image detection information corresponding to the above - marked images with rectangles, including the surgical instrument category, the four vertex coordinates of the rectangle, the length and width of the rectangle; images without marks do not generate image detection information;
[0088] Finally, the marked images marked based on the original image set and the image detection information corresponding one - to - one with the marked images constitute the detection dataset;
[0089] (2) Construct the segmentation dataset:
[0090] Based on the original image set obtained in step 2.1, each image is marked as follows: ① If more than half of the tip part of the surgical instrument in the image is shown, then draw a closed contour box along the contour of the visible part of the image, see Figure 6 ; ② If less than half of the tip part of the surgical instrument in the image is shown, then do not mark; thus, a segmentation dataset composed of marked images is obtained;
[0091] Generate image segmentation information corresponding to the above-mentioned images marked with rectangular frames, including the surgical instrument category and the multi-point coordinates that can accurately depict the contour frame; unmarked images do not generate image segmentation information;
[0092] Finally, the marked images obtained based on the original image set and the image detection information corresponding one-to-one to the marked images constitute the detection data set.
[0093] In this embodiment, since both the detection data set and the segmentation data set are obtained by marking based on the original image set, the applicant constructed the surgical instrument image data set 1 by using the method of separately constructing the detection data set and the segmentation data set, and also used the method of simultaneously marking rectangular frames and contour frames in each image based on the original image set, but generating image detection information and image segmentation information corresponding to the images respectively to construct the surgical instrument image data set 2; after comparison, in the subsequent network training, the surgical instrument image data sets obtained by these two methods have the same network training effect.
[0094] Step 3. Perform data augmentation on the surgical instrument image data set obtained in Step 2:
[0095] Specifically, the implementation steps of this Step 3 are as follows:
[0096] Step 3.1. Refer to Figure 7 , and use the Mosaic data augmentation method to perform data augmentation on the surgical instrument image data set obtained in Step 3; specifically, after performing Mosaic data augmentation, the detection data set is enriched, a lot of small target information is added, and the information that the network can learn is improved. At the same time, in the training process, the BN layer calculates 4 images each time, indirectly increasing the batch_size and reducing the memory requirement. In addition to the 4-Mosaic data augmentation, YOLOv5 also newly adds a 9-Mosaic data augmentation, and the principle is the same as that of the 4-Mosaic data augmentation, except that the number of images used becomes 9; however, there are also some defects in the Mosaic data augmentation. When there are already a large number of small targets in the data set, the small targets will become smaller after data augmentation, resulting in a poor generalization ability of the model;
[0097] Step 3.2: Use the Albumentations image enhancement library to perform data enhancement on the surgical instrument image dataset obtained in Step 3. Specifically, the Albumentations image enhancement library includes various data enhancement methods, including random color, contrast transformation, brightness transformation, noise perturbation, random cropping, scaling, translation, and rotation, a total of eight methods. Among them, in the color transformation method, the color is set to be randomly transformed; in the contrast transformation method, the contrast enhancement value is set to be 1.5 times the original image, and the reduction value is set to be 0.9 times the original image; in the brightness transformation method, the brightness enhancement value is set to be 1.5 times the original image, and the reduction value is set to be 0.9 times the original image; in the noise perturbation method, Gaussian noise is added; in the image rotation method, the rotation angle is set to 100°. As Figure 8 shown in the figure is the effect display diagram before and after the data enhancement process of Step 3.2 on the image in this embodiment. The methods used from top to bottom in the figure are the brightness enhancement method, the random cropping method, and the contrast transformation method;
[0098] In this Step 3, all the images in the surgical instrument image dataset have completed the data enhancement process of the images by combining various methods, further expanding the amount of image data in the dataset, avoiding overfitting of the trained image detection network, and effectively improving the generalization ability of the image detection network. In this embodiment, the number of dataset images after various data enhancement methods is 5059.
[0099] Step 4: Use the surgical instrument image dataset obtained in Step 3 to train the image detection network constructed in Step 1 to achieve the purpose of being able to detect the pose of surgical instruments based on images in real time;
[0100] Specifically, the implementation steps of this Step 4 are as follows:
[0101] Step 4.1: Divide the surgical instrument image dataset obtained in Step 3 into a training set, a test set, and a validation set according to a ratio of 8:1:1;
[0102] Step 4.2: Input the training set obtained by the division in Step 4.1 into the image detection network constructed in Step 1. First, train the image detection network with the labeled images in the detection dataset as the input and the image detection information as the output, and then train the image detection network with the labeled images in the segmentation dataset as the input and the image segmentation information as the output. At the same time, during the training process, use the test set to perform ablation experiment verification on the network after each round of training. Finally, take the weights corresponding to the round with the best recognition accuracy during training as the optimal weights and save them, and the training is completed.
[0103] In this embodiment, the training parameters are shown in Table 1 below.
[0104] Table 1:
[0105]
[0106]
[0107] Step 5: Input the images of surgical instruments collected in real time into the image detection network trained in Step 4, and then the marked images with the accurately detected postures of surgical instruments can be output in real time.
[0108] Furthermore, in order to evaluate the effectiveness of the method for detecting the posture of surgical instruments, the detection accuracy and speed of the method of this application are respectively evaluated; specifically, the evaluation indicators include Precision, Recall, Average Precision, mean Average Precision (mAP), and the number of images detected per second (FPS).
[0109] As Figure 10 shown is a schematic diagram of the confusion matrix for binary classification; in the figure, T and F refer to the positive and negative samples of the true value, and P and N are the positive and negative of the predicted value; therefore, there are a total of four classification results; among them, TP (True Positive) means that the true value is a positive sample and the prediction is also a positive sample; FP (False Positive) means that the true value is a negative sample and the prediction is a positive sample; FN (False Negative) means that the true value is a positive sample and the prediction is a negative sample; TN (True Negative) means that the true value is a negative sample and the prediction is also a negative sample. It is not difficult to see that FN and FP are two types of prediction errors. Therefore, generally, we hope that the values of TP and TN are as large as possible, and the values of FN and FP are as small as possible.
[0110] Precision is the proportion of the correctly predicted surgical instrument samples among all the samples predicted as surgical instruments. It is used to measure the ability of the model to find the correct ones, and its expression:
[0111]
[0112] Recall is the proportion of the correctly predicted surgical instrument samples among all the actual surgical instrument samples. It is used to measure the ability of the model to find all, and its expression:
[0113]
[0114] During the training process of the model, multiple sets of P and R values are generated. By statistically analyzing these values and plotting the precision on the vertical axis and the recall on the horizontal axis, the P-R curve (Precision-Recall curve) can be obtained. See Figure 11 The P-R curve can be used to reflect the superiority of the detection performance of the detection network for a single surgical instrument category.
[0115] The average precision refers to the mean precision of the surgical instruments of a certain category, that is, the area enclosed by the curve of this category and the coordinate axes. Its expression is:
[0116]
[0117] The mean average precision is the average of the mean precisions of all categories of surgical instruments, that is, the average of the areas enclosed by the curves of all categories and the coordinate axes. Its expression is:
[0118]
[0119] In the formula, N refers to the total number of surgical instrument categories.
[0120] To comprehensively evaluate the precision and recall, in addition to mAP, the F1 score (F1Score) can also be used. Its expression is:
[0121]
[0122] The F1 score is the harmonic mean of the precision and recall. Its value ranges from 0 to 1, and the closer it is to 1, the better the model performance.
[0123] Semantic segmentation is mainly for pixel-level detection. Therefore, the concept of intersection over union is first explained. The intersection over union (IoU) is the ratio of the intersection area to the union area of the ground truth box and the predicted box of the target, which is used to measure the overlapping degree of the ground truth box and the predicted box. The expression of IoU is:
[0124]
[0125] Among them, the threshold of IoU is set to 0.5; when IoU is greater than 0.5, it is considered that the predicted box overlaps with the ground truth box, that is, the detection result is correct; when IoU is less than 0.5, the predicted box does not overlap with the ground truth box, that is, the detection result is incorrect; as Figure 12 shown, the red box in the figure is the ground truth box A, and the blue box is the predicted box B.
[0126] The most important evaluation metrics in semantic segmentation are the Mean Intersection over Union (MIoU) and the Dice coefficient. The MIoU is the average of the sum of the intersection over union between the predicted values and the ground truth values for each surgical instrument category. The Dice coefficient is a measure of set similarity and is commonly used to calculate the similarity between two samples. The values of both of the above metrics range from 0 to 1, and the closer to 1, the better the segmentation performance. Their expressions are as follows:
[0127]
[0128]
[0129] In the formula, k is the number of surgical instrument categories;
[0130] To verify that the method of this application has better effects compared with the prior art, while evaluating the above metrics, other technical methods are evaluated for metrics to form a comparative reference.
[0131] First, to judge and prove the effectiveness of the method of this application, ablation experiments are respectively carried out using the same dataset of this application; among them, as shown in Table 2, the network models are specifically as follows: "YOLOv5" means using the YOLOv5 network as the image detection network; "YOLOv5+GAM" means using the network improved from the YOLOv5 network in step 1.2 of this application as the image detection network; "YOLOv5+SiLU" means using the network improved from the YOLOv5 network only in step 1.1 of this application as the image detection network; "YOLOv5+GAM+SiLU" means using the network improved from the YOLOv5 network only in step 1.1 and step 1.2 of this application as the image detection network; "YOLOv5+GAM+SiLU+SIoU" means using the network improved from the YOLOv5 network only in step 1.1, step 1.2 and step 1.3 of this application as the image detection network; "YOLOv5+GAM+SiLU+SIoU+PSPNet segmentation head" means the image detection network of this application, that is, using the network improved from the YOLOv5 network in step 1.1, step 1.2, step 1.3 and step 1.4 of this application as the image detection network;
[0132] The specific test results of this ablation experiment are shown in Table 2 below.
[0133] Table 2:
[0134]
[0135]
[0136] From the experimental results in Table 2, it can be seen that the independent introduction of the GAM attention mechanism can effectively enhance the network's ability to extract important features, and its precision, recall, mAP, and FPS have all improved; when all CBL modules are replaced with CBS modules alone, the problem of gradient explosion is avoided, that is, the instability of training is alleviated; and based on the introduction of the GAM attention mechanism and at the same time replacing all CBL modules with CBS modules, although the precision has decreased to some extent, the other indicators have all improved; when the original loss function in the prediction network is further replaced with the SIoU loss function, it shows that the convergence speed of the network is accelerated; and further combined with the PSPNet semantic segmentation head, due to predicting the contour of the surgical instrument, the FPS has decreased to some extent, but the prediction is more accurate and the other indicators have all improved; finally, based on the above four ways of improving the YOLOv5 network, compared with the original YOLOv5 network, the precision of the image detection network of this application has increased by 1%, the recall has increased by 5.5%, the mAP has increased by 2%, and the FPS has increased by 29.5%.
[0137] Furthermore, in order to visually observe the detection effect of the improved algorithm, some detection results before and after improvement are visualized and compared.
[0138] As shown in Figure 13(a), it is a comparison diagram of the detection results where the YOLOv5 network has a missed detection problem compared with this application; among them, the left figure is the detection result obtained by training with the YOLOv5 network, and the right figure is the detection result obtained by training the image detection network of this application; from this group of comparison diagrams, it can be seen that the original YOLOv5 has a missed detection problem, and the improved image detection network not only solves these problems, but also reaches a higher level in detection accuracy.
[0139] As shown in Figure 13(b), it is a comparison diagram of the detection results where the YOLOv5 network has a misdetection problem compared with this application; among them, the left figure is the detection result obtained by training with the YOLOv5 network, and the right figure is the detection result obtained by training the image detection network of this application; from the above two groups of comparison diagrams, it can be seen that the original YOLOv5 has a misdetection problem, and the improved image detection network not only solves these problems, but also reaches a higher level in detection accuracy.
[0140] As Figure 14 shown is the confusion matrix of various types of surgical instruments, which is a summary of the classification results of the model for all types of surgical instruments in the dataset, showing which type of surgical instrument the detection network will be confused about when making predictions, not only being able to understand the errors generated by the model in classification, but also further knowing the types of errors that occur.
[0141] As shown in Table 3, the detection results of the precision, recall rate, and average detection accuracy of the seven surgical instruments included in the surgical instrument dataset using the method of this application are presented.
[0142] Table 3:
[0143] Surgical instrument category Precision(%) Recall(%) AP(%) Grasper 93.4 92.8 96.7 Bipolar 97.5 96.7 96.6 Hook 99.2 100 99.5 Scissors 96.4 100 97.8 Clipper 99.5 100 99.5 Irrigator 99.7 100 99.5 SpecimenBag 95.7 94.8 95.9 mAP(%) NA NA 97.9
[0144] From the detection results in Table 3, it can be seen that the precision of the seven surgical instruments is between 93.4% and 99.7%, the recall rate is between 92.8% and 100%, and the average detection accuracy is between 95.9% and 99.5%.
[0145] As shown in Figure 15(a), it is the precision curve graph of the seven categories of surgical instruments; it can be seen from the figure that the detection accuracy of the surgical instruments is very high and can be effectively applied to image processing.
[0146] As shown in Figure 15(b), it is the confidence curve graph of the seven categories of surgical instruments; it can be seen from the figure that the IoU effect between each detected bounding box and the matching ground truth is good, and the category and contour of the surgical instruments can be detected more accurately.
[0147] As shown in Figure 16(a), it is the F1 curve graph of the seven categories of surgical instruments; among them, the F1 score is the harmonic mean of the precision rate and the recall rate, and its value is between 0 and 1. The closer it is to 1, the better the effect of the model; it can be seen from the figure that the image detection network of this application has good detection performance for each category of surgical instruments.
[0148] As shown in Figure 16(b), it is the P-R curve of the seven categories of surgical instruments; it can be seen from the figure that all seven surgical instruments show excellent detection performance.
[0149] As Figure 17 (a) shows the change curve of the bounding box loss function (a) obtained based on the training set using the SIoU loss function; as Figure 17 (b) shows the change curve of the object detection loss function (b) obtained based on the training set using the SIoU loss function; as Figure 17 (c) shows the change curve of the classification loss function (c) obtained based on the training set using the SIoU loss function; these three pictures respectively represent the mean value of the bounding box loss function, the mean value of the object detection loss function, and the mean value of the classification loss function of the training set; specifically, in this application, Box_loss uses the SIoU loss function, Box is speculated to be the mean value of the SIoU loss function, and the smaller the box, the more accurate it is; Objectness_loss is speculated to be the mean value of the object detection loss, and the smaller it is, the more accurate the object detection is; Classification_loss is speculated to be the mean value of the classification loss, and the smaller it is, the more accurate the classification is.
[0150] As shown in Figure 18 (a) is the change curve of the bounding box loss function obtained by using the SIoU loss function based on the validation set (a); As shown in Figure 18 (b) is the change curve of the object detection loss function obtained by using the SIoU loss function based on the validation set (b); As shown in Figure 18 (c) is the change curve of the classification loss function obtained by using the SIoU loss function based on the validation set (c); These three pictures respectively represent the mean values of the bounding box loss function, object detection loss function, and classification loss function of the training set; Specifically, val Box_loss is used for the bounding box loss of the validation set; val Objectness_loss is used for the mean value of the object detection loss of the validation set; val classification_loss is used for the mean value of the classification loss of the validation set.
[0151] In addition, due to the existence of interference factors such as instrument occlusion, instrument jitter, incomplete display, insufficient ambient light, or instrument reflection in the real surgical scenario, for the surgical instrument detection algorithm, whether it can accurately detect surgical instruments in the presence of interference factors is an important consideration index, which can help doctors more easily identify surgical instruments when encountering interference, reduce distraction, and focus more on surgical operations.
[0152] As shown in Fig. 19(a) is the detection result diagram of the presence of tissue and organ occlusion in the image; As shown in Fig. 19(b) is the detection result diagram of the presence of smoke occlusion in the image; As shown in Fig. 19(c) is the detection result diagram of the incomplete display of the instrument in the image; As shown in Fig. 19(d) is the detection result diagram of the instrument jitter in the image; As shown in Fig. 19(e) is the detection result diagram of the insufficient light in the image; As shown in Fig. 19(f) is the detection result diagram of the reflection in the image; The detection result diagrams obtained under the above different conditions can fully show that the method of the present application can still correctly detect surgical instruments under the interference, and has a high detection accuracy, that is, it shows the high adaptability of the method in the complex and changeable surgical environment, which is of great significance for real clinical surgeries and the medical field.
[0153] As shown in Table 4 below are the detection results of the surgical instrument pose detection method of the present application and other surgical instrument detection methods after training with the same data set.
[0154] Table 4:
[0155]
[0156] It can be clearly seen from the comparison results shown in Table 4 that, compared with other algorithms, the image detection effect achieved by the method of this application has reached the highest level, with an accuracy of 97.3%, a recall rate of 97.8%, and mAP and FPS reaching 97.9% and 133 frames per second respectively.
[0157] In summary, compared with the previous surgical instrument detection methods, the highest level has been achieved in all evaluation indicators, improving the detection accuracy while ensuring very high real-time performance. This depends on the excellent network structure of the selected YOLOv5, and more importantly, the several improvement strategies proposed in this paper play a crucial role. It not only reaches a very high level in surgical instrument detection, but also, due to the addition of the PSPNet semantic segmentation head, this method can also perform the segmentation of surgical instruments, combining detection and segmentation to complete the detection of the pose of surgical instruments.
[0158] Among them, the results of this method in surgical instrument segmentation are shown in Table 5 below.
[0159] Table 5:
[0160] Image detection network MIoU(%) Dice(%) YOLOv5 + GAM + SiLU + SIoU + PSPNet segmentation head 85.7 86.6
[0161] It can be seen from the detection results in Table 5 that the algorithm invented in this application has reached a relatively high level in the two important indicators of MIoU and Dice coefficient, and can accurately detect when the pose of surgical instruments changes (such as the opening and closing of surgical instruments), proving the feasibility and efficiency of the method of combining object detection and semantic segmentation for pose detection.
[0162] As Figure 20 shown, in step S5 of the method of this application, the surgically acquired instrument image collected in real time is input into the image detection network trained in step 4, and a marked image accurately detecting the pose of the surgical instrument can be output in real time. Among them, part (a) in the figure is the pose diagram of the grasper and the sample bag; part (b) in the figure is the pose diagram of the surgical hook; part (c) in the figure is the pose diagram of the irrigator and the grasper; part (d) in the figure is the pose diagram of the bipolar forceps; part (e) in the figure is the pose diagram of the surgical scissors and the grasper; part (f) in the figure is the pose diagram of the surgical forceps; part (g) in the figure is the pose diagram of the grasper; part (h) in the figure is the pose diagram of the irrigator and the grasper.
[0163] In summary, this application proposes a surgical instrument pose detection algorithm. Based on the original YOLOv5 algorithm, the GAM attention mechanism is introduced into its backbone network to improve the network's feature extraction ability. The SiLU activation function is adopted in the BottlenneckCSP module to avoid gradient explosion. The unstable training situation is alleviated. The SIoU loss function is adopted in the prediction part to reduce the degree of freedom of regression, accelerate the network convergence and further improve the regression accuracy. Finally, combined with the PSPNet semantic segmentation head, the surgical instrument pose detection is realized.
[0164] The above are only the preferred specific embodiments of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A real-time detection method for the pose of surgical instruments based on images, characterized in that, The steps are as follows: Step 1: Construct an image detection network suitable for real-time pose recognition of surgical instruments, and the structure includes: (1) An image input end; (2) A backbone network, which consists of a first CBS module, a second CBS module, a first CSP1_1 module, a third CBS module, a CSP1_2 module, a fourth CBS module, a CSP1_3 module, a fifth CBS module, a second CSP1_1 module, an SPPF module, and a GAM module connected in sequence; (3) A neck network, which consists of a first CBS module, a first upsampling module, a first Concat module, a first CSP2_1 module, a second CBS module, a second upsampling module, a second Concat module, a second CSP2_1 module, a third CBS module, a third Concat module, a third CSP2_1 module, a fourth CBS module, a fourth Concat module, and a fourth CSP2_1 module connected in sequence; Among them, the input end of the first Concat module of the neck network is also connected to the output end of the CSP1_3 module of the backbone network, the input end of the second Concat module of the neck network is also connected to the output end of the first CSP1_2 module of the backbone network, the input end of the third Concat module of the neck network is also connected to the output end of the second CBS module of the neck network, and the input end of the fourth Concat module of the neck network is also connected to the output end of the first CBS module of the neck network; (4) A semantic segmentation network PSPNet, whose input ends are respectively connected to the output ends of the first CBS module of the backbone network, the second CBS module of the backbone network, and the second CSP2_1 module of the neck network, and whose output ends are respectively connected to the output ends of the first Conv module, the second Conv module, and the third Conv module of the prediction network; (5) A prediction network, which consists of a first Conv module, a second Conv module, and a third Conv module; The loss function in the prediction network adopts the SIoU loss function; Among them, the CBS module consists of a convolutional layer, a BN layer, and a SiLU activation function connected in sequence; The output end of the GAM module is connected to the input end of the first CBS module of the neck network; Step 2: Construct a surgical instrument image dataset for training the image detection network, which consists of a detection dataset and a segmentation dataset; Among them, the detection dataset is obtained by annotating the tips of surgical instruments with the smallest rectangular boxes based on images containing different types of surgical instruments, and each annotated image corresponds to generated image detection information, including the surgical instrument category, the four vertex coordinates of the rectangular box, the length and width of the rectangular box; The segmentation dataset is obtained by annotating the contour of the tip of the surgical instrument with a closed contour box based on images containing different types of surgical instruments, and each annotated image corresponds to generated image segmentation information, including the surgical instrument category and the multi-point coordinates accurately depicting the contour box; Step 3: Perform data augmentation processing on the surgical instrument image dataset obtained in Step 2 to expand the data volume of the surgical instrument image dataset; Step 4: Use the surgical instrument image dataset obtained in Step 3 to train the image detection network constructed in Step 1, including: Step 4.1: Randomly divide the surgical instrument image dataset into a training set and a test set; Step 4.2: Input the training set into the image detection network. First, use the labeled images in the detection dataset as the input and the image detection information as the output to train and validate the image detection network. Then, use the labeled images in the segmentation dataset as the input and the image segmentation information as the output to train and validate the image detection network; Step 5: Input the surgically instrument images collected in real time into the image detection network trained in Step 4, and the labeled images accurately detecting the postures of the surgical instruments can be output in real time.
2. The real-time detection method for the pose of a surgical instrument based on an image according to claim 1, wherein The specific implementation steps of Step 2 are as follows: Step 2.1: Collect images of different surgical instruments during the operation to obtain the original image set; among them, the number of images in the original image set should meet the requirements of network training, and the proportion of the number of images of different surgical instruments in the total number of images is the same or approximately the same; Step 2.2: Use the image annotation software labelme to manually annotate each image obtained in Step 1.1 in turn to obtain the surgical instrument image dataset; the surgical instrument image dataset consists of a detection dataset and a segmentation dataset, where (1) The construction method of the detection dataset is as follows: Based on the original image set obtained in Step 2.1, each image is marked as follows: ① If more than half of the tip part of the surgical instrument in the image is shown, then circle the visible part of the image in the form of the smallest rectangle; ② If less than half of the tip part of the surgical instrument in the image is shown, then do not mark it; Generate the image detection information corresponding to the images marked with the rectangle, including the surgical instrument category, the four vertex coordinates of the rectangle, the length and width of the rectangle; the unmarked images do not generate image detection information; Finally, the marked images marked based on the original image set and the image detection information corresponding to the marked images one by one constitute the detection dataset; (2) The construction method of the segmentation dataset is as follows: Based on the original image set obtained in Step 2.1, each image is marked as follows: ① If more than half of the tip part of the surgical instrument in the image is shown, then draw a closed contour box along the contour of the visible part of the image; ② If less than half of the tip part of the surgical instrument in the image is shown, then do not mark it; Then a segmentation dataset composed of marked images is obtained; Generate the image segmentation information corresponding to the images marked with the rectangle, including the surgical instrument category and the multi-point coordinates that can accurately depict the contour box; the unmarked images do not generate image segmentation information; Finally, the marked images marked based on the original image set and the image detection information corresponding to the marked images one by one constitute the detection dataset.
3. The real-time detection method for the posture of a surgical instrument based on an image according to claim 1, characterized in that, The specific implementation manner of Step 3 is as follows: Step 3.1: Use the Mosaic data augmentation method to perform data augmentation on the surgical instrument image dataset obtained in Step 3; Step 3.2: Using the Albumentations image enhancement library, perform data enhancement processing on the surgical instrument image dataset obtained in Step 3 by using at least one of the methods of random color, contrast transformation, brightness transformation, noise perturbation, random cropping, scaling, translation, and rotation in sequence.
Citation Information
Patent Citations
Real-time target detection method based on full-channel space positioning attention mechanism
CN115082737A
Image-based surgical instrument position detection method
CN115690096A