A real-time surgical instrument detection method based on endoscopic images

By improving the YOLOV3 network, adopting the hybrid attention mechanism of MobileNetV3 and CBAM, combining PANet feature fusion, and using DIoU strategy, the real-time and accuracy of surgical instrument detection in endoscopic images are solved, efficient detection of surgical instruments, and improving surgical quality evaluation and safety.

CN116342869BActive Publication Date: 2025-07-08CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310309370.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2025-07-08
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

The prior art has problems with insufficient real-time and accuracy in endoscopic images, especially in the case of blood stains, smoke, occlusion and artifacts, which lead to missed and missed detection.

Method used

The MobileNetV3 network is used to replace the DarkNet53 feature extractor, combined with the CBAM hybrid attention mechanism and the PANet feature fusion module, and DIoU is used as a non-maximum suppression criterion to build a lightweight object detection model to improve detection speed and accuracy.

Benefits of technology

Real-time and accurate detection of surgical instruments in endoscopic images is achieved, the recall rate and detection speed of the model are improved, the number of falsely deleted target boxes is reduced, and the accuracy of surgical safety and quality evaluation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342869B_ABST
    Figure CN116342869B_ABST
Patent Text Reader

Abstract

The present invention discloses a real-time surgical instrument detection method based on endoscopic images. The DarkNet53 feature extractor in the benchmark YOLOV3 network is replaced with a MobileNetV3 network, and the SE channel attention mechanism in the MobileNetV3 network is replaced with a CBAM hybrid attention mechanism. Based on the above improvement methods, in the feature extraction module, the model parameters and computational amount can be effectively reduced, and the use of the hybrid attention module can make the key feature expressions more prominent and suppress useless features; in the feature fusion module, the improved PANet can effectively alleviate the problem of information loss while improving the network training speed; in the classification module of the network, a decay function associated with the overlap degree of the target box is set, which can alleviate the situation of misdeletion caused by overlap, reduce the number of misdeleted target boxes, and improve the model recall rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of clinical and surgical robotics, and particularly to a real-time surgical instrument detection method based on endoscopic images. Background Art

[0002] Minimally invasive surgery refers to a type of surgery in which specific surgical instruments are inserted into a patient's body through small incisions and related operations are performed. Minimally invasive surgery is favored by patients because of its small trauma area, low pain index, and fast recovery speed. Currently, laparoscopic surgery is the most widely used type of surgery in minimally invasive surgery. In the computer-aided intervention system for minimally invasive surgery, the application of computers can be reflected in two aspects: on the one hand, during the surgery, specifically including surgical robot-assisted physician operation and surgical operation abnormality reminder to improve the quality and safety of surgery; on the other hand, after the surgery, specifically including surgical quality assessment, prognosis analysis, education and training for young doctors, etc., which has extremely high value.

[0003] Surgical instruments, as the most prominent and iconic target objects in endoscopic images, are considered the main factors that dominate the entire surgical process. Surgical instrument detection is the primary or main task in the application of computer technology before and after surgery. In recent years, surgical instrument detection based on computer vision has received great attention. During the surgery, the detection based on surgical instruments helps to identify surgical actions and surgical stages, and thus can better realize the automation and intelligence of surgical robots. In addition, due to factors such as limited endoscopic vision and improper doctor operation, surgical instruments may damage human tissues. The real-time and accurate detection of surgical instruments will improve surgical safety. After the surgery, surgical instrument detection is an important part of the surgical quality assessment of the surgeon, and is of great significance for reducing postoperative complications caused by poor surgical quality of doctors.

[0004] In actual endoscopic images, surgical instrument detection faces the following challenges: (1) The complexity or unclarity of endoscopic images caused by blood stains, smoke generated instantaneously during electrocoagulation and cutting, unstable light illumination conditions, etc. during the surgical process; (2) The problem of missed detection caused by mutual occlusion between surgical instruments and diseased tissues; (3) The deformation of surgical instruments and the high artifacts increase the difficulty of recognition. In addition, when performing robot-assisted surgery and anomaly detection tasks, the cognition and processing of endoscopic images often require real-time requirements. Therefore, researching a surgical instrument detection method that simultaneously meets real-time and high-precision requirements has important value in the medical field. Summary of the Invention

[0005] The purpose of the present invention is to provide a real-time surgical instrument detection method based on endoscopic images to solve the problems raised in the above background art.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A real-time surgical instrument detection method based on endoscopic images, comprising the following steps:

[0008] Step 1: Replace the DarkNet53 feature extractor in the baseline YOLOV3 network with the MobileNetV3 network, and use depthwise separable convolutions to replace the standard convolution operations to construct a lightweight object detection model;

[0009] Step 2: Replace the SE channel attention mechanism in the MobileNetV3 network with the CBAM hybrid attention mechanism to fuse the spatial relationship and channel dimension information in the input image;

[0010] Step 3: Replace the FPN feature fusion module in the YOLOV3 network with the PANet network to fuse features at different levels in a bottom-up and top-down dual-path parallel manner, alleviating the problem of feature loss. In addition, use depthwise separable convolutions to replace the standard convolution operations in PANet, while alleviating the loss of shallow features, reducing the number of model parameters, and improving the network training speed;

[0011] Step 4: To solve the problem of low recall rate caused by the YOLOV3 network, use DIoU as the non-maximum suppression criterion to effectively solve the problem of missed detection caused by target occlusion.

[0012] Based on the above improvement methods, the model parameters and computational complexity can be effectively reduced in the feature extraction module, and the use of the hybrid attention module can more prominently express key features and suppress useless features; in the feature fusion module, the improved PANet can effectively alleviate the problem of information loss while improving the network training speed; in the classification module of the network, a decay function associated with the overlap degree of the target box is set, which can alleviate the misdeletion situation caused by overlap, reduce the number of misdeleted target boxes, and improve the model recall rate.

[0013] Compared with the prior art, the beneficial effects of the present invention are:

[0014] The present invention realizes a real-time surgical instrument detection method for endoscopic images. Through the lightweight network framework constructed by the present invention, real-time and accurate detection of surgical instruments in endoscopic images can be achieved. Compared with the widely used YOLOV3 network, it can achieve improvements in speed and accuracy. Therefore, the present invention has important application values in aspects such as real-time understanding of the surgical scene, guiding the operation of surgical robots for auxiliary operations, and evaluating the surgical quality in the surgical scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a framework diagram of the object detection network constructed by the present invention.

[0016] Figure 2 It is a schematic diagram of the hybrid attention mechanism. Specific implementation manners

[0017] The technical solutions in the embodiments of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0018] Embodiment 1, as Figure 1 shown, a real-time surgical instrument detection method based on endoscopic images replaces the original DarkNet53 network with Figure 1 the MobileNetV3 network shown in the figure to construct a lightweight network architecture. The MobileNetV3 network replaces the standard convolution operation with depthwise separable convolution. The depthwise separable convolution replaces the convolution operation with depth convolution and point convolution. Table 1 summarizes the time complexity and space complexity analysis of the two convolution methods. Among them, the size of the input image is n×n×c, c is the number of input channels, c′ is the number of output channels, the convolution kernel size of the standard convolution is k×k, and the size of the finally generated feature map is n′×n′×c′. It can be seen that the computational amount of the depthwise separable convolution is much smaller than that of the standard convolution, and the time complexity and space complexity of the network model are reduced.

[0019] Table 1: Analysis table of time complexity and space complexity of standard convolution and depthwise separable convolution;

[0020]

[0021]

[0022] The SE channel attention mechanism in the above MobileNetV3 network is replaced by the CBAM hybrid attention mechanism. The origin of the attention mechanism can be described as a mechanism for reallocating weights. Different weights correspond to the degree of attention to the corresponding positions of the image, and different degrees of attention correspond to the degree of association between the features extracted by the network and the target. Therefore, adopting the attention mechanism in the network can make the image pay more attention to the target areas related to the task and improve the model detection performance. The convolution operation is essentially to fuse spatial and channel information to extract features. The SE attention mechanism belongs to the channel attention mechanism and lacks the ability to focus on key features in the spatial relationship, while the spatial attention mechanism can solve the problem of "where is the important information". The hybrid attention mechanism CBAM attention mechanism adopted in the present invention is as Figure 2As shown, this mechanism can simultaneously fuse spatial relationships and channel dimension information, and simultaneously focus on target information in both channels and space, and jointly calculate the target feature weights.

[0023] In the YOLOV3 network, an FPN feature fusion structure is used to fuse features of different scales. The FPN fusion network fuses shallow features in a bottom-up manner. When transmitting features in a deeper network, it is easy to cause the problem of loss of shallow features. As Figure 1 As shown in the middle part, the present invention uses a PANet network to replace the FPN fusion structure, extracts feature information in a dual-path parallel manner of bottom-up and top-down, effectively retains features at different levels, and finally fuses these features. In addition, the number of layers of the PANet network is smaller, alleviating the problem of feature loss caused by the transmission of shallow features.

[0024] Many studies have shown that the YOLOV3 network has a low recall rate in the overall detection effect. This is because when performing feature classification, the NMS strategy will cause the phenomenon of misdeleting occluded targets due to a large overlapping area of the targets. The process of the NMS method for processing candidate boxes can be summarized as follows:

[0025] (1) Put all the generated candidate boxes into the set E, and set the confidence of each candidate box as s i ;

[0026] (2) Put the candidate box M with the highest confidence in the set E into the final set G;

[0027] (3) Calculate the IoU value between the remaining candidate boxes in E and M. If the calculated value is greater than the set threshold N t , then set its confidence to 0 and remove it from E;

[0028] (4) Repeat the above steps until E is empty.

[0029] However, when there is an overlap between two targets, the low-confidence target will be misdeleted, resulting in missed detection by the model. To solve this problem, the present invention uses DIoU as a new non-maximum suppression strategy to replace the above-mentioned way of directly setting the confidence to 0 violently. This strategy can be described by the following formula:

[0030]

[0031]

[0032] Among them, ρ is the distance between the centers of two candidate boxes, and c is the distance between the diagonals of the minimum circumscribed rectangles of two candidate boxes. It can be seen from the above formula that if the overlapping area between two candidate boxes is larger, the DIoU value is larger, s iThe greater the attenuation, the greater the possibility of being suppressed, and vice versa, the closer it is to the actual situation. Compared with the original NMS method, the improved strategy does not directly set the object with a large overlapping area to 0, effectively alleviating the situation of misdeleting the bounding box due to overlapping.

[0033] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention.

[0034] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A real-time surgical instrument detection method based on endoscopic images, characterized in that, It includes the following steps: Step 1: Replace the DarkNet53 feature extractor in the baseline YOLOV3 network with the MobileNetV3 network, and use depthwise separable convolutions to replace standard convolution operations to build a lightweight object detection model; Step 2: Replace the SE channel attention mechanism in the MobileNetV3 network with the CBAM hybrid attention mechanism to fuse the spatial relationship and channel dimension information in the input image; Step 3: Replace the FPN feature fusion module in the YOLOV3 network with the PANet network to fuse features at different levels in a bottom-up and top-down dual-path parallel manner to alleviate the problem of feature loss. In addition, use depthwise separable convolutions to replace the standard convolution operations in PANet, which can alleviate the loss of shallow features while reducing the number of model parameters and improving the network training speed; Step 4: Use DIoU as the non-maximum suppression criterion to solve the problem of missed detection caused by object occlusion.

2. A real-time surgical instrument detection method based on endoscopic images according to claim 1, characterized in that, The baseline YOLOV3 network consists of three parts: feature extraction, feature fusion, and classification.

3. A real-time surgical instrument detection method based on endoscopic images according to claim 1, characterized in that, The MobileNetV3 network uses depthwise separable convolutions to replace standard convolution operations, and depthwise separable convolutions replace convolution operations with depthwise convolutions and pointwise convolutions.

4. A real-time surgical instrument detection method based on endoscopic images according to claim 3, characterized in that The essence of the attention mechanism is a mechanism for reallocating weights. Different weights correspond to the degree of attention to the corresponding positions in the image, and different degrees of attention correspond to the correlation between the features extracted by the network and the target.

5. A real-time surgical instrument detection method based on endoscopic images according to claim 1, characterized in that, The DIoU is described by the following formula: ; ; where ρ is the distance between the centers of two candidate bounding boxes, and c is the distance between the diagonals of the minimum enclosing rectangles of the two candidate bounding boxes.

Citation Information

Patent Citations

  • Vibration damper detection method based on deep learning algorithm

    CN112906654A

  • Large-breadth SAR image ship target detection and identification method based on fine segmentation

    CN113409325A