Surgical procedure and surgical instrument combined detection method, medium and electronic device
By combining image feature extraction and temporal information fusion networks in a dual-task detection method, the problem of low accuracy in surgical procedure and instrument identification is solved, achieving high-precision and real-time detection of surgical procedures and instruments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies have low accuracy in identifying surgical stages and instruments, and often employ a single-task identification method, resulting in insufficient precision.
A video-based dual-task joint detection method is adopted, which combines image domain and time domain features through an image feature extraction network and a temporal information fusion network, and uses a strip attention causal mask and constraint network to jointly detect surgical procedures and surgical instruments.
It improves the accuracy and robustness of surgical procedures and instrument detection, and can predict the surgical process in real time, with high accuracy and real-time performance.
Smart Images

Figure CN115187908B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a video processing method, in particular to a video-based surgical procedure and surgical instrument joint detection method, medium and electronic device. BACKGROUND
[0002] Surgical video analysis aims to evaluate and track the surgical process in real time, allowing surgeons to track detailed information of various events occurring in the operating room (OR) to improve the quality and safety of surgical operations. In the process of automatic surgical procedure analysis, surgical stage recognition and surgical instrument detection are the basis for understanding various intraoperative states in the surgical process. By recognizing the surgical stage and surgical instruments being performed through video, various relevant information can be provided to the doctor to optimize preoperative planning, intraoperative guidance and postoperative evaluation. However, the prior art usually adopts a single task recognition method to recognize the surgical stage and surgical instruments separately, which has low accuracy. SUMMARY
[0003] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a surgical procedure and surgical instrument joint detection method, medium and electronic device, which solves the problem of low accuracy of surgical stage and surgical instrument recognition in the prior art.
[0004] To achieve the above-mentioned purpose and other related purposes, the first aspect of the present application provides a video-based surgical procedure and surgical instrument joint detection method, which comprises: acquiring a surgical video; processing the surgical video by using an image feature extraction network to obtain the image domain features of the surgical video; and fusing the image domain features with the time domain features of the surgical video by using a time domain information fusion network and detecting the surgical procedure and surgical instrument.
[0005] In an embodiment of the first aspect, for a frame of surgical image in the surgical video: acquiring the image domain features of the frame of surgical image comprises processing the frame of surgical image by using the image feature extraction network to obtain the image domain features of the frame of surgical image; and detecting the surgical procedure and surgical instrument of the frame of surgical image comprises processing the image domain features of the frame of surgical image and the image domain features of several previous frames of surgical image by using the time domain information fusion network to realize the fusion of the image domain features with the time domain features of the frame of surgical image and detect the surgical procedure and surgical instrument of the frame of surgical image.
[0006] In an embodiment of the first aspect, the time domain information fusion network adopts a bar attention causal mask.
[0007] In an embodiment of the first aspect, the training method of the image feature extraction network comprises: constructing the image feature extraction network; obtaining a training video, the training video comprising a plurality of training images; preprocessing the training images; and training the image feature extraction network using the preprocessed training images.
[0008] In an embodiment of the first aspect, the training method of the image feature extraction network further comprises: obtaining a test image; processing the test image using a center crop method; and testing the performance of the feature extraction network using the test image.
[0009] In an embodiment of the first aspect, the training method of the time domain information fusion network comprises: constructing the time domain information fusion network and configuring a constraint network for the time domain information fusion network; obtaining training data; training the time domain information fusion network using the training data, and using the constraint network to constrain the distance between the prediction result of the time domain information fusion network and the gold standard during the training process.
[0010] In an embodiment of the first aspect, using the constraint network to constrain the distance between the prediction result of the time domain information fusion network and the gold standard comprises: using the constraint network to encode the prediction result of the time domain information fusion network and the gold standard into a plurality of Gaussian distributions, and constraining the distance between the prediction result of the time domain information fusion network and the gold standard to be minimum.
[0011] In an embodiment of the first aspect, before obtaining the image domain features of the surgical video, the joint detection method further comprises: preprocessing the surgical video.
[0012] The second aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the joint detection method of any one of the first aspect of the present application.
[0013] The third aspect of the present application provides an electronic device, comprising: a memory storing a computer program; a processor communicatively connected to the memory, and invoking the computer program to execute the joint detection method of any one of the first aspect of the present application.
[0014] As described above, the surgical procedure and surgical instrument joint detection method, medium and electronic device provided in one or more embodiments of the present application have the following beneficial effects:
[0015] The joint detection method simultaneously recognizes surgical instruments and surgical procedures using a double task, and in this process, the correlation between surgical procedures and surgical instruments can be fully considered, so that the method has high accuracy.
[0016] In addition, the joint detection method utilizes a two-stage network to perform feature extraction and fusion from the image domain and the time domain respectively, which can further improve the accuracy of surgical procedure and surgical instrument detection, and the recognition result has good robustness.
[0017] Furthermore, the time domain information fusion network used in the joint detection method can use a bar-shaped attention causal mask, which prevents the later video images from being used to predict the surgical procedure and surgical instrument of the previous video images, and thus can be applied to real-time prediction.
[0018] Further, the information fusion network used in the joint detection method can use a constraint mechanism during training, which is conducive to further improving the accuracy of surgical procedure and surgical instrument detection. BRIEF DESCRIPTION OF DRAWINGS
[0019] Figure 1 A flowchart showing the joint detection method of the present application in a specific embodiment.
[0020] Figure 2A A partial structure diagram of the time domain information fusion network in the joint detection method of the present application in a specific embodiment.
[0021] Figure 2B A schematic diagram of the attention causal mask in the joint detection method of the present application in a specific embodiment.
[0022] Figure 3A A flowchart showing the training method of the image feature extraction network in the joint detection method of the present application in a specific embodiment.
[0023] Figure 3B A flowchart showing the test method of the image feature extraction network in the joint detection method of the present application in a specific embodiment.
[0024] Figure 4 A flowchart showing the training method of the time domain information fusion network in the joint detection method of the present application in a specific embodiment.
[0025] Figure 5 A structure diagram of the electronic device of the present application in a specific embodiment.
[0026] ELEMENT NUMBER EXPLANATION
[0027] 500 electronic device
[0028] 510 memory
[0029] 520 processor
[0030] 530 display
[0031] S11-S13 steps
[0032] S31a-S34a steps
[0033] S31b-S33b steps
[0034] S41-S43 steps DETAILED DESCRIPTION
[0035] Other advantages and benefits of the present application will become apparent to those skilled in the art upon consideration of the disclosure or can be learned by practice of the application. The application can be realized and achieved by means of the structures and combinations of features described in this specification and claims. Various modifications and changes can be made thereto without departing from the spirit and scope of the application. It is to be understood that the following examples and features thereof can be combined with each other, if not in conflict.
[0036] It is to be understood that the drawings shown in the following examples are only schematic and are non-limiting examples. In the drawings, the size, the relative sizes, and the proportions of the parts shown are not necessarily to scale and, specifically, the dimensions of the parts, of the relative sizes, and the proportions are not necessarily to scale and, specifically, the dimensions of the parts, of the drawings serving merely to illustrate concepts. Like numbers refer to like or similar elements throughout the several views.
[0037] By recognizing the ongoing surgical stage and surgical instruments from surgical videos, a variety of relevant information can be provided to the surgeon to optimize preoperative planning, intraoperative guidance, and postoperative evaluation. However, the prior art usually adopts a single-task recognition method to recognize the surgical stage and surgical instruments respectively, which has a low accuracy.
[0038] To at least address the above problems, the present application provides a surgical procedure and surgical instrument joint detection method. The joint detection method adopts a dual-task method to simultaneously recognize the surgical instruments and the surgical procedure, and in this process, the correlation between the surgical procedure and the surgical instruments can be fully considered, thus having a high accuracy.
[0039] Next, the joint detection method provided by the present application will be described in detail by specific examples in conjunction with the drawings.
[0040] Figure 1 The flowchart shown is the joint detection method in an embodiment of the present application. As shown in the figure, the joint detection method includes the following steps: Figure 1As shown, the joint detection method provided in this embodiment includes the following steps S11 to S13.
[0041] S11, Acquire surgical video, wherein the surgical video includes multiple frames of surgical images. This surgical video can be obtained, for example, by capturing images of the surgical site using an image acquisition device.
[0042] S12, an image feature extraction network is used to process the surgical video to obtain its image domain features. This image feature extraction network can be, for example, a trained deep learning network. Deep learning is a type of machine learning, its concept originating from research on artificial neural networks. Deep learning discovers distributed feature representations of data by combining low-level features to form more abstract high-level representations of attribute categories or features.
[0043] S13, a temporal information fusion network is used to fuse the image domain features and temporal features of the surgical video to detect the surgical procedure and surgical instruments. The temporal information fusion network can be, for example, another trained deep learning network. In step S13, the image domain features of the surgical video can be input into the temporal information fusion network for processing. During processing, the temporal information fusion network can incorporate the temporal features of the surgical video into the image domain features. After processing, the output of the temporal information fusion network includes the detection results of the surgical procedure and surgical instruments.
[0044] As described above, the joint detection method provided in this embodiment utilizes a dual-task approach to simultaneously identify surgical instruments and surgical procedures, thus fully considering the correlation between surgical procedures and surgical instruments, resulting in high accuracy. Furthermore, the joint detection method provided in this embodiment employs a two-stage network to extract and fuse features from the image domain and the time domain respectively, which can further improve the accuracy of surgical procedure and surgical instrument detection, and the detection results exhibit good robustness.
[0045] In one embodiment of the present invention, for a surgical image frame in a surgical video, obtaining the image domain features of the surgical image frame includes: processing the surgical image frame using an image feature extraction network to obtain the image domain features of the surgical image frame. Detecting the surgical procedure and surgical instruments in the surgical image frame includes: processing the image domain features of the surgical image frame and the image domain features of several previous surgical images using a temporal domain information fusion network to achieve the fusion of the image domain features and temporal domain features of the surgical image frame, thereby outputting the detection results of the surgical procedure and surgical instruments in the surgical image frame.
[0046] Specifically, the single-frame surgery image carries image domain features, and the continuous multiple frames of surgery images carry time domain features. Thus, the time domain information fusion network can obtain the time domain features according to the image domain features of the single-frame surgery image and the previous several frames of surgery images, and then fuse the time domain features with the image domain features, so as to improve the accuracy of the output surgery procedure and surgery instrument detection result.
[0047] Preferably, the joint detection method provided in the embodiment can realize real-time processing of the surgery video. Specifically, in step S11, the image feature extraction network can obtain the image domain features of the single-frame surgery image in real time, and the time domain information fusion network can output the surgery procedure and the surgery instrument of the single-frame surgery image in real time according to the image domain features of the single-frame surgery image and the previous several frames of surgery images. Through real-time processing of the surgery video, the current surgery procedure and the surgery instrument can be obtained in time, which has good real-time performance.
[0048] In an embodiment of the present application, the time domain information fusion network comprises a plurality of cascaded Transformer layers. Figure 2A The structure diagram of two cascaded Transformer layers is shown. Figure 2A As shown in the figure, each Transformer layer comprises a strip attention causal mask layer, an addition & normalization layer (Add & Norm), a linear layer (FNN) and another addition & normalization layer connected in sequence. Figure 2B The strip attention causal mask in the embodiment is shown. The strip attention causal mask makes the time domain information fusion network only trace back to the video information of the previous limited frames at the current time, wherein the number of traced frames is determined by the bandwidth of the strip attention causal mask. Since the surgery procedure and the surgery instrument are only related to the surgery images in the previous period of time, compared with the mask that traces back to all previous frames of information, the strip attention causal mask is beneficial to improve the prediction accuracy.
[0049] In an embodiment of the present application, before obtaining the image domain features of the surgery video, the joint detection method further comprises: pre-processing the surgery video.
[0050] Optionally, the pre-processing of the surgery video comprises resampling, removing black edges and / or normalization and the like. Through the pre-processing of the surgery video, the data amount of the surgery video can be reduced, and the processing efficiency can be improved.
[0051] Figure 3A The flowchart of the training method of the image feature extraction network in an embodiment of the present application is shown. Figure 3A As shown in the figure, the training method of the image feature extraction network in the embodiment comprises the following steps S31a to S34a.
[0052] S31a, constructing an initial model of the image feature extraction network.
[0053] S32a, obtaining a training video, the training video comprising a plurality of training images.
[0054] S33a, pre-processing the training images. Optionally, the pre-processing of the training images comprises shuffling the temporal order of the training images and / or augmenting the training images. The augmenting of the training images can be achieved by randomly performing any one or more of the following operations on the training images: random cropping, color shift, horizontal flipping, random rotation.
[0055] S34a, training the image feature extraction network using the pre-processed training images. During the training process, a weighted cross-entropy loss function and a binary cross-entropy loss function can be used to train the procedure recognition task and the instrument detection task, respectively.
[0056] Referring to Figure 3B , the training method of the image feature extraction network in the embodiment can further comprise the following steps S31b to S33b.
[0057] S31b, obtaining a test image, which can be a plurality of images in one or more test videos.
[0058] S32b, processing the test image using a center cropping method. Specifically, during the surgery video acquisition process, important surgical behaviors, such as the operation of surgical instruments on organs, are often placed in the center of the image. In order to retain as much important action information as possible, the center cropping method can be used to process the test image in step S32b, that is, the center point of the test image is taken as the starting point, and a certain number of pixel points are retained above, below, left and right for cropping.
[0059] S33b, testing the performance of the feature extraction network using the test image.
[0060] After the image feature extraction network is trained using the above steps S31a-S34a and S31b-S33b, the image feature extraction network can process an input surgery image and output the image domain features of the surgery image.
[0061] Figure 4 A training method flowchart of a time domain information fusion network in an embodiment of the present application is shown. As Figure 4 shown, the training method of the time domain information fusion network in the embodiment comprises the following steps S41 to S43.
[0062] S41, construct a time domain information fusion network, and configure a constraint network for the time domain information fusion network. The constraint network is cascaded after the time domain information fusion network.
[0063] S42, obtain training data, which includes image domain features of multiple frames of training images and corresponding gold standards. The image domain features of the training images can be obtained by an image feature extraction network. The gold standards corresponding to the training images include annotation results of surgical procedures and surgical instruments in the training images, and can be obtained by manual annotation or the like.
[0064] S43, train the time domain information fusion network by using the training data, and constrain the distance between the prediction result of the time domain information fusion network and the gold standard by using the constraint network during the training process. Specifically, during the training process, the time domain information fusion network outputs TxC1 surgical procedure prediction labels and TxC2 surgical instrument prediction labels, and the concatenation of the two can obtain a prediction set with a size of Tx(C1+C2). The prediction set and the gold standard are input into the constraint network together, so that the constraint network encodes the prediction set and the gold standard into multiple Gaussian distributions and constrains the distance between them. Wherein, T, C1 and C2 are positive integers, T represents the total number of frames of training images, C1 represents the number of categories of surgical procedures, and C2 represents the number of categories of surgical instruments.
[0065] Optionally, constraining the distance between the prediction result of the time domain information fusion network and the gold standard by using the constraint network includes: encoding the prediction result of the time domain information fusion network and the gold standard into multiple Gaussian distributions by using the constraint network, and constraining the distance between the prediction result of the time domain information fusion network and the gold standard to be minimum. For example, a KL divergence loss function can be used to constrain the distance between the corresponding Gaussian distributions to be minimum.
[0066] Optionally, before step S43, the training method of the time domain information fusion network can further include: pre-processing the training data, for example, performing augmentation processing on the training data, and / or stretching the training data in the time domain. Wherein, stretching the training data in the time domain means randomly stretching the size of the image domain features of the training images from TxD to TxD', which can be realized by linear interpolation or the like. Wherein, D and D' can be configured according to actual requirements.
[0067] It should be noted that the embodiment introduces the constraint network in the training phase of the time domain information fusion network, and does not use the constraint network in the testing phase and actual application phase of the time domain information fusion network.
[0068] In some embodiments, the Cholec80 cholecystectomy dataset is utilized to train and verify the image feature extraction network and the time domain information fusion network. Among them, the test set verification result surgical procedure recognition accuracy is 93.12±4.71%, the precision is 89.25±5.49%, the recall rate is 90.10±5.45%, the Jaccard index is 81.11±7.62%, and the F1 index is 88.26±5.37%, and the surgical instrument recognition target detection index (Mean Average Precision, mAP) is 95.15±3.87%, all of which exceed the current benchmark algorithm. It can be seen that the training method and the test method provided by the application have good performance.
[0069] Based on the above description of the joint detection method, the application further provides a computer readable storage medium having a computer program stored thereon. The computer program is executed by a processor to implement the joint detection method of any one of the above embodiments.
[0070] In the application, any combination of one or more storage media can be used. The storage medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may, for example, be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of computer readable storage media include: electrical connections having one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component.
[0071] The application further provides an electronic device. Figure 5 The structure schematic diagram of the electronic device 500 in an embodiment of the application is shown. As shown in the figure, the electronic device 500 in the embodiment includes a memory 510 and a processor 520. Figure 5
[0072] The memory 510 is used for storing a computer program; preferably, the memory 510 includes: ROM, RAM, disk, U disk, memory card or optical disk and various media that can store program codes.
[0073] In particular, the memory 510 can include a computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. The electronic device 500 can further include other removable / non-removable, volatile / non-volatile computer system storage media. The memory 510 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the application.
[0074] The processor 520 is connected with the memory 510, and is configured to execute the computer program stored in the memory 510, so that the electronic device 500 executes the joint detection method shown in the figure. Figure 1 The processor 520 is connected with the memory 510, and is configured to execute the computer program stored in the memory 510, so that the electronic device 500 executes the joint detection method shown in the figure.
[0075] Preferably, the processor 520 can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; and can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0076] Preferably, the electronic device 500 in the embodiment can further include a display 530. The display 530 is connected in communication with the memory 510 and the processor 520, and is configured to display the relevant GUI interaction interface of the joint detection method.
[0077] The protection scope of the joint detection method described in the application is not limited to the execution order of the steps listed in the embodiment, and any scheme realized by increasing, reducing or replacing the steps of the prior art according to the principle of the application is included in the protection scope of the application.
[0078] In summary, the application provides a joint detection method. The joint detection method can simultaneously recognize surgical instruments and surgical procedures by using a double task, and the correlation between the surgical procedures and the surgical instruments can be fully considered in the process, so that the accuracy is higher. In addition, the joint detection method uses a two-stage network to extract and fuse features from the image domain and the time domain, which can further improve the accuracy of surgical procedure and surgical instrument detection, and the recognition result has good robustness. Furthermore, the time domain information fusion network used in the joint detection method can use a bar attention causal mask, which makes the later video images not used to predict the surgical procedures and surgical instruments of the previous video images, so it can be applied to real-time prediction. Further, the information fusion network used in the joint detection method can use a constraint mechanism in the training process, which is conducive to further improving the accuracy of surgical procedure and surgical instrument detection. Therefore, the application effectively overcomes the shortcomings in the prior art and has high industrial utilization value.
[0079] The above embodiments only exemplarily illustrate the principles and effects of the application, and are not used to limit the application. Any person skilled in the art can modify or change the above embodiments without departing from the spirit and scope of the application. Therefore, all equivalent modifications or changes completed by those skilled in the art without departing from the spirit and technical idea disclosed by the application should be covered by the claims of the application.
Claims
1. A video-based method for joint detection of surgical procedures and surgical instruments, characterized in that, The joint detection method includes: Obtain the surgical video; The surgical video is processed using an image feature extraction network to obtain image domain features of the surgical video; The surgical video's image domain features and time domain features are fused using a time-domain information fusion network to detect the surgical procedure and surgical instruments; The training method of the time-domain information fusion network includes: constructing the time-domain information fusion network and configuring a constraint network for the time-domain information fusion network; acquiring training data; training the time-domain information fusion network using the training data, and using the constraint network to constrain the distance between the prediction results of the time-domain information fusion network and the gold standard during the training process; wherein, during the training process, the time-domain information fusion network outputs surgical procedure prediction labels of T×C1 and surgical instrument prediction labels of T×C2, which are concatenated to obtain a prediction set of size T×(C1+C2), and the prediction set and the gold standard are input into the constraint network together, so that the constraint network encodes the prediction set and the gold standard into multiple Gaussian distributions and constrains the distance between them, where T, C1, and C2 are all positive integers, T represents the total number of training images, C1 represents the number of surgical procedure categories, and C2 represents the number of surgical instrument categories.
2. The combined detection method according to claim 1, characterized in that, For a single frame of surgical image in the surgical video: Obtaining the image domain features of the surgical image frame includes: processing the surgical image frame using the image feature extraction network to obtain the image domain features of the surgical image frame; The detection of the surgical procedure and surgical instruments in the surgical image frame includes: using a time-domain information fusion network to process the image domain features of the surgical image frame and the image domain features of several previous surgical images, so as to achieve the fusion of the image domain features and time domain features of the surgical image frame and detect the surgical procedure and surgical instruments in the surgical image frame.
3. The combined detection method according to claim 2, characterized in that, The time-domain information fusion network employs a striped attention causal mask.
4. The combined detection method according to claim 1, characterized in that, The training method for the image feature extraction network includes: Construct the image feature extraction network; Acquire a training video, which includes multiple frames of training images; The training images are preprocessed; The image feature extraction network is trained using the preprocessed training images.
5. The combined detection method according to claim 4, characterized in that, The training method for the image feature extraction network also includes: Obtain the test image; The test image was processed using a center-cropping method; The performance of the feature extraction network was tested using the test images.
6. The combined detection method according to claim 1, characterized in that, Using the constraint network to constrain the distance between the prediction results of the time-domain information fusion network and the gold standard includes: The constraint network is used to encode the prediction results of the time-domain information fusion network and the gold standard into multiple Gaussian distributions, and the distance between the prediction results of the time-domain information fusion network and the gold standard is constrained to be minimized.
7. The combined detection method according to any one of claims 1 to 6, characterized in that, Before acquiring the image domain features of the surgical video, the joint detection method further includes preprocessing the surgical video.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the joint detection method according to any one of claims 1 to 7.
9. An electronic device, characterized in that, The electronic device includes: A memory that stores a computer program; The processor, which is communicatively connected to the memory, executes the joint detection method according to any one of claims 1 to 7 when it invokes the computer program.
Citation Information
Patent Citations
Real-time action recognition method based on spatial-temporal feature fusion
CN113052059A