Operation triple recognition method combining frame-level and video-level activation guidance

By combining frame-level and video-level class activation-guided surgical triplet recognition methods, and utilizing backbone networks and temporal relationship modeling, the problem of existing technologies failing to effectively utilize global temporal information and multi-scale temporal features is solved, thereby improving the accuracy and robustness of surgical behavior recognition.

CN120953864APending Publication Date: 2025-11-14WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510809178.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing surgical triplet recognition methods fail to effectively utilize global temporal information and multi-scale temporal features in surgical videos, resulting in insufficient generalization ability and low recognition accuracy in complex surgical scenarios.

Method used

A surgical triplet recognition method combining frame-level and video-level class activation guidance is proposed. Frame-level spatial features are extracted through a backbone network, and component class activation mapping and temporal relationship modeling are performed. Spatiotemporal feature guidance is then applied by combining video-level spatial feature sequences to achieve frame-level and video-level relationship modeling.

Benefits of technology

It improves the accuracy and robustness of surgical video behavior recognition, effectively utilizes spatiotemporal features for surgical behavior triplet recognition, reduces confusion caused by similar behaviors and potential task conflicts between components and triples, and improves the recognition accuracy in complex surgical scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953864A_ABST
    Figure CN120953864A_ABST
Patent Text Reader

Abstract

The invention discloses an operation triple recognition method combining frame-level and video-level activation guidance, and belongs to the technical field of target recognition. Comprising the following steps: inputting each piece of video frame data into a backbone network to obtain a frame-level spatial feature, and inputting the frame-level spatial feature into a first independent branch network to obtain a frame-level component class activation mapping sequence; operation triple features are obtained based on the frame-level spatial features and the component class activation mapping sequence, a target video is input into a backbone network to obtain a video-level spatial feature sequence, and sequential relation modeling is carried out to obtain an operation triple spatial-temporal feature sequence; inputting the operation triple spatial-temporal feature sequence into a second independent branch network to obtain a video-level component class activation mapping sequence; and performing video-level component relation modeling based on the video-level component class activation mapping sequence and the operation triple spatial-temporal feature sequence to obtain an operation triple recognition result of each video frame. According to the method, the accuracy of operation video behavior recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of target recognition technology, and in particular relates to a surgical triplet recognition method that combines frame-level and video-level class activation guidance. Background Technology

[0002] With the rapid development of medical technology, video surgical behavior recognition is becoming increasingly important in the field of surgical data science. This technology not only provides comprehensive information for surgical procedure analysis and surgical scenario understanding, but also supports safety alerts and computer-aided decision-making within the operating room, providing surgeons with intraoperative context-aware support and improving surgical safety. Therefore, automated surgical behavior recognition has become a crucial problem that urgently needs to be solved.

[0003] Existing surgical triplet recognition methods are mainly based on multi-task learning frameworks. Some methods use spatial features from a single frame or several adjacent frames to generate class activation maps of triplet components, thereby obtaining triplet predictions. However, these methods only model triplet component relationships at the frame level, failing to effectively utilize global temporal information and multi-scale temporal features in surgical videos, and neglecting video-level triplet relationship modeling. This limits the model's generalization ability in complex surgical scenarios. Furthermore, some methods address the uneven class distribution problem through semantic enhancement strategies, but still rely on a simple multi-task framework, simultaneously predicting triples and components. When there is a conflict between component class prediction and overall triplet prediction, the model's learning performance is affected, weakening its ability to model triplet component relationships.

[0004] Existing technologies neglect the relationships between components and between components and triples in surgical behavior triples, lacking global temporal information and multi-scale features. As a result, they suffer from low triple recognition accuracy and insufficient generalization ability in surgical triple recognition. Summary of the Invention

[0005] This application aims to address at least one of the technical problems existing in the prior art. To this end, this application proposes a surgical triplet recognition method that combines frame-level and video-level class activation guidance, which improves the accuracy of surgical video behavior recognition.

[0006] In a first aspect, this application provides a surgical triplet recognition method that combines frame-level and video-level class activation guidance, the method comprising: The target video is acquired, and multiple video frame data are obtained based on the target video. Each video frame data is input into the backbone network to obtain frame-level spatial features. The frame-level spatial features are then input into the first independent branch network to obtain a frame-level component class activation mapping sequence. Based on the frame-level spatial features and frame-level component class activation mapping sequence, frame-level component relationship modeling guided by spatial features is performed to obtain surgical triplet features, wherein the surgical triplet includes three components: instrument, action, and target. The backbone network is trained based on the surgical triplet features. The target video is input into the trained backbone network to obtain a video-level spatial feature sequence. Based on the video-level spatial feature sequence, temporal relationship modeling is performed to obtain the surgical triplet spatiotemporal feature sequence. The surgical triplet spatiotemporal feature sequence is then input into the second independent branch network to obtain the video-level component class activation mapping sequence. Based on the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence, spatiotemporal feature-guided video-level component relationship modeling is performed to obtain the surgical triplet recognition result for each video frame in the target video.

[0007] According to one embodiment of this application, the step of inputting each video frame data into the backbone network to obtain frame-level spatial features, and inputting the frame-level spatial features into the first independent branch network to obtain a frame-level component class activation mapping sequence includes: Each video frame data is input into the swin-transformer network for spatial feature extraction to obtain frame-level spatial features; The frame-level spatial features are input into a first independent branch network for dimensionality transformation to obtain a frame-level component class activation mapping sequence. The first independent branch network includes a first convolutional network, a second convolutional network, and a third convolutional network.

[0008] According to one embodiment of this application, the step of performing frame-level component relationship modeling guided by spatial features based on the frame-level spatial features and frame-level component class activation mapping sequence to obtain surgical triplet features includes: The frame-level spatial features and frame-level component class activation mapping sequence are mapped to a unified feature space to generate a serialized component token set, and the component token set is stacked into a frame-level relation sequence. The frame-level relation sequence is subjected to multi-head self-attention operation to obtain a frame-level relation representation sequence, which includes a self-attention representation token of the component token set, a first cross-attention representation token between components, and a first cross-attention representation token between a component and a surgical triplet. The frame-level relation representation sequence is input into the hybrid expert model. Different linear modulation experts are dynamically activated according to the routing network to process different relation representation tokens in the frame-level relation representation sequence. Spatial feature-guided frame-level component relation modeling is performed to obtain surgical triplet features.

[0009] According to one embodiment of this application, the step of dynamically activating different linear modulation expert processing frame-level relation representation tokens in the routing network to perform spatial feature-guided frame-level component relation modeling to obtain surgical triplet features includes: Based on each relation representation token in the frame-level relation representation sequence, the softmax scores of different linear modulation experts are calculated and sorted according to the routing network; The linear modulation experts ranked first and second in the softmax score are selected as the target experts. Frame-level component relationship modeling guided by target experts and computation of affine transformation operators are performed. Based on the affine transformation operator, the features of different relation representation tokens in the frame-level relation representation sequence are calculated to obtain the surgical triplet features.

[0010] According to one embodiment of this application, training the backbone network based on the surgical triplet features, and inputting the target video into the trained backbone network to obtain a video-level spatial feature sequence, includes: The component prediction result is obtained based on the frame-level component class activation mapping sequence, and the component classification loss function and contrastive learning loss function are obtained based on the comparison between the component prediction result and the true label. Based on the surgical triplet features, a frame-level initial surgical triplet identification result is obtained, and a frame-level triplet classification loss function is obtained based on the comparison between the initial surgical triplet identification result and the true label. The first loss function is obtained based on the component classification loss function, the contrastive learning loss function, and the frame-level triplet classification loss function. Based on the surgical triplet features, the backbone network is trained according to the first loss function to obtain a trained backbone network. The target video is input into the trained backbone network to obtain frame-level features, which are then stacked along the time dimension to obtain a video-level spatial feature sequence.

[0011] According to one embodiment of this application, the step of performing temporal relationship modeling based on the video-level spatial feature sequence to obtain a surgical triplet spatiotemporal feature sequence, and inputting the surgical triplet spatiotemporal feature sequence into a second independent branch network to obtain a video-level component class activation mapping sequence includes: The video-level spatial feature sequence is input into a temporal convolutional network to model temporal relationships, resulting in a surgical triplet spatiotemporal feature sequence. The spatiotemporal feature sequence of the surgical triplet is input into the second independent branch network for dimensionality transformation to obtain a video-level component class activation mapping sequence. The second independent branch network includes a four-convolutional network, a fifth-convolutional network, and a sixth-convolutional network.

[0012] According to one embodiment of this application, the video-level component relationship modeling guided by spatiotemporal features based on the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence to obtain the surgical triplet recognition result for each video frame in the target video includes: The video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence are concatenated along the feature dimension to obtain the concatenated feature. The splicing features are subjected to multi-head self-attention operation to obtain a video-level surgical triplet relation representation, which includes a self-attention representation token of the component, a second cross-attention representation token between components, and a second cross-attention representation token between the component and the surgical triplet. The video-level surgical triplet relation representation is input into the hybrid expert model. Different linear modulation experts are dynamically activated according to the routing network to process different relation representation tokens in the video-level surgical triplet relation representation. Spatiotemporal feature-guided video-level component relation modeling is performed to obtain the surgical triplet video-level features. Based on the video-level features of the surgical triplet, the surgical triplet identification result is obtained for each video frame in the target video.

[0013] Secondly, this application provides a surgical triplet recognition device that combines frame-level and video-level class activation guidance, the device comprising: The acquisition module is used to acquire the target video, obtain multiple video frame data based on the target video, input each video frame data into the backbone network to obtain frame-level spatial features, and input the frame-level spatial features into the first independent branch network to obtain the frame-level component class activation mapping sequence. The first processing module is used to perform frame-level component relationship modeling guided by spatial features based on the frame-level spatial features and the frame-level component class activation mapping sequence to obtain surgical triplet features, wherein the surgical triplet includes three components: instrument, action, and target. The second processing module is used to train the backbone network based on the surgical triplet features, input the target video into the trained backbone network, and obtain a video-level spatial feature sequence. The third processing module is used to perform temporal relationship modeling based on the video-level spatial feature sequence to obtain the surgical triplet spatiotemporal feature sequence, and input the surgical triplet spatiotemporal feature sequence into the second independent branch network to obtain the video-level component class activation mapping sequence. The recognition module is used to perform spatiotemporally feature-guided video-level component relationship modeling based on the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence, so as to obtain the surgical triplet recognition result for each video frame in the target video.

[0014] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the surgical triplet recognition method combining frame-level and video-level class activation guidance as described in the first aspect above.

[0015] Fourthly, this application provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the surgical triplet recognition method combining frame-level and video-level class activation guidance as described in the first aspect above.

[0016] Fifthly, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the surgical triplet recognition method combining frame-level and video-level class activation guidance as described in the first aspect.

[0017] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the surgical triplet recognition method combining frame-level and video-level class activation guidance as described in the first aspect above.

[0018] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application.

[0019] The present invention provides a surgical triplet recognition method that combines frame-level and video-level class activation guidance, which has the following advantages over the prior art: (1) This invention acquires target video and extracts multiple video frame data, inputs each video frame into the backbone network to obtain frame-level spatial features, combines frame-level spatial features and component class activation mapping sequence to perform frame-level component relationship modeling guided by spatial features, performs temporal relationship modeling based on video-level spatial feature sequence, and combines surgical triplet spatiotemporal feature sequence and video-level component class activation mapping sequence to perform spatiotemporal feature-guided component relationship modeling, thus realizing frame-level and video-level relationship modeling, and simultaneously realizing feature refinement of triplet recognition in the spatiotemporal dimension. It realizes both frame-level spatial feature extraction and video-level multi-scale temporal feature modeling of surgical behavior, effectively utilizes spatiotemporal features for surgical behavior triplet recognition, and improves the recognition accuracy and robustness of triplets in complex surgical scenarios.

[0020] (2) This invention effectively extracts the spatiotemporal features of surgical triples by inputting the video-level spatial feature sequence into a temporal convolutional network to model temporal relationships. By extracting the spatiotemporal features of surgical triples and generating component class activation mapping sequences, component class activation mapping can be used to achieve effective learning of components and triples, fully learn the intrinsic relationship between components and triples, reduce confusion caused by similar behaviors and possible task conflicts between components and triples, and improve the accuracy of triple recognition.

[0021] (3) This invention maps frame-level spatial features and frame-level component class activation mapping sequences to a unified feature space, generating a serialized set of component tokens. A frame-level relation representation sequence is obtained through multi-head self-attention operations, effectively extracting frame-level surgical triplet features. Furthermore, the frame-level relation representation sequence is input into a hybrid expert model. By dynamically activating different linear modulation experts to process different relation representation tokens, the complex and diverse component features are effectively integrated. The model fully learns the relationships between components and between components and triples, achieving frame-level component relation modeling guided by spatial features. This reduces confusion caused by similar behaviors and potential task conflicts between components and triples, improving the accuracy of triplet recognition. Attached Figure Description

[0022] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is one of the flowcharts of the surgical triplet recognition method combining frame-level and video-level class activation guidance provided in the embodiments of this application; Figure 2 This is the second flowchart of the surgical triplet recognition method combining frame-level and video-level class activation guidance provided in the embodiments of this application; Figure 3This is a schematic diagram of the structure of the surgical triplet recognition device that combines frame-level and video-level class activation guidance provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0024] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0025] The following description, in conjunction with the accompanying drawings, details the surgical triplet recognition method, device, electronic device, and readable storage medium that combine frame-level and video-level class activation guidance provided in this application, through specific embodiments and application scenarios.

[0026] Among them, the surgical triplet recognition method that combines frame-level and video-level class activation guidance can be applied to the terminal, specifically executed by the hardware or software in the terminal.

[0027] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0028] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.

[0029] The surgical triplet recognition method combining frame-level and video-level class activation guidance provided in this application embodiment can be executed by an electronic device or a functional module or entity within an electronic device capable of implementing the surgical triplet recognition method combining frame-level and video-level class activation guidance. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The following description uses an electronic device as an example to illustrate the surgical triplet recognition method combining frame-level and video-level class activation guidance provided in this application embodiment.

[0030] Figure 1 This is one of the flowcharts illustrating the surgical triplet recognition method combining frame-level and video-level class activation guidance provided in this application embodiment, such as... Figure 1 As shown, the surgical triplet recognition method combining frame-level and video-level class activation guidance includes steps 110, 120, 130, 140, and 150.

[0031] Step 110: Obtain the target video, obtain multiple video frame data based on the target video, input each video frame data into the backbone network to obtain frame-level spatial features, and input the frame-level spatial features into the first independent branch network to obtain a frame-level component class activation mapping sequence. For example, the endoscopic video of a laparoscopic cholecystectomy is obtained as the target video, the target video is split into multiple video frame data, and given a complete surgical video, the surgical behavior triplet appearing in each frame of the video is identified.

[0032] In some embodiments, the step of inputting each video frame data into the backbone network to obtain frame-level spatial features, and inputting the frame-level spatial features into the first independent branch network to obtain a frame-level component class activation mapping sequence, includes: Each video frame data is input into the swin-transformer network for spatial feature extraction to obtain frame-level spatial features; The frame-level spatial features are input into a first independent branch network for dimensionality transformation to obtain a frame-level component class activation mapping sequence. The first independent branch network includes a first convolutional network, a second convolutional network, and a third convolutional network.

[0033] For example, each video frame is cropped to 384*384 pixels and input into the Swin-transformer network to obtain frame-level spatial features. Frame-level spatial features include triplet information. Triplets consist of three components: instrument, action, and target. Instruments have 6 categories, specifically including grippers, bipolar electrocoagulation forceps, monopolar hooks, scissors, clamps, and flushers; actions have 10 categories, specifically including grasping, pulling, separating, coagulating, clamping, cutting, suctioning, flushing, packing, and no action; targets have 15 categories, specifically including gallbladder, gallbladder bed, cystic duct, cystic artery, gallbladder pedicle, blood vessel, fluid, abdominal wall cavity, liver, adhesions, omentum, peritoneum, intestine, specimen bag, and no target. There are a total of 100 categories in triplets, which is an effective combination of instrument, action, and target.

[0034] The frame-level spatial features are input into the first independent branch network for dimensionality transformation to obtain the frame-level component class activation mapping sequence. The first independent branch network includes a first convolutional network, a second convolutional network, and a third convolutional network, which are used to enhance the frame-level spatial features to obtain the class activation mapping of triple components.

[0035] Taking six categories of instrument components as an example, the first independent branch network outputs an activation map of six channels, with each channel corresponding to a category. This map represents the response intensity of the feature at that location in each of the six categories, helping to identify which instrument the feature belongs to.

[0036] In this embodiment, by inputting each video frame data into the swin-transformer network for spatial feature extraction, more accurate frame-level spatial features can be obtained. These features are then further input into the first independent branch network for dimensional transformation to obtain a frame-level component class activation mapping sequence, thereby improving the robustness and accuracy of subsequent surgical triplet recognition.

[0037] Step 120: Based on the frame-level spatial features and the frame-level component class activation mapping sequence, perform frame-level component relationship modeling guided by spatial features to obtain surgical triplet features, wherein the surgical triplet includes three components: instrument, action, and target. In some embodiments, the frame-level component relationship modeling guided by spatial features based on the frame-level spatial features and the frame-level component class activation mapping sequence to obtain surgical triplet features includes: The frame-level spatial features and frame-level component class activation mapping sequence are mapped to a unified feature space to generate a serialized component token set, and the component token set is stacked into a frame-level relation sequence. The frame-level relation sequence is subjected to multi-head self-attention operation to obtain a frame-level relation representation sequence, which includes a self-attention representation token of the component token set, a first cross-attention representation token between components, and a first cross-attention representation token between a component and a surgical triplet. The frame-level relation representation sequence is input into the hybrid expert model. Different linear modulation experts are dynamically activated according to the routing network to process different relation representation tokens in the frame-level relation representation sequence. Spatial feature-guided frame-level component relation modeling is performed to obtain surgical triplet features.

[0038] It's easy to understand that class activation mapping refers to a feature map containing only features of specific components. A hybrid expert model (MoE) is a deep learning architecture that integrates multiple specialized sub-models (experts) based on a dynamic routing mechanism. The class activation mapping-guided hybrid expert model takes component class activation maps and triplet features as input, combines the dynamic learning capabilities of MoE (Mixture of Experts) to dynamically model features of different components, optimize triplet prediction, and simultaneously employs linear feature modulation techniques to design a linearly modulated expert structure that can handle diverse component features, generated from the input. Two operators, implemented by linear functions, and based on the Hadamard product (using...). An affine transformation is performed so that the mapping can be computed based on the input, to accommodate inputs from different components. The computation formula is shown below: in, As a feature modulation expert, As input to the hybrid expert model, It is an affine transformation operator generated from the input. All are linear transformation networks.

[0039] Given a frame-level component class activation mapping sequence and frame-level spatial features h, where , Representing instruments, actions, and targets respectively, the intrinsic relationships between components and between components and triples are learned, mapped to a unified feature space, generating serialized token representations, and stacking these tokens into a frame-level relation sequence. D is the dimension of the feature space, and n is the sequence length, which is related to the mapping method.

[0040] Relation sequence Perform multi-head self-attention operations to obtain self-attention representations of component tokens, cross-attention between components, and cross-attention representations between components and triples in the feature space, resulting in a sequence of relation representations. , The sequence length is given.

[0041] In some embodiments, the step of dynamically activating different linear modulation expert processing frame-level relation representation tokens in the routing network to perform spatial feature-guided frame-level component relation modeling to obtain surgical triplet features includes: Based on each relation representation token in the frame-level relation representation sequence, the softmax scores of different linear modulation experts are calculated and sorted according to the routing network; The linear modulation experts ranked first and second in the softmax score are selected as the target experts. Frame-level component relationship modeling guided by target experts and computation of affine transformation operators are performed. Based on the affine transformation operator, the features of different relation representation tokens in the frame-level relation representation sequence are calculated to obtain the surgical triplet features.

[0042] Furthermore, the relation is represented as a sequence. The input is fed into a hybrid expert model, which then adjusts the input based on the routing network. Dynamically activate different linear modulation experts This process handles relation representations formed by different components, performs relation modeling, and outputs triplet features after relation modeling. The specific process of triplet relation modeling at the full frame level is as follows: According to the sequence Each token in Calculate the softmax score for each token, and select the two experts with the highest scores to assign a linear modulation expert to the token. The linear modulation expert first calculates the operator for performing affine transformation on the features based on the input token. and Then, operators are used to compute on the input tokens. For the sequence... Each token in Calculate and preserve the original relative positions, and finally aggregate the outputs of the activated experts. This generates the final decoded output. , Given the sequence length, the output of CAMOE has the same dimension as the input, as expressed in the following formula: , Where S is the processed frame-level surgical triplet feature sequence. For the nth surgical triplet, Let tokens represent sequences of relations. CAMoE is a module that processes component features and models the relationships between components and between components and triples. R is a routing network that selects appropriate experts. As a feature modulation expert, As input to the hybrid expert model, It is an affine transformation operator generated from the input. All are linear transformation networks.

[0043] In this embodiment, by dynamically activating different linear modulation experts based on the routing network to process different relation representation tokens in the frame-level relation representation sequence, more accurate spatial feature-guided frame-level component relation modeling can be achieved. By calculating and ranking the softmax scores of different linear modulation experts, the first and second ranked experts are selected as target experts. Then, affine transformations and feature calculations of the surgical triplet features are performed through the target experts. The frame-level class activation mapping-guided hybrid expert model guides spatial feature learning through class activation mapping, enabling effective representation of component features. Utilizing the multi-expert collaborative processing of different components in the hybrid expert model, component identification is more effective, improving the recognition effect and accuracy of triples.

[0044] In this embodiment, by mapping frame-level spatial features and frame-level component class activation mapping sequences to a unified feature space, a serialized set of component tokens is generated. A frame-level relation representation sequence is then obtained through multi-head self-attention operations, effectively extracting frame-level surgical triplet features. Furthermore, the frame-level relation representation sequence is input into a hybrid expert model. By dynamically activating different linear modulation experts to process different relation representation tokens, the complex and diverse component features are effectively integrated. This fully learns the relationships between components and between components and triples, achieving frame-level component relation modeling guided by spatial features. This reduces confusion caused by similar behaviors and potential task conflicts between components and triples, improving the accuracy of triplet recognition.

[0045] Step 130: Train the backbone network based on the surgical triplet features, input the target video into the trained backbone network, and obtain a video-level spatial feature sequence; In some embodiments, training the backbone network based on the surgical triplet features, and inputting the target video into the trained backbone network to obtain a video-level spatial feature sequence, includes: The component prediction result is obtained based on the frame-level component class activation mapping sequence, and the component classification loss function and contrastive learning loss function are obtained based on the comparison between the component prediction result and the true label. Based on the surgical triplet features, a frame-level initial surgical triplet identification result is obtained, and a frame-level triplet classification loss function is obtained based on the comparison between the initial surgical triplet identification result and the true label. The first loss function is obtained based on the component classification loss function, the contrastive learning loss function, and the frame-level triplet classification loss function. Based on the surgical triplet features, the backbone network is trained according to the first loss function to obtain a trained backbone network. The target video is input into the trained backbone network to obtain frame-level features, which are then stacked along the time dimension to obtain a video-level spatial feature sequence.

[0046] What's easy to understand is that the frame-level triplet relationship modeling results are used to generate frame-level triplet recognition results, which are then compared with the true labels to generate a frame-level triplet classification loss function. The component prediction results are obtained by class activation mapping and compared with the true labels to generate component classification loss. It also utilizes contrastive learning to enhance imbalanced categories, generating contrastive learning loss. This yields enhanced spatial features.

[0047] Frame-level triplet classification loss function Based on the baseline, the component classification loss was determined experimentally. and contrastive learning loss Setting parameters The first loss function is obtained and used to optimize the parameters of the backbone network, resulting in a robust and efficient backbone network. The formula for calculating the first loss function is shown below: in, For the first loss function, For frame-level triple classification loss function, To compare learning loss, For component classification loss, As the first weighting factor, It is the second weighting factor.

[0048] Finally, the complete video is input frame by frame into the trained backbone network, transforming the original video into a video-level feature sequence. Each frame of the target video first passes through the backbone network to obtain frame-level features, which are then stacked along the time dimension to obtain a spatial feature sequence that is consistent with the original video in terms of order and length.

[0049] In this embodiment, a first loss function is constructed by combining the component classification loss function, the contrastive learning loss function, and the frame-level triplet classification loss function, thereby optimizing the training process of the backbone network and improving its feature extraction capability. By inputting the target video into the trained backbone network and stacking it along the time dimension to obtain a video-level spatial feature sequence, more accurate frame-level features can be obtained effectively, improving the recognition accuracy of surgical triplet features and enhancing the backbone network's ability to process complex video data.

[0050] Step 140: Perform temporal relationship modeling based on the video-level spatial feature sequence to obtain the surgical triplet spatiotemporal feature sequence, and input the surgical triplet spatiotemporal feature sequence into the second independent branch network to obtain the video-level component class activation mapping sequence. In some embodiments, the step of performing temporal relationship modeling based on the video-level spatial feature sequence to obtain a surgical triplet spatiotemporal feature sequence, and inputting the surgical triplet spatiotemporal feature sequence into a second independent branch network to obtain a video-level component class activation mapping sequence includes: The video-level spatial feature sequence is input into a temporal convolutional network to model temporal relationships, resulting in a surgical triplet spatiotemporal feature sequence. The spatiotemporal feature sequence of the surgical triplet is input into the second independent branch network for dimensionality transformation to obtain a video-level component class activation mapping sequence. The second independent branch network includes a four-convolutional network, a fifth-convolutional network, and a sixth-convolutional network.

[0051] The process is straightforward: video-level spatial feature sequences are input into a temporal convolutional network for temporal relationship modeling. The resulting triplet spatiotemporal feature sequence contains rich spatiotemporal information. Each frame in the spatiotemporal feature sequence input into the temporal convolutional network incorporates features from all frames of the complete video, meaning each frame contains its associated temporal features, enabling frame-by-frame prediction. This triplet spatiotemporal feature sequence is then transformed through a second independent branch network to obtain video-level activation mapping sequences for each component class.

[0052] In this embodiment, by inputting the video-level spatial feature sequence into a temporal convolutional network to model temporal relationships, the spatiotemporal features of surgical triples are effectively extracted. By extracting the spatiotemporal features of surgical triples and generating component class activation mapping sequences, component class activation mapping can be used to achieve effective learning of components and triples, fully learn the intrinsic relationship between components and triples, reduce confusion caused by similar behaviors and possible task conflicts between components and triples, and improve the accuracy of triple recognition.

[0053] Step 150: Based on the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence, perform spatiotemporal feature-guided video-level component relationship modeling to obtain the surgical triplet recognition result for each video frame in the target video.

[0054] Figure 2 This is the second flowchart illustrating the surgical triplet recognition method combining frame-level and video-level class activation guidance provided in this application embodiment, as shown below. Figure 2 As shown, Figure 2 The upper part performs frame-level processing, taking each frame of the image as input, mapping the images into a sequence, and inputting it into the hybrid expert model to implement class activation guidance operations. Figure 2 The lower part represents video-level processing, which processes the video into a spatial feature sequence as input. This sequence is then processed through a hybrid expert model to perform video-level class activation guidance. Based on the video-level component class activation mapping sequence and the spatiotemporal feature sequence of surgical triples, spatiotemporal feature-guided video-level component relationship modeling is performed to obtain the surgical triple identification result for each video frame in the target video. The surgical triple identification result represents the category of the triples in each frame of the target video. For example, no triple corresponds to category 0, and <hook, hook, organ> corresponds to category 1.

[0055] According to the surgical triplet recognition method combining frame-level and video-level class activation guidance provided in the embodiments of this application, the target video is acquired and multiple video frame data are extracted. Each video frame is input into the backbone network to obtain frame-level spatial features. The frame-level spatial features and component class activation mapping sequence are combined to perform frame-level component relationship modeling guided by spatial features. Temporal relationship modeling is performed based on video-level spatial feature sequence. The spatiotemporal feature sequence of surgical triplet and video-level component class activation mapping sequence are combined to perform spatiotemporal feature-guided component relationship modeling. This realizes frame-level and video-level relationship modeling and simultaneously achieves feature refinement for triplet recognition in the spatiotemporal dimension. It realizes both frame-level spatial feature extraction and video-level multi-scale temporal feature modeling of surgical behavior. It effectively utilizes spatiotemporal features for surgical behavior triplet recognition, improving the recognition accuracy and robustness of triplets in complex surgical scenarios.

[0056] In some embodiments, the video-level component relationship modeling guided by spatiotemporal features based on the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence to obtain the surgical triplet recognition result for each video frame in the target video includes: The video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence are concatenated along the feature dimension to obtain the concatenated feature. The splicing features are subjected to multi-head self-attention operation to obtain a video-level surgical triplet relation representation, which includes a self-attention representation token of the component, a second cross-attention representation token between components, and a second cross-attention representation token between the component and the surgical triplet. The video-level surgical triplet relation representation is input into the hybrid expert model. Different linear modulation experts are dynamically activated according to the routing network to process different relation representation tokens in the video-level surgical triplet relation representation. Spatiotemporal feature-guided video-level component relation modeling is performed to obtain the surgical triplet video-level features. Based on the video-level features of the surgical triplet, the surgical triplet identification result is obtained for each video frame in the target video.

[0057] The straightforward approach is to concatenate the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence along the feature dimension to obtain concatenated features. Multi-head self-attention is then applied to these concatenated features to obtain representations of component self-attention, cross-attention between components, and cross-attention between components and triples in the feature space, thus yielding the video-level surgical triplet relationship representation. , This represents the number of video frames.

[0058] Furthermore, the video-level surgical ternary relationship is represented... The input is fed into a hybrid expert model, which dynamically activates different linear modulation experts to process the relational representations formed by different components, thus providing a sequence. Each token in Assigning linear modulation experts effectively facilitates relation modeling. The linear modulation expert first computes operators for affine transformation of the features based on the input tokens. and Then, operators are used to compute on the input tokens. For the sequence... Each token in Calculate and preserve the original relative positions, and finally aggregate the outputs of the activated experts. This generates the final decoded output. , Given the sequence length, the output of CAMOE has the same dimension as the input, and the calculation formula is as follows: , Where ST is the surgical triplet video-level feature sequence with a length of n. For the nth surgical triplet video-level feature, For video-level surgery, triples are used to represent tokens; CAMoE is a proposed module for processing component features and modeling relationships between components and between components and triples; R is a routing network for selecting appropriate experts. As a feature modulation expert, As input to the hybrid expert model, It is an affine transformation operator generated from the input. All are linear transformation networks.

[0059] Based on the video-level features of surgical triplets, the surgical triplet recognition result for each video frame in the target video is obtained, and the calculation formula is as follows: in, The results of the surgical triplet identification are shown, with sigmoid as the activation function, ensuring that the value is non-negative. This is a function that takes the maximum value among the channels representing the category, corresponding to the most likely category.

[0060] It should be noted that by inputting the spatiotemporal feature sequence of triples into an independent classification network, the component class activation mapping sequence and component prediction results are obtained, resulting in a video-level component classification loss. Video-level triplet classification loss Experimentally, the video-level component classification loss was set. Weights relative to video-level triple classification loss The parameters are The total loss is obtained and used to optimize the overall model parameters, thereby achieving effective surgical behavior triplet recognition. The formula for calculating the total loss is as follows: in, For the total loss, For video-level triplet classification loss, For video-level component classification loss, It is the third weighting factor.

[0061] It is worth noting that the method is evaluated using the mean accuracy (mAP) of triplet recognition, in addition to triplet-level metrics. ), and also includes component-specific metrics ( , , ) and inter-component relationship metrics ( and Average precision is a core metric for measuring the detection or recognition performance of a model. It is used to comprehensively evaluate the model's performance at different confidence thresholds by calculating the area under the precision-recall curve. The formula for calculating total precision is shown below: Where p(r) represents the precision when the recall is r. For total accuracy.

[0062] In this embodiment, by concatenating the video-level component class activation mapping sequence and the spatiotemporal feature sequence of surgical triples along the feature dimension, a concatenated feature is obtained. Multi-head self-attention is then performed to obtain a video-level surgical triple relation representation. By inputting this surgical triple relation representation into a hybrid expert model, and combining it with a routing network to dynamically activate different linear modulation experts, different relation tokens in the video-level surgical triple relation representation are effectively handled, achieving spatiotemporal feature-guided video-level component relation modeling. Through linear feature modulation technology to dynamically integrate features of different components, and class activation mapping to guide the relation modeling technology between components and between components and triples, effective learning of component-triples is achieved using component class activation mapping. This fully learns the intrinsic connections between components and triples, reducing confusion caused by similar behaviors and potential task conflicts between components and triples, thus improving the accuracy of triple recognition. This approach can be applied to scenarios such as intelligent surgical navigation, robot-assisted surgery, and surgical scene understanding.

[0063] The surgical triplet recognition method combining frame-level and video-level class activation guidance provided in this application can be executed by a surgical triplet recognition device combining frame-level and video-level class activation guidance. This application uses the surgical triplet recognition device combining frame-level and video-level class activation guidance executing the surgical triplet recognition method as an example to illustrate the surgical triplet recognition device combining frame-level and video-level class activation guidance provided in this application.

[0064] This application also provides a surgical triplet recognition device that combines frame-level and video-level class activation guidance, such as... Figure 3 As shown, the surgical triplet recognition device combining frame-level and video-level class activation guidance includes: an acquisition module 310, a first processing module 320, a second processing module 330, a third processing module 340, and a recognition module 350.

[0065] The acquisition module 310 is used to acquire the target video, obtain multiple video frame data based on the target video, input each video frame data into the backbone network to obtain frame-level spatial features, and input the frame-level spatial features into the first independent branch network to obtain a frame-level component class activation mapping sequence. The first processing module 320 is used to perform frame-level component relationship modeling guided by spatial features based on the frame-level spatial features and the frame-level component class activation mapping sequence to obtain surgical triplet features, wherein the surgical triplet includes three components: instrument, action, and target. The second processing module 330 is used to train the backbone network based on the surgical triplet features, input the target video into the trained backbone network, and obtain a video-level spatial feature sequence. The third processing module 340 is used to perform temporal relationship modeling based on the video-level spatial feature sequence to obtain the surgical triplet spatiotemporal feature sequence, and input the surgical triplet spatiotemporal feature sequence into the second independent branch network to obtain the video-level component class activation mapping sequence. The recognition module 350 is used to perform spatiotemporal feature-guided video-level component relationship modeling based on the video-level component class activation mapping sequence and surgical triplet spatiotemporal feature sequence to obtain the surgical triplet recognition result for each video frame in the target video.

[0066] According to the surgical triplet recognition method combining frame-level and video-level class activation guidance provided in the embodiments of this application, the target video is acquired and multiple video frame data are extracted. Each video frame is input into the backbone network to obtain frame-level spatial features. The frame-level spatial features and component class activation mapping sequence are combined to perform frame-level component relationship modeling guided by spatial features. Temporal relationship modeling is performed based on video-level spatial feature sequence. The spatiotemporal feature sequence of surgical triplet and video-level component class activation mapping sequence are combined to perform spatiotemporal feature-guided component relationship modeling. This realizes frame-level and video-level relationship modeling and simultaneously achieves feature refinement for triplet recognition in the spatiotemporal dimension. It realizes both frame-level spatial feature extraction and video-level multi-scale temporal feature modeling of surgical behavior. It effectively utilizes spatiotemporal features for surgical behavior triplet recognition, improving the recognition accuracy and robustness of triplets in complex surgical scenarios.

[0067] The surgical triplet recognition device combining frame-level and video-level class activation guidance provided in this application embodiment can achieve… Figures 1 to 2 The various processes implemented in the embodiment of the surgical triplet recognition method that combines frame-level and video-level class activation guidance will not be described in detail here to avoid repetition.

[0068] In some embodiments, such as Figure 4As shown, this application embodiment also provides an electronic device 400, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the program is executed by the processor 401, it implements the various processes of the above-described surgical triplet recognition method embodiment combining frame-level and video-level class activation guidance, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0069] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0070] This application also provides a non-transitory computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described embodiment of the surgical triplet recognition method combining frame-level and video-level class activation guidance, and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0071] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0072] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described surgical triplet recognition method combining frame-level and video-level class activation guidance.

[0073] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0074] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described embodiment of the surgical triplet recognition method combining frame-level and video-level class activation guidance, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0075] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a device-level chip, device chip, chip device, or on-chip device chip, etc.

[0076] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0077] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the surgical triplet recognition method combining frame-level and video-level class activation guidance of the various embodiments of this application.

[0078] In the description of this application, "first feature" and "second feature" may include one or more of the features.

[0079] In the description of this application, "multiple" means two or more.

[0080] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0081] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0082] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

Claims

1. A surgical triplet recognition method combining frame-level and video-level class activation guidance, characterized in that, The method includes: The target video is acquired, and multiple video frame data are obtained based on the target video. Each video frame data is input into the backbone network to obtain frame-level spatial features. The frame-level spatial features are then input into the first independent branch network to obtain a frame-level component class activation mapping sequence. Based on the frame-level spatial features and frame-level component class activation mapping sequence, spatial feature-guided frame-level component relationship modeling is performed to obtain surgical triplet features, wherein the surgical triplet includes three components: instrument, action, and target. The backbone network is trained based on the surgical triplet features. The target video is input into the trained backbone network to obtain a video-level spatial feature sequence. Based on the video-level spatial feature sequence, temporal relationship modeling is performed to obtain the surgical triplet spatiotemporal feature sequence. The surgical triplet spatiotemporal feature sequence is then input into the second independent branch network to obtain the video-level component class activation mapping sequence. Based on the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence, spatiotemporal feature-guided video-level component relationship modeling is performed to obtain the surgical triplet recognition result for each video frame in the target video.

2. The surgical triplet recognition method combining frame-level and video-level class activation guidance according to claim 1, characterized in that, The process of inputting each video frame data into the backbone network to obtain frame-level spatial features, and then inputting the frame-level spatial features into the first independent branch network to obtain a frame-level component class activation mapping sequence, includes: Each video frame data is input into the win-transformer network for spatial feature extraction to obtain frame-level spatial features; The frame-level spatial features are input into a first independent branch network for dimensionality transformation to obtain a frame-level component class activation mapping sequence. The first independent branch network includes a first convolutional network, a second convolutional network, and a third convolutional network.

3. The surgical triplet recognition method combining frame-level and video-level class activation guidance according to claim 1, characterized in that, The frame-level component relationship modeling guided by spatial features based on the frame-level spatial features and frame-level component class activation mapping sequence yields surgical triplet features, including: The frame-level spatial features and frame-level component class activation mapping sequence are mapped to a unified feature space to generate a serialized component token set, and the component token set is stacked into a frame-level relation sequence. The frame-level relation sequence is subjected to multi-head self-attention operation to obtain a frame-level relation representation sequence, which includes a self-attention representation token of the component token set, a first cross-attention representation token between components, and a first cross-attention representation token between a component and a surgical triplet. The frame-level relation representation sequence is input into the hybrid expert model. Different linear modulation experts are dynamically activated according to the routing network to process different relation representation tokens in the frame-level relation representation sequence. Spatial feature-guided frame-level component relation modeling is performed to obtain surgical triplet features.

4. The surgical triplet recognition method combining frame-level and video-level class activation guidance according to claim 3, characterized in that, The process involves dynamically activating different linear modulation expert processing tokens in the frame-level relation representation sequence based on the routing network, performing spatial feature-guided frame-level component relation modeling, and obtaining surgical triplet features, including: Based on each relation representation token in the frame-level relation representation sequence, the softmax scores of different linear modulation experts are calculated and sorted according to the routing network; The linear modulation experts ranked first and second in the softmax score are selected as the target experts. Frame-level component relationship modeling guided by target experts and computation of affine transformation operators are performed. Based on the affine transformation operator, the features of different relation representation tokens in the frame-level relation representation sequence are calculated to obtain the surgical triplet features.

5. The surgical triplet recognition method combining frame-level and video-level class activation guidance according to claim 1, characterized in that, The process involves training the backbone network based on the surgical triplet features, inputting the target video into the trained backbone network, and obtaining a video-level spatial feature sequence, including: The component prediction result is obtained based on the frame-level component class activation mapping sequence, and the component classification loss function and contrastive learning loss function are obtained based on the comparison between the component prediction result and the true label. Based on the surgical triplet features, a frame-level initial surgical triplet identification result is obtained, and a frame-level triplet classification loss function is obtained based on the comparison between the initial surgical triplet identification result and the true label. The first loss function is obtained based on the component classification loss function, the contrastive learning loss function, and the frame-level triplet classification loss function. Based on the surgical triplet features, the backbone network is trained according to the first loss function to obtain a trained backbone network. The target video is input into the trained backbone network to obtain frame-level features, which are then stacked along the time dimension to obtain a video-level spatial feature sequence.

6. The surgical triplet recognition method combining frame-level and video-level class activation guidance according to claim 1, characterized in that, The temporal relationship modeling based on the video-level spatial feature sequence yields a surgical triplet spatiotemporal feature sequence. This surgical triplet spatiotemporal feature sequence is then input into a second independent branch network to obtain a video-level component class activation mapping sequence, including: The video-level spatial feature sequence is input into a temporal convolutional network to model temporal relationships, resulting in a surgical triplet spatiotemporal feature sequence. The spatiotemporal feature sequence of the surgical triplet is input into the second independent branch network for dimensionality transformation to obtain a video-level component class activation mapping sequence. The second independent branch network includes a four-convolutional network, a fifth-convolutional network, and a sixth-convolutional network.

7. The surgical triplet recognition method combining frame-level and video-level class activation guidance according to claim 1, characterized in that, The video-level component relationship modeling guided by spatiotemporal features, based on the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence, yields the surgical triplet recognition result for each video frame in the target video, including: The video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence are concatenated along the feature dimension to obtain the concatenated feature. The splicing features are subjected to multi-head self-attention operation to obtain a video-level surgical triplet relation representation, which includes a self-attention representation token of the component, a second cross-attention representation token between components, and a second cross-attention representation token between the component and the surgical triplet. The video-level surgical triplet relation representation is input into the hybrid expert model. Different linear modulation experts are dynamically activated according to the routing network to process different relation representation tokens in the video-level surgical triplet relation representation. Spatiotemporal feature-guided video-level component relation modeling is performed to obtain the surgical triplet video-level features. Based on the video-level features of the surgical triplet, the surgical triplet identification result is obtained for each video frame in the target video.

8. A surgical triplet recognition device combining frame-level and video-level class activation guidance, implemented using the surgical triplet recognition method combining frame-level and video-level class activation guidance as described in any one of claims 1 to 7, characterized in that, The device includes: The acquisition module is used to acquire the target video, obtain multiple video frame data based on the target video, input each video frame data into the backbone network to obtain frame-level spatial features, and input the frame-level spatial features into the first independent branch network to obtain the frame-level component class activation mapping sequence. The first processing module is used to perform frame-level component relationship modeling guided by spatial features based on the frame-level spatial features and the frame-level component class activation mapping sequence to obtain surgical triplet features, wherein the surgical triplet includes three components: instrument, action, and target. The second processing module is used to train the backbone network based on the surgical triplet features, input the target video into the trained backbone network, and obtain a video-level spatial feature sequence. The third processing module is used to perform temporal relationship modeling based on the video-level spatial feature sequence to obtain the surgical triplet spatiotemporal feature sequence, and input the surgical triplet spatiotemporal feature sequence into the second independent branch network to obtain the video-level component class activation mapping sequence. The recognition module is used to perform spatiotemporally feature-guided video-level component relationship modeling based on the video-level component class activation mapping sequence and the surgical triplet spatiotemporal feature sequence, so as to obtain the surgical triplet recognition result for each video frame in the target video.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the surgical triplet recognition method combining frame-level and video-level class activation guidance as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the surgical triplet recognition method combining frame-level and video-level class activation guidance as described in any one of claims 1 to 7.