Surgical evaluation method and system based on multimodal representation and causal reasoning

By introducing multimodal representation and causal reasoning methods in surgical activity analysis, tool movement trajectory and position heat map are generated, and combined with visual and trajectory characteristics to perform motion recognition and skill evaluation, the problem of insufficient efficiency and universality of surgical activity analysis in the existing technology is solved, and high-performance surgical action recognition and skill evaluation is achieved.

CN119418164BActive Publication Date: 2025-05-06SHANDONG UNIV QILU HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411522302.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-05-06
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

The prior art lacks efficiency and repeatability in surgical activity analysis, especially in common minimally invasive surgical scenarios, where traditional methods are less versatile, and most studies use only single-modal data, limiting the performance of surgical skills evaluation.

Method used

Using surgical evaluation methods based on multimodal representation and causal reasoning, a network structure based on causal reasoning is designed for joint prediction by obtaining surgical videos, generating tool motion trajectories and location heat maps, combining visual and trajectory features, and inputting them into the video classification model for action recognition and skill evaluation, and designing a network structure based on causal reasoning for joint prediction.

Benefits of technology

It improves the accuracy and generalization ability of surgical activity analysis, and can effectively perform surgical action recognition and skill evaluation in ordinary minimally invasive surgical scenarios, with good versatility and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418164B_ABST
    Figure CN119418164B_ABST
Patent Text Reader

Abstract

The present invention discloses a surgical evaluation method and system based on multimodal representation and causal reasoning, relates to the field of artificial intelligence technology, and conducts in-depth research on surgical skill analysis and evaluation based on deep learning in common surgical scenarios. Starting from the two perspectives of data and algorithm framework, the present invention proposes two core contents to improve the evaluation performance of surgical skills. First, several new technologies have been developed to provide observation data of three different modalities in common minimally invasive surgery, including video, trajectory and language. On this basis, a cross-modal comparative learning strategy is designed, which can learn discriminative features from data of three different modalities. Secondly, a unified framework for joint surgical action recognition and skill evaluation is proposed, which designs a prediction structure based on causal reasoning to model the causal relationship between the two tasks, thereby achieving higher performance action recognition and skill evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a surgical operation evaluation method and system based on multimodal representation and causal reasoning. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] Related studies have shown that postoperative outcomes are closely related to the surgeon's intraoperative surgical activities. Therefore, it is very important to analyze surgical activities so that factors leading to poor surgical outcomes can be identified and surgical feedback can be provided in a timely manner. Traditionally, the analysis of surgical activities relies on experienced surgeons who evaluate technical performance based on predetermined standards or their own surgical experience. However, this evaluation is usually time-consuming and labor-intensive, lacking efficiency and repeatability.

[0004] Over the past few decades, artificial intelligence (AI) has achieved great success in many fields, inspiring many researchers to analyze surgical activities through deep learning. Generally, the understanding of surgical activities is mainly divided into two aspects: what operations are performed during the operation and how well these operations are performed. Therefore, previous studies usually treat surgical activity analysis as a classification task, including the two questions of "what actions are performed" and "how the actions are performed". Accordingly, either the surgical video is identified as a specific action type, or the surgical skills are evaluated at a specific level.

[0005] Years of practice have shown that it is not easy to conduct a refined analysis of surgical activities, and data is the key to these deep learning-based surgical activity analysis methods. Thanks to surgical robot systems such as the da Vinci robot, in robot-assisted minimally invasive surgery, there are multiple modalities of data available for analyzing surgical skills, among which robot motion trajectories and video recordings are the two most commonly used data. Some recent studies have shown that the use of multimodal data (combining motion trajectories and videos) can improve the performance of surgical activity analysis, because different modalities can provide complementary information and promote deeper analysis. However, in actual surgical cases, traditional minimally invasive surgery is a more common choice, especially in hospitals in underdeveloped countries and regions. In this case, standard robot motion trajectory data is no longer available, and video data is the only option, which limits the applicability of this type of method in general surgical scenarios. Some researchers have tried to extract optical flow features from surgical videos to provide certain motion information in the absence of trajectory data. However, the motion information contained in the optical flow features is very limited. Therefore, a natural question is, in this case, what other modalities of data are helpful for surgical activity analysis, and how to use these modal data? Another issue worth noting is that most existing studies choose to validate their methods on the JIGSAWS dataset, which has several limitations that prevent it from being applied to more general surgical scenarios, including limited dataset size, poor data diversity, and the data only containing simulated environments.

[0006] Existing works usually use expert ratings (ordinal scale) or proficiency scores (interval scale) to evaluate surgical skill levels. A recent survey showed that the most widely used skill level assessment scheme is three different expert levels, namely novice, intermediate and expert. The technical schemes of existing methods for surgical skill level assessment are mostly the same. They first extract features from surgical data and then predict the skill level based on the extracted features. From the perspective of data modality, most studies either use only motion trajectory data or only surgical video data, and only a few studies use data from both modalities at the same time. The introduction of multimodal data can provide more information, thereby improving the performance of skill assessment. However, a major limitation of existing methods is that motion trajectory data is usually only applicable to robot-assisted surgery scenarios, but not to general minimally invasive surgery scenarios, resulting in poor versatility of existing methods.

[0007] Most previous studies only focus on one of the two core issues in surgical activity analysis. In contrast, recent studies have attempted to combine the two tasks in one framework. Intuitively, it is more reasonable to treat surgical action recognition and skill assessment as a whole rather than separately. This is because the evaluation criteria for different types of surgical actions are not exactly the same, and surgical skills should be evaluated according to their corresponding action types. Nevertheless, these methods usually just add two parallel prediction branches after the feature encoder to achieve a so-called unified framework, which is used to predict action type and skill level, respectively. However, the causal correlation between the two tasks has not been deeply studied and analyzed. Summary of the invention

[0008] To overcome the above-mentioned deficiencies of the prior art, the present invention provides a surgical evaluation method and system based on multimodal representation and causal reasoning, which solves the problem of how to model the correlation between surgical action recognition and skill evaluation, and solves these two tasks simultaneously within a unified framework.

[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0010] In a first aspect, the present invention provides a surgical evaluation method based on multimodal representation and causal reasoning, comprising:

[0011] Obtain surgical videos to be evaluated;

[0012] Input the surgical video into a visual encoder for feature extraction to obtain visual features; input the surgical video into a tool trajectory generation module to generate a tool motion trajectory and a position heat map sequence;

[0013] The motion trajectory is input into a trajectory encoder for feature extraction to obtain trajectory features; the position heat map sequence and the visual features are fused and linearly mapped to obtain local visual features; the visual features are linearly mapped to obtain global visual features;

[0014] The trajectory features, local visual features, and global visual features are input into a video classification model for processing. The visual classification model includes a block embedding module, an encoder, a parallel action recognition prediction branch, and a skill level assessment prediction branch that are connected in sequence. The action recognition prediction branch outputs a predicted action type, and the skill level assessment prediction branch outputs a predicted skill level.

[0015] According to a further technical solution, the visual encoder adopts a visual encoder based on a convolutional neural network.

[0016] According to a further technical solution, the tool trajectory generation module uses a surgical tool detector based on YOLOv8 to detect the operating tool, and generates a two-dimensional trajectory through the position of the detected operating tool boundary box, converts the two-dimensional trajectory into a three-dimensional trajectory based on a depth change estimation strategy, and generates a tool motion trajectory.

[0017] A further technical solution is to project the area where the operating tool is located onto an empty background based on the detected operating tool boundary box, thereby creating a position heat map sequence.

[0018] According to a further technical solution, the trajectory encoder is constructed based on a bidirectional LSTM recurrent neural network.

[0019] According to a further technical solution, a network structure based on causal reasoning is designed in the visual classification model.

[0020] A further technical solution is that the network structure based on causal reasoning is specifically:

[0021] Converting the predicted action type into an action feature;

[0022] In the skill level assessment prediction branch, the encoder output features, the linear mapping layer output features in the action recognition prediction branch and the action features are concatenated and input into the linear mapping layer to obtain implicit features; subsequently, the action features and the implicit features are concatenated and input into the linear mapping layer and the classification layer in turn to finally output the predicted skill level.

[0023] In a second aspect, the present invention provides a surgical evaluation system based on multimodal representation and causal reasoning, comprising:

[0024] A video acquisition module is configured to: acquire a surgical video to be evaluated;

[0025] A video processing module is configured to: input the surgical video into a visual encoder for feature extraction to obtain visual features; input the surgical video into a tool trajectory generation module to generate a tool motion trajectory and a position heat map sequence;

[0026] The multimodal feature extraction module is configured to: input the motion trajectory into a trajectory encoder for feature extraction to obtain trajectory features; perform linear mapping after fusing the position heat map sequence and the visual features to obtain local visual features; perform linear mapping on the visual features to obtain global visual features;

[0027] The joint prediction module is configured to: input the trajectory features, local visual features, and global visual features into a video classification model for processing, wherein the visual classification model includes a block embedding module, an encoder, a parallel action recognition prediction branch, and a skill level assessment prediction branch connected in sequence, wherein the action recognition prediction branch outputs a predicted action type, and the skill level assessment prediction branch outputs a predicted skill level.

[0028] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the surgical evaluation method based on multimodal representation and causal reasoning as described in the first aspect.

[0029] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the surgical evaluation method based on multimodal representation and causal reasoning as described in the first aspect are implemented.

[0030] One or more of the above technical solutions have the following beneficial effects:

[0031] The present invention aims to further improve the accuracy and generalization ability of deep learning-based surgical activity analysis. Specifically, the present invention proposes two core contents to improve the performance of surgical skill evaluation from the perspectives of data and algorithm framework. First, the present invention introduces three different modalities of data for surgical activity analysis in ordinary minimally invasive surgery scenarios, including video, trajectory and language. In ordinary minimally invasive surgery scenarios, trajectory and language data themselves cannot be obtained. To this end, the present invention proposes a motion trajectory generation method based on visual detection and a language description generation method based on a large language model. In addition, the present invention also designs a cross-modal contrast learning strategy that can learn compact and discriminative multimodal data representations from raw data. Secondly, the present invention proposes a unified framework for joint surgical action recognition and skill evaluation using multimodal data representations. Among them, the present invention designs a network structure based on causal reasoning to model the causal correlation between the two tasks, and on this basis, joint prediction is performed to achieve higher performance action recognition and skill evaluation. Features of multiple modalities and different scales are processed by a conversion encoder and then input into a network structure based on causal reasoning, thereby simultaneously predicting surgical actions and skill levels.

[0032] The effectiveness of the proposed method is verified by experiments, and it is proved that it has the most advanced performance in both surgical action recognition and skill assessment tasks. In addition, the algorithm is tested on the RARP-45 dataset, and the results show that the method of the present invention can be used across different surgeons and hospitals, and has good generalization. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0034] Figure 1 is a schematic diagram of surgical tool detection results and generated trajectories in the evaluation method of an embodiment of the present invention;

[0035] Figure 2 is a schematic diagram of a video description generation method based on a large language model in an embodiment of the present invention;

[0036] Figure 3 It is a unified framework for surgical action recognition and skill level assessment in an embodiment of the present invention;

[0037] Figure 4 It is a schematic diagram of the design structure of the track encoder in an embodiment of the present invention;

[0038] Figure 5 is the data distribution of the data set in the embodiment of the present invention in terms of surgical action categories and skill levels;

[0039] Figure 6 is the normalized confusion matrix of the evaluation method of the embodiment of the present invention on the surgical action recognition task;

[0040] Figure 7 is a normalized confusion matrix of the evaluation method of the embodiment of the present invention on the skill level evaluation task;

[0041] Figure 8 is a comparative experimental result using different motion trajectory representations in the embodiments of the present invention;

[0042] Fig. 9 These are two commonly used prediction network structures in the prior art in the embodiments of the present invention;

[0043] Fig.10 It is the comparison result of classification accuracy of different network structures in the embodiment of the present invention. DETAILED DESCRIPTION

[0044] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0045] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0046] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0047] Terminology explanation:

[0048] The purpose of surgical skill assessment is to analyze the surgeon's technical performance during surgery, that is, to evaluate the surgeon's surgical skill level. Surgical skill assessment is very important for the education of novice surgeons and the improvement of surgical quality.

[0049] Embodiment 1

[0050] like Figure 3 As shown, a surgical evaluation method based on multimodal representation and causal reasoning includes the following steps:

[0051] S1: Obtain the surgical video to be evaluated;

[0052] In this embodiment, the aim is to improve the performance of deep learning-based surgical activity analysis tasks by applying multimodal data to general minimally invasive surgery scenarios.

[0053] Compared with robot-assisted surgery, ordinary minimally invasive surgery has a much wider range of applications. In such ordinary surgical scenarios, surgical videos may be the only available data modality. In order to quickly build a multimodal dataset, the present invention reuses data from the RARP-45 dataset because it provides real surgical videos with fine action category annotations. Although the RARP-45 dataset was collected from robot-assisted surgery, the relevant motion trajectory data has not yet been publicly released. In view of the fact that motion trajectories are very helpful for surgical activity analysis, the present invention designs a method to generate corresponding motion trajectories by visually tracking surgical tools in surgical videos. In addition to video and trajectory data, a language description of the surgical video is also introduced to provide rich semantic information to further improve the performance of surgical activity analysis.

[0054] S2: Input the surgical video into a visual encoder for feature extraction to obtain visual features; input the surgical video into a tool trajectory generation module to generate a tool motion trajectory and a position heat map sequence;

[0055] In this embodiment, the visual encoder adopts a visual encoder based on a convolutional neural network. First, M frames are uniformly sampled from a surgical video V of length N, and then input into a visual encoder based on a convolutional neural network to extract a visual feature φ(V) of dimension (M, C, w, h).

[0056] In this embodiment, the tool trajectory generation module uses a surgical tool detector based on YOLOv8 to detect the operating tool, and generates a two-dimensional trajectory through the position of the detected operating tool boundary box, converts the two-dimensional trajectory into a three-dimensional trajectory based on the depth change estimation strategy, and generates the tool motion trajectory. Specifically, all N frames of images in the video clip V will be input into the tool trajectory generation module, and finally the generated trajectory T = {trajL, trajR} is obtained.

[0057] Tool motion trajectory generation specifically includes:

[0058] In the tool trajectory generation module, a tool motion trajectory generation method based on visual detection is proposed, which includes two parts: surgical tool detection and trajectory interpolation sampling, to obtain motion trajectories from general surgical scenarios.

[0059] (1) Surgical tool testing

[0060] Given a surgical video clip V containing N frames of images, we first need to accurately detect the surgical tools appearing in the video. To this end, the present invention builds a surgical tool detector based on YOLOv8. YOLOv8 has been proven to be a powerful and easily scalable object detection framework in the field of computer vision. Figure 1 As shown, the detection targets of the present invention are two patient-side operating tools (PSM), namely the left operating tool and the right operating tool. For each frame of the video segment V, the bounding boxes of the two PSMs are detected by the object detector (surgical tool detector), and the center points of the bounding boxes are used to represent the current positions of the two PSMs, expressed as normalized two-dimensional coordinates: posL = (x l ,y l ) and posR=(x r ,y r ). Figure 1 The surgical tool detection results and the generated trajectories are shown in .

[0061] In practical applications, two-dimensional position information often cannot effectively reflect the motion trajectory of surgical tools in a three-dimensional workspace. Generally, a depth estimation algorithm can be used to convert the two-dimensional position in an image into a three-dimensional space, but this is not applicable to the scenario of the present invention. Because most existing depth estimation algorithms and related data sets focus on general scenarios, conventional depth estimation algorithms often have poor results when applied in surgical scenarios. According to observations, the key to motion analysis of surgical tools is to analyze the laws of their motion trajectories rather than to obtain their absolute positions in the three-dimensional world. Inspired by this, the present invention proposes a simple and effective "depth change" estimation strategy. In a surgical video clip V, the present invention assumes that the camera's viewing angle is fixed, then the surgical tools in the video will follow the rule of "near large and far small": when the tool moves toward the camera, the detected bounding box will become larger, and vice versa.

[0062] The area of ​​the left and right tool bounding boxes detected in the first frame of the video clip V is used as the standard, denoted as areaL1=w l1 ×h l1 and areaR1=w r1 ×h r1 For the i-th frame image in the video segment V, the detected left and right tool bounding box areas are recorded as areaL i =w li ×h li and areaR i =w ri ×h ri Assuming that the position of the camera is the origin of the z-axis, the present invention can calculate the z-axis coordinates of the positions of the left and right tools as and Accordingly, the three-dimensional position coordinates of the left and right tools can be expressed as: posL = (x l ,y l ,z l ) and posR=(x r ,y r ,z r ). For a surgical video clip V containing N frames of images, its complete tool motion trajectory can be expressed as: trajL = {posL1, posL2, ..., posL N} and trajR={posR1,posR2,...,posR N}.

[0063] (2) Trajectory interpolation based on TCB splines

[0064] In real surgical scenarios, some interference factors may cause surgical tools to be missed, such as lighting changes, tissue bleeding, and visual field obstruction. In order to complete these missed trajectory points, the present invention adopts a trajectory interpolation strategy based on TCB splines. TCB splines are a classic spline interpolation algorithm that uses variable splines to create a smooth curve through a series of given points. TCB splines provide three hyperparameters to flexibly adjust the shape of the fitting curve, including tension τ k ∈[-1,1], continuity γ k ∈[-1,1] and bias β k ∈[-1,1]. Since the moving speed of the surgical tool is variable, the distribution of trajectory points in the workspace is not uniform. The present invention roughly divides these points into tightly distributed points and loosely distributed points, and sets different hyperparameters (τ k ,γ k ,β k ), where the former is set to (0.5,0.5,0) and the latter is set to (-0.5,0.2,0).

[0065] In this embodiment, in ordinary minimally invasive surgery scenarios, language data itself cannot be obtained, and the present invention proposes a language description generation method based on a large language model.

[0066] Specifically, if a surgical intern seeks advice on surgical techniques from an experienced surgeon, they usually do so through verbal communication. It can be said that natural language descriptions are often the most convenient and effective medium of communication. Language descriptions can provide semantic details of surgical activities, which is potentially helpful for improving the performance of neural network models in understanding surgical activities. However, compared with the annotation of surgical actions, it is more challenging to annotate high-quality language descriptions for surgical videos. On the one hand, the collection of language descriptions is very laborious and requires a lot of manpower. On the other hand, the collected language descriptions are usually colloquial language, which contains many redundant and erroneous words. In recent years, large language models (LLMs) have demonstrated powerful understanding and analysis capabilities. Inspired by this, the present invention utilizes the powerful understanding and reasoning capabilities of LLMs to design a video description generation method based on large language models.

[0067] Figure 2The proposed method for generating video descriptions based on a large language model is presented. The main advantage of this method is that it can provide structured and semantically rich language descriptions for surgical videos with very little human intervention. The specific process is as follows: experienced surgeons are organized to watch surgical videos, and they can make comments in spoken language during the viewing process. The speech recognition model Whisper developed by OpenAI is used to convert the spoken comments into textual language descriptions in real time. After the viewing, GPT-4 is used as an interviewer: it will ask the surgeon a series of questions based on the surgical evaluation criteria established by the Japanese Society of Endoscopic Surgery. The surgeon's answers will also be converted into textual language. Please note that the present invention realizes the spoken dialogue between the surgeon and GPT-4 through a text-to-speech (TTS) engine. For each surgical video, the present invention will input all its corresponding textual language descriptions into GPT-4 for language polishing. Since GPT-4 is a general large language model, the present invention fine-tunes it through prompt engineering to make it well adapted to surgical scenarios. Finally, GPT-4 is able to generate a structured and semantically rich language description for each surgical video.

[0068] According to the above method, the present invention obtains data of three different modalities in a common minimally invasive surgery scenario for surgical activity analysis, including video, trajectory and language.

[0069] S3: inputting the motion trajectory into a trajectory encoder for feature extraction to obtain trajectory features; fusing the position heat map sequence and the visual features and performing linear mapping to obtain local visual features; performing linear mapping on the visual features to obtain global visual features;

[0070] In this embodiment, the goal is to build a unified framework that uses multimodal features as input to simultaneously perform surgical action recognition and skill level assessment. Its mathematical representation is: given a surgical video clip V containing N frames of images, V = {I1, I2, ..., I N}, the goal is to simultaneously identify its surgical action type g∈G and rate its skill level l∈L. Among them, G={G0,G1,G2,G3,G4,G5,G6,G7} represents 8 different surgical action types, and L={L0,L1,L2} represents 3 different skill levels.

[0071] like Figure 3 As shown in the figure, this embodiment proposes a unified framework based on ViViT (Video Vision Transformer) to jointly perform surgical action recognition and skill level assessment using multimodal data features.

[0072] First, the data input into the unified framework based on ViViT is processed to extract the features of multimodal data for input. Specifically:

[0073] The tool motion trajectory T generated by the tool trajectory generation module is input into the trajectory encoder for feature extraction, and finally the trajectory feature F is obtained. T .like Figure 4 As shown in Figure 1, the trajectory encoder is built based on a bidirectional LSTM recurrent neural network. During the trajectory generation process, surgical tool detection is performed on N frames of images. Based on the detection results, i.e., the detected operating tool boundary boxes, the areas where the left and right tools are located are projected onto an empty background, thereby creating a position heat map sequence H of length M. Figure 3 As shown in the figure, the position heat map highlights the location information of the surgical tools in the video, which enhances the representation ability of visual features from a local perspective.

[0074] In order to ensure consistency with the visual feature dimension φ(V), this embodiment scales the size of the position heat map to a uniform size of w×h, and expands its number of channels from 1 dimension to C dimension. The heat map sequence H and the visual feature φ(V) are inner-producted to obtain the fusion feature. Then, the local visual feature F is obtained by linear mapping through several layers of neural networks. Vl Similarly, the visual feature φ(V) is input into several layers of neural networks for linear mapping, and the global visual feature F can be obtained. Vg In essence, F Vl and F Vg They represent the local and global visual information of a surgical video respectively. Finally, all the features, namely the trajectory features F T , local visual features F Vl , global visual features F Vg They are all sent to the ViViT visual classification model for processing and then input into different prediction branches.

[0075] S4: Input the trajectory features, local visual features, and global visual features into a video classification model for processing, wherein the visual classification model includes a block embedding module, an encoder, a parallel action recognition prediction branch, and a skill level assessment prediction branch connected in sequence, wherein the action recognition prediction branch outputs a predicted action type, and the skill level assessment prediction branch outputs a predicted skill level.

[0076] In this embodiment, if Figure 3As shown in the figure, the visual classification model is a Transformer-based video classification model (ViViT). It includes a patch embedding module, an encoder, a parallel action recognition prediction branch, and a skill level assessment prediction branch connected in sequence. The patch embedding module divides the input data into multiple fixed-size blocks by convolution; the action recognition prediction branch includes a linear mapping layer and a Softmax classification layer connected in sequence; the skill level assessment prediction branch includes a linear mapping layer, a linear mapping layer, and a Softmax classification layer connected in sequence.

[0077] When dealing with joint prediction tasks, existing methods usually simply add two parallel prediction branches after the feature encoder. In contrast, this embodiment designs a network structure based on causal reasoning to model the causal relationship between surgical action recognition and skill level assessment.

[0078] The network structure based on causal reasoning is specifically as follows: the predicted action type is converted into action features; in the skill level assessment prediction branch, the encoder output features, the linear mapping layer output features in the action recognition prediction branch and the action features are concatenated and input into the linear mapping layer to obtain implicit features; subsequently, the action features and the implicit features are concatenated and input into the linear mapping layer and the classification layer in turn, and finally the predicted skill level is output.

[0079] Further, such as Figure 3 As shown, in the surgical action recognition prediction branch, the encoder output features are sequentially input into the linear mapping layer and the Softmax classification layer, and the predicted action type is finally output. Assuming that the output feature of the ViViT encoder is x1, and the feature after linear mapping processing in the action recognition prediction branch is x2, this embodiment converts the action category label g into the action feature x3 through 0-1 encoding. Then in the skill level assessment prediction branch, the features x1, x2, and x3 are concatenated, and then input into the linear mapping layer to obtain the implicit feature x4. Subsequently, the features x3 and x4 are concatenated, and then sequentially input into the linear mapping layer and the Softmax classification layer, and the predicted skill level is finally output.

[0080] In addition, this embodiment designs a cross-modal contrastive learning strategy, the core principle of which is to use semantically rich language descriptions to enhance the discriminability and representativeness of visual and trajectory features. As mentioned above, this embodiment generates a text language description for each surgical video. Figure 3 As shown, the language description is input into a pre-trained text encoder, which is implemented using CLIP to extract language features F LSo far, this embodiment has obtained three different modal features: trajectory feature F T 、Language Features F L And fusion video features F V (F Vl and F Vg ). The goal of cross-modal contrastive learning is to align video-trajectory data in the video-trajectory common space and video-language data in the video-language common space. The paired data used for training is obtained from the constructed multimodal dataset. Positive samples refer to different modal data corresponding to the same action category, and negative samples refer to modal data corresponding to any two different action categories. The overall loss function of cross-modal contrastive learning is defined as:

[0081] L=λ VT *NCE(F V ,F T )+λ VL *MIL-NCE(F V ,F L ) (1)

[0082]

[0083] in, Contains positive samples, Contains negative samples, τ is a hyperparameter used to adjust the degree of distinction between positive and negative samples. VT and λ VL It is to adjust the weighted weights of two different loss functions.

[0084] A two-stage training strategy is also designed to optimize the constructed network model. In the first stage, cross-modal contrastive learning is performed using the data of three different modalities using the loss function defined in Formula 1. The goal of this training stage is to improve the discriminability and representativeness of visual and trajectory features. In the second stage, the entire network is trained based on the weights obtained in the first stage. In this process, the classification cross entropy loss function of the action recognition and skill assessment tasks is calculated as the basis for network update. The overall loss function target of this training stage is a weighted combination of the cross entropy loss functions corresponding to the two tasks.

[0085] Experimental analysis

[0086] (1) Dataset definition

[0087] As mentioned above, this embodiment creates a multimodal dataset by reusing the surgical video data provided by the RARP-45 dataset. The RARP-45 dataset provides a total of 45 dorsal vascular complex suture surgery records, which are manually segmented into 1634 surgical video clips according to 8 pre-defined action categories. Considering that the RARP-45 dataset does not provide labels for surgical skill levels, this embodiment re-labels the surgical videos according to three pre-defined skill levels. In addition, a dataset is created for training the surgical tool detector mentioned above.

[0088] Specifically, this embodiment follows the action category definition used in the RARP-45 dataset, that is, there are 7 fine-grained surgical actions and one background category. Since the RARP-45 dataset was established for surgical action recognition, no labels are provided for surgical skill levels. In the present invention, this embodiment uses three different skill levels to measure the surgical skills of surgeons, namely novice, intermediate and expert. The detailed definitions of surgical action categories and skill levels are shown in Table 1. In order to label each surgical video with a specific skill level, the present invention has formed a well-trained annotation team, all of whom are experienced urologists who have been engaged in urological surgery for more than two years.

[0089] Data augmentation is a strategy widely used in the field of deep learning to improve model performance, especially in applications with small data sets. Therefore, this embodiment uses a variety of data augmentation strategies to expand the data set, including image scaling, cropping, rotation, flipping, shearing, blurring and sharpening. Through data augmentation, 18,780 data samples were finally obtained. The data set is divided into a training set and a validation set in a ratio of 8:2. Figure 5 The data distribution of the entire dataset in terms of surgical action categories and skill levels is shown. In addition, in order to create a training dataset for the surgical tool detector, this embodiment randomly samples images from the collected surgical videos and manually labels these images. The labeling process is performed according to the standard steps of YOLOv8.

[0090] Table 1. Detailed definitions of surgical action categories and skill level categories

[0091]

[0092] (2) Algorithm Evaluation

[0093] 1. Comparison Method

[0094] In order to fully test the performance of the algorithm proposed in the present invention, this embodiment selects 9 comparison methods in the field of surgical activity analysis based on deep learning for comparison. These comparison methods are representative methods proposed in the past five years, covering surgical action recognition tasks and skill level assessment tasks. According to the different task types, the comparison methods can be divided into three groups: (1) only performing surgical action recognition tasks; (2) only performing skill level assessment tasks; (3) performing both tasks at the same time.

[0095] 2. Evaluation Indicators

[0096] This embodiment uses two common indicators to evaluate the performance of surgical action recognition and skill level assessment tasks. The first is accuracy, which is usually used to measure the ability of a classification model to correctly predict results. The other is the F1 score, which is usually used to measure the overall accuracy of binary or multi-classification tasks. It takes into account both the accuracy and recall of the classification model, and can be regarded as a weighted average of the model's accuracy and recall. Due to differences in the experimental settings of the comparison methods, this embodiment uses a special variant of the F1 score, the F1@10 score, to evaluate those comparison methods that use frame-level experimental settings. The F1@10 score is calculated by calculating the proportion of the intersection between the category of each predicted segment and the category of the real segment that is greater than a certain threshold. In order to make a fair comparison, this embodiment simultaneously calculates the F1 score and the F1@10 score of the method proposed in the present invention.

[0097] 3. Experimental Results

[0098] The results of the comparative experiment are shown in Table 2. From the results in Table 2, it can be seen that the method proposed in the present invention outperforms all the comparative methods in the tasks of surgical action recognition and skill level assessment. These results demonstrate the effectiveness of the method proposed in the present invention and reflect the importance of exploring multimodal scene representation and integrating surgical action recognition and skill level assessment into a unified framework. This embodiment plots the normalized confusion matrix of the surgical action recognition and skill level assessment tasks to demonstrate the classification performance of the method of the present invention in each subcategory. The images are shown in Figure 2. Figure 6 and Figure 7 . Based on the results shown in the picture, there are two main observations in this embodiment. First, the unbalanced data distribution leads to performance differences between different categories of surgical action categories and skill level grades. Generally speaking, categories with less data usually achieve higher accuracy. The possible reason is that the model overfits on fewer data samples. Secondly, in the surgical action recognition task, the prediction error is mainly concentrated at the boundaries of different action categories. This may be because there is a certain overlap between adjacent action categories, which makes it difficult for the algorithm to effectively distinguish them.

[0099] Table 2. Comparative experimental results of different methods on surgical action recognition and skill level assessment tasks

[0100]

[0101] (3) Ablation experiment

[0102] 1. Multimodal Representation

[0103] The introduction of multimodal data can improve the accuracy of surgical activity analysis because data of different modalities can provide complementary information. Motion trajectory data and video data are the two most widely used modalities, but they are often only available in robot-assisted surgery scenarios. In the present invention, this embodiment introduces three different modalities of data, including video, trajectory, and language, in a general minimally invasive surgery scenario. In order to verify whether the introduction of these three modal data brings performance improvement and which modality contributes more, the present invention conducts ablation experiments on multimodal representation. Five test variants, namely V1-V5 listed in Table 3, are created by different combinations of the three modal data. The complete method of this embodiment uses all three modalities and is referred to as V6. In Table 3, this embodiment calculates the accuracy of different variants in surgical action recognition and skill level assessment tasks. The following conclusions can be drawn from the results in Table 3. First, variant V6 performs better than all other variants, which shows that the introduction of data of three different modalities in this article is effective. Secondly, the comparison results between V1 and V3, V2 and V4, and V5 and V6 clearly show that the introduction of language modality data and the designed cross-modality contrast learning strategy can significantly enhance the discriminability of features, thereby improving the overall performance. Thirdly, by observing the comparison results between V1 and V2, and V3 and V4, this embodiment finds that the representation of video modality is more effective than the representation of motion trajectory.

[0104] In conventional minimally invasive surgery scenarios, the robot's motion trajectory is usually not available. For this reason, this embodiment proposes a motion trajectory generation method based on visual detection. However, this method estimates the motion trajectory by detecting the surgical tools in the video stream. Does this estimated motion trajectory have the same information representation capability as the standard motion trajectory collected by the robot platform? In addition, some studies have proposed that in the absence of the robot's motion trajectory, optical flow can be extracted from the video to simulate motion information. Is the optical flow method a better choice? In order to answer these questions, this embodiment adopts a simple framework of a bidirectional multi-layer independent RNN model + a deep convolutional neural network model (DCNN) built based on the VGG architecture, and uses three different input data for surgical action recognition, namely the standard robot motion trajectory, the motion trajectory estimated from the video, and the optical flow extracted from the video. The action recognition experiment was conducted on the suturing task of the JIGSAWS dataset because this dataset provides standard robot motion trajectory data. The experimental comparison results are shown in the figure. Figure 8As shown in the figure, it can be seen that the task accuracy of the variants using the robot's standard motion trajectory data and the estimated motion trajectory data is very close, and is significantly higher than the accuracy of the variant using optical flow data. These results show that the motion trajectory generation method based on visual detection proposed in the present invention is very effective in modeling motion information, which provides a promising solution for common surgical scenarios where the robot's standard motion trajectory cannot be obtained.

[0105] Table 3. Comparison of the accuracy of action recognition and skill assessment tasks using different data variants

[0106]

[0107] 2. Surgical Activity Analysis Consent Framework

[0108] In the present invention, this embodiment designs a network structure based on causal reasoning to model the causal correlation between the surgical action recognition task and the skill level assessment task. In order to demonstrate the technical advantages of this structure, this embodiment compares it with the existing methods. Fig. 9 Two prediction network structures commonly used in existing work are demonstrated. The first is a "multiplexing structure", which can simultaneously complete surgical action recognition and skill level assessment tasks in one network. However, in one reasoning, the network can only be used to independently perform one of the two tasks, and cannot perform them at the same time. The second is a "parallel structure", which adds two parallel prediction branches after the feature encoder, which are used to predict action type and skill level respectively. This embodiment replaces the causal reasoning-based network structure designed in this paper with these two different structures, while keeping all other settings unchanged, so as to conduct an ablation comparison experiment. The comparison results are shown in Figure 2. Fig.10 As shown in the results, it can be clearly seen that the network structure based on causal reasoning designed by the present invention is more effective than the comparison method.

[0109] (4) Generalization ability test experiment

[0110] Whether it has the ability to generalize in multiple scenarios is an important indicator to measure the performance of an algorithm model. Most previous studies were conducted on the JIGSAWS dataset, but the dataset is data collected in a simulated environment. Therefore, these previous methods often perform poorly in real surgical scenarios. Recently, some researchers created the RARP-45 dataset, which is a more realistic surgical activity analysis dataset. However, this dataset only contains surgical records collected in the same hospital, thus limiting the general ability in more general scenarios. To this end, this embodiment conducted an external experiment to test whether the method of the present invention can be extended to external surgical videos from different surgeons and hospitals. In this embodiment, a total of 300 surgical video clips were collected from YouTube and annotated. These videos cover 8 different hospitals and 10 surgeons. Finally, the method proposed in the present invention achieved an accuracy of 70.7% and 79.3% in surgical action recognition and skill level assessment tasks, respectively. The results show that the method of the present invention has good general generalization ability.

[0111] In summary, the present invention deeply studies several key elements of surgical activity analysis based on deep learning, including data and framework, and contributes two key innovations on this basis. First, a novel technology is designed to introduce three different modalities of data observation for conventional minimally invasive surgery scenarios, including video, trajectory and language. Accordingly, a cross-modal comparative learning strategy is designed, which can learn discriminative multimodal representations from data of three modalities. Secondly, a unified framework is proposed to jointly perform surgical action recognition and skill level assessment using multimodal representations. In this framework, a prediction structure based on causal reasoning is designed, which can effectively model the causal correlation between the two tasks of surgical action recognition and skill level assessment. Finally, through a large number of experimental tests, it is proved that the method of the present invention is superior to the mainstream methods in the current field and has good generalization and versatility.

[0112] Embodiment 2

[0113] This embodiment provides a surgical evaluation system based on multimodal representation and causal reasoning, including:

[0114] A video acquisition module is configured to: acquire a surgical video to be evaluated;

[0115] A video processing module is configured to: input the surgical video into a visual encoder for feature extraction to obtain visual features; input the surgical video into a tool trajectory generation module to generate a tool motion trajectory and a position heat map sequence;

[0116] The multimodal feature extraction module is configured to: input the motion trajectory into a trajectory encoder for feature extraction to obtain trajectory features; perform linear mapping after fusing the position heat map sequence and the visual features to obtain local visual features; perform linear mapping on the visual features to obtain global visual features;

[0117] The joint prediction module is configured to: input the trajectory features, local visual features, and global visual features into a video classification model for processing, wherein the visual classification model includes a block embedding module, an encoder, a parallel action recognition prediction branch, and a skill level assessment prediction branch connected in sequence, wherein the action recognition prediction branch outputs a predicted action type, and the skill level assessment prediction branch outputs a predicted skill level.

[0118] Embodiment 3

[0119] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method of embodiment 1 when executing the program.

[0120] Embodiment 3

[0121] The purpose of this embodiment is to provide a computer-readable storage medium, a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps of the method of embodiment 1 are performed.

[0122] The steps involved in the apparatus of the above embodiments 3 and 4 correspond to the method embodiment 1, and the specific implementation method can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0123] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0124] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

[0125] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A surgical evaluation method based on multimodal representation and causal reasoning, characterized in that: include: Obtain surgical videos to be evaluated; Inputting the surgical video into a visual encoder for feature extraction to obtain visual features; Inputting the surgical video into a tool trajectory generation module to generate a tool motion trajectory and a position heat map sequence; Inputting the motion trajectory into a trajectory encoder for feature extraction to obtain trajectory features; The position heat map sequence and the visual features are fused and linearly mapped to obtain local visual features; the visual features are linearly mapped to obtain global visual features; The trajectory features, local visual features, and global visual features are input into a visual classification model for processing, wherein the visual classification model includes a block embedding module, an encoder, a parallel action recognition prediction branch, and a skill level assessment prediction branch connected in sequence, wherein the action recognition prediction branch outputs a predicted action type, and the skill level assessment prediction branch outputs a predicted skill level; The visual classification model is designed with a network structure based on causal reasoning, and the network structure based on causal reasoning is specifically as follows: the predicted action type is converted into action features; in the skill level assessment prediction branch, the encoder output features, the linear mapping layer output features in the action recognition prediction branch and the action features are spliced ​​and input into the linear mapping layer to obtain implicit features; Subsequently, the action features and the latent features are concatenated and input into the linear mapping layer and the classification layer in sequence, and finally the predicted skill level is output.

2. The surgical evaluation method based on multimodal representation and causal reasoning as claimed in claim 1, characterized in that: The visual encoder adopts a visual encoder based on convolutional neural network.

3. The surgical evaluation method based on multimodal representation and causal reasoning as claimed in claim 1, characterized in that: The tool trajectory generation module uses a surgical tool detector based on YOLOv8 to detect the operating tool, and generates a two-dimensional trajectory through the position of the detected operating tool boundary box, converts the two-dimensional trajectory into a three-dimensional trajectory based on a depth change estimation strategy, and generates a tool motion trajectory.

4. The surgical evaluation method based on multimodal representation and causal reasoning as claimed in claim 3, characterized in that: Based on the detected operating tool bounding box, the area where the operating tool is located is projected onto an empty background, thereby creating a position heat map sequence.

5. The surgical evaluation method based on multimodal representation and causal reasoning as claimed in claim 1, characterized in that: The trajectory encoder is built based on a bidirectional LSTM recurrent neural network.

6. A surgical evaluation system based on multimodal representation and causal reasoning, characterized in that: include: A video acquisition module is configured to: acquire a surgical video to be evaluated; A video processing module is configured to: input the surgical video into a visual encoder for feature extraction to obtain visual features; input the surgical video into a tool trajectory generation module to generate a tool motion trajectory and a position heat map sequence; The multimodal feature extraction module is configured to: input the motion trajectory into a trajectory encoder for feature extraction to obtain trajectory features; perform linear mapping after fusing the position heat map sequence and the visual features to obtain local visual features; perform linear mapping on the visual features to obtain global visual features; A joint prediction module is configured to: input the trajectory features, local visual features, and global visual features into a visual classification model for processing, wherein the visual classification model includes a block embedding module, an encoder, a parallel action recognition prediction branch, and a skill level assessment prediction branch connected in sequence, wherein the action recognition prediction branch outputs a predicted action type, and the skill level assessment prediction branch outputs a predicted skill level; The visual classification model is designed with a network structure based on causal reasoning, and the network structure based on causal reasoning is specifically as follows: the predicted action type is converted into action features; in the skill level assessment prediction branch, the encoder output features, the linear mapping layer output features in the action recognition prediction branch and the action features are spliced ​​and input into the linear mapping layer to obtain implicit features; Subsequently, the action features and the latent features are concatenated and input into the linear mapping layer and the classification layer in sequence, and finally the predicted skill level is output.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the surgical evaluation method based on multimodal representation and causal reasoning as described in any one of claims 1 to 5 are implemented.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the surgical assessment method based on multimodal representation and causal reasoning as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Visual SLAM front-end pose estimation method based on deep learning

    CN111127557A

  • System and method for the evaluation of or improvement of minimally invasive surgery skills

    US20140287393A1