Laparoscopic surgery full-granularity recognition system and method based on feature extraction and task segmentation

Through the laparoscopic full-grain recognition system of feature extraction and task segmentation, the multi-frame spatial feature extraction, single-frame feature restoration and temporal feature fusion modules are used to solve the computing resource problem of multi-task recognition, realize the efficient output of multi-grained surgical scenario information, and improve the application of intelligent algorithms in surgical operations.

CN120451578APending Publication Date: 2025-08-08THE AFFILIATED HOSPITAL OF QINGDAO UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510595324.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing surgical scenario understanding method requires the construction of multiple deep learning models to identify single-grained surgical scenarios, which leads to excessive computing resources and the inability to use each other between tasks, hindering the application process of intelligent algorithms in surgical operations.

Method used

A laparoscopic surgery full-grain recognition system based on feature extraction and task segmentation is adopted. Through a multi-frame spatial feature extraction module, a single-frame feature reduction module and a multi-frame time feature fusion module, an encoder-decoder-form image segmentation deep learning model is used to realize unified recognition of surgical stage, surgical steps, surgical action ternary body and surgical scene segmentation.

Benefits of technology

It reduces the use of computing resources, and outputs multiple granularity surgical scenario information at one time through multi-task coordination, improving the application efficiency of intelligent algorithms in surgical operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451578A_ABST
    Figure CN120451578A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical information processing, and relates to a laparoscopic surgery full-granularity recognition system and method based on feature extraction and task segmentation. Firstly, a video stream composed of multiple frames of images of an endoscope is input into the system, and then the spatial feature of each frame of image is obtained through a multi-frame spatial feature extraction module; the spatial features of all the frames are then input into a multi-frame time feature fusion module to obtain time features fused with different time lengths, and then the time features are used for outputting operation stages, operation steps and operation ternary body labels; meanwhile, only the spatial features of the current frame are extracted from the spatial features of all the frames, and an operation scene segmentation label is output through a single-frame feature restoration module. According to the laparoscopic surgery full-granularity recognition system and method based on feature extraction and task segmentation, computing resources can be greatly reduced, and compared with an existing single-granularity surgery scene technology, the laparoscopic surgery full-granularity recognition system and method based on feature extraction and task segmentation can output surgery scene information of multiple granularities at a time by means of the mutual coordination effect of multiple tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical information processing, and relates to a full-granularity recognition system and method for laparoscopic surgery based on feature extraction and task segmentation. Background Art

[0002] The surgical scene includes location information such as the status and position of surgical instruments and patient tissue, as well as action information such as the surgeon's surgical actions. Understanding the real-time surgical scene during surgery can empower surgical robots with intelligence, enabling them to perceive the surgical scene and make decisions about their next steps. It can also help surgeons learn surgical procedures, optimize surgical quality, and predict potential surgical risks.

[0003] Surgical scene understanding can be categorized into four areas, from coarse to fine granularity, including surgical stage recognition, surgical step recognition, surgical action triad recognition, and surgical scene segmentation. Existing surgical scene understanding methods often only recognize one granularity, such as surgical steps or surgical scene segmentation. To recognize all four granularities, four deep learning models must be built to identify surgical scenes at a single granularity and then combined. Multiple deep learning models consume significant computing resources, and the independent tasks prevent them from leveraging each other. This significantly hinders the application of intelligent algorithms in surgical procedures. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems in the prior art and provide a full-granularity recognition system and method for laparoscopic surgery based on feature extraction and task segmentation. Only one network can be used to identify tasks of four granularities: surgical stage, surgical steps, surgical action ternary, and surgical scene segmentation; it solves the problem of how to effectively coordinate the four tasks of different granularities and how the network can take into account the four tasks during the training process.

[0005] The technical solution provided by the present invention is a full-granularity recognition system for laparoscopic surgery based on feature extraction and task segmentation, which consists of a multi-frame spatial feature extraction module, a single-frame feature restoration module, and a multi-frame temporal feature fusion module; wherein the multi-frame spatial feature extraction module is used to extract the spatial features of each frame of the endoscopic video stream composed of the input multi-frame images; the single-frame feature restoration module is used to restore the spatial features of the single frame to the original image size through upsampling and generate surgical scene segmentation labels; the multi-frame temporal feature fusion module is used to extract temporal features of different time lengths from the multi-frame spatial features, and fuse them to output labels of the surgical stage, surgical steps and surgical triads; the multi-frame spatial feature extraction module and the single-frame feature restoration module adopt an image segmentation deep learning model in the form of an encoder-decoder.

[0006] Preferably, the training process of the encoder-decoder form of the image segmentation deep learning model is carried out according to the following steps: training of the surgical stage recognition task, using the loss of the surgical stage as the total loss function of the model to train the model until the surgical stage recognition task converges; training of the surgical step recognition task, using the sum of the losses of the surgical stage and the surgical step as the total loss function of the model to train the model until the surgical step recognition task converges; training of the surgical triplet recognition task, using the sum of the losses of the surgical stage, surgical step and surgical triplet as the total loss function of the model to train the model until the surgical triplet recognition task converges; training of the surgical scene segmentation task, using the sum of the losses of the surgical stage, surgical step, surgical triplet and surgical scene segmentation as the total loss function of the model to train the model until the surgical scene segmentation task converges.

[0007] Preferably, the multi-frame spatial feature extraction module is selected from any one of the encoders of Unet, Swin-Unet, SAM, and SAM2.

[0008] Preferably, the single-frame feature restoration module is selected from any one of the decoders of Unet, Swin-Unet, SAM, and SAM2.

[0009] The present invention also provides a full-granularity recognition method for laparoscopic surgery based on feature extraction and task segmentation, which includes the following steps: (1) Spatial feature extraction The endoscopic video stream composed of multiple frames of images is input into the multi-frame spatial feature extraction module to extract the spatial features of each frame of image; (2) Surgical scene segmentation label extraction Extract the spatial features of the current frame from all the frame spatial features output by the multi-frame spatial feature extraction module, and input them into the single-frame feature restoration module to output the surgical scene segmentation label; (3) Temporal feature extraction and fusion The multi-frame spatial features output from the multi-frame spatial feature extraction module are input to the multi-frame temporal feature fusion module to extract and fuse temporal features.

[0010] Preferably, the temporal feature extraction and fusion specifically include: (1) The multi-frame spatial features are first passed through the average pooling layer to obtain the pooled multi-frame spatial features; (2) The pooled multi-frame spatial features are repeated 4 times, where the first 3 repeated multi-frame spatial features are combined with 3 different class labels respectively, and the last repeated multi-frame spatial features are not processed; the class label is used to represent the temporal features later; (3) The multi-frame spatial features obtained in step (2) are respectively passed through different mask layers to obtain masked multi-frame spatial features; the mask layers include an empty mask layer, a mask layer of 1 / 2 input frame length, a mask layer of 3 / 4 input frame length, and a mask layer of other input frame lengths except the current frame; (4) The masked multi-frame spatial features are respectively subjected to the self-attention layer to extract temporal features; the class label is extracted as the temporal feature extracted from the multi-frame spatial features; the temporal features include the temporal features of the complete time series length, the temporal features of the total length of the previous frame being 1 / 2 of the input frames, the temporal features of the total length of the previous frame being 1 / 4 of the input frames, and the spatial features of the current frame; (5) The extracted time features are spliced together through the connection operation and fused into the final time features; (6) The final temporal features are processed through four independent linear layers to obtain the final surgical stage label, surgical step label, surgical triplet label, and surgical scene segmentation label.

[0011] The full-granularity recognition system and method for laparoscopic surgery based on feature extraction and task segmentation provided by the present invention can significantly reduce computing resources and utilize the mutual coordination of multiple tasks. Compared with the existing single-granularity surgical scene technology, it can output surgical scene information of multiple granularities at one time. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a flowchart of the operation of a full-granularity recognition system for laparoscopic surgery based on feature extraction and task segmentation in an embodiment of the present invention; Figure 2 This is the operation flow chart of the multi-frame time feature fusion module. DETAILED DESCRIPTION

[0013] The present invention is further described below with reference to specific embodiments and the accompanying drawings. More details are set forth in the following description to facilitate a full understanding of the present invention. However, the present invention can obviously be implemented in a variety of other ways different from the description. For those skilled in the art, any replacement, improvement or transformation made to the embodiments of the present invention is within the scope of protection of the present invention, and the scope of protection of the present invention should not be limited by the content of this specific embodiment.

[0014] The present invention provides a full-granularity recognition system for laparoscopic surgery based on feature extraction and task segmentation, such as Figure 1 As shown in the figure, it consists of a multi-frame spatial feature extraction module, a single-frame feature restoration module, and a multi-frame temporal feature fusion module. The four granularity tasks share a multi-frame spatial feature extraction module, and then pass through the single-frame feature restoration module and the multi-frame temporal feature fusion module to obtain the four granularity task outputs.

[0015] The multi-frame spatial feature extraction module is used to extract the spatial features of each frame image. The input is the endoscopic video stream composed of multiple frames of images, and the output is the spatial features of each frame image.

[0016] The single-frame feature restoration module is used to restore the spatial features of a single frame to the original image size through upsampling and generate surgical scene segmentation labels. The input is the spatial features of a single-frame image, and the output is the surgical scene label of the single-frame image.

[0017] The multi-frame temporal feature fusion module is used to extract temporal features of different time lengths from multi-frame spatial features, fuse them, and then output the labels of surgical stages, surgical steps, and surgical ternaries.

[0018] The multi-frame spatial feature extraction module and the single-frame feature restoration module can be any encoder-decoder deep learning model for image segmentation, such as Unet, Swin-Unet, SAM, SAM2, etc. The multi-frame spatial feature extraction module is the encoder part of the deep learning model for image segmentation, and the single-frame feature restoration module is the decoder part of the deep learning model for image segmentation.

[0019] Figure 1 This is the entire process of the present invention. First, the video stream composed of multiple frames of endoscopic images is input into the system of the present invention. Then, the spatial features of each frame of the image are obtained through the multi-frame spatial feature extraction module. The spatial features of all frames are then input into the multi-frame temporal feature fusion module to obtain the fusion of temporal features of different time lengths, which are then used to output the surgical stage, surgical step, and surgical ternary label. At the same time, only the spatial features of the current frame are extracted from the spatial features of all frames, and the surgical scene segmentation label is output through the single-frame feature restoration module.

[0020] (1) Spatial feature extraction The endoscopic video stream consisting of multiple frame images is input into the multi-frame spatial feature extraction module to extract the spatial features of each frame image.

[0021] (2) Surgical scene segmentation label extraction Only the spatial features of the current frame are extracted from all the frame spatial features output by the multi-frame spatial feature extraction module, and fed into the single-frame feature restoration module to output the surgical scene segmentation label.

[0022] (3) Temporal feature extraction and fusion The multi-frame spatial features output from the multi-frame spatial feature extraction module are fed into the multi-frame temporal feature fusion module to extract and fuse temporal features. Figure 2 As shown: 3.1 Multi-frame spatial features (a) are first passed through the average pooling layer to obtain the pooled multi-frame spatial features (b). The average pooling layer can reduce the dimensionality of the multi-frame spatial features (a), retaining only the most representative spatial features, thereby reducing the computational complexity of subsequent temporal feature processing.

[0023] 3.2 The multi-frame spatial feature (b) is repeated four times for subsequent processing. The first three repeated multi-frame spatial features (b) are combined with three different class labels to form multi-frame spatial features (c1), (c2), and (c3). The last repeated multi-frame spatial feature (b) remains unprocessed and becomes multi-frame spatial feature (c4). The class labels are used to represent temporal features.

[0024] 3.3 Multi-frame spatial features (c1), multi-frame spatial features (c2), multi-frame spatial features (c3) and multi-frame spatial features (c4) are respectively obtained through different mask layers to obtain masked multi-frame spatial features (d1), (d2), (d3) and (d4). Specifically: the mask layer of the multi-frame spatial feature (c1) is an empty mask layer, so the multi-frame spatial feature (d1) retains the spatial features of all frames; the mask layer of the multi-frame spatial feature (c2) masks the time series of 1 / 2 length, so the multi-frame spatial feature (d2) retains the spatial features of the current frame and the total length of 1 / 2 input frames before the current frame; the mask layer of the multi-frame spatial feature (c3) masks the time series of 3 / 4 length, so the multi-frame spatial feature (d3) retains the spatial features of the current frame and the total length of 1 / 4 input frames before the current frame; the mask layer of the multi-frame spatial feature (c4) masks the time series of multiple frames except the current frame, so the multi-frame spatial feature (d4) only retains the spatial features of the current frame; therefore, the multi-frame spatial features (d1), (d2), (d3) and (d4) respectively retain the spatial feature information of all frames, the total length of 1 / 2 input frames before the current frame, the total length of 1 / 4 input frames before the current frame and the current frame.

[0025] 3.4 Multi-frame spatial features (d1), (d2), and (d3) are respectively subjected to the self-attention layer for temporal feature extraction. The self-attention layer is divided into seven steps: Step 1: Add position encoding to the multi-frame spatial features (d1), (d2), and (d3) respectively; Step 2: Divide the result of the first step into multiple heads and perform subsequent processing on each head; Step 3: Multiply the features of each head by the learnable weights to generate the query matrix Q, key matrix K, and value matrix V respectively; Step 4: Calculate the similarity or correlation between the query matrix Q and the key matrix K; Step 5: Normalize the original scores from step 4 to obtain weight coefficients, and perform weighted summation on the value matrix V according to the weight coefficients. The calculation formula for each head's self-attention is: ; Among them, X is the feature of each head, W Q 、W K and W V is a learnable weight matrix, and the query matrix Q = W Q X, key matrix K = W K X, value matrix V = W V X,d k represents the dimension of matrix Q, softmax represents the softmax function; Step 6: Concatenate the self-attention outputs of each head together through a concatenation operation and multiply them by the learnable weights to obtain the result of the multi-head self-attention. Step 7: Input the results of multi-head self-attention into the multi-layer perceptron for further processing to obtain the final time features.

[0026] All three self-attention layers of the present invention share parameters, which can greatly improve computational efficiency and avoid overfitting.

[0027] 3.5 After the multi-frame spatial features (d1), (d2), and (d3) pass through the self-attention layer, the present invention only uses their respective class labels as the temporal features (e1), (e2), and (e3) extracted from the multi-frame spatial features (d1), (d2), and (d3). At the same time, the present invention retains the spatial features of the current frame in the multi-frame spatial features (d4) as (e4). Temporal features (e1), (e2), (e3), and (e4) respectively represent the temporal features of the complete time series length, the temporal features of the total length of 1 / 2 input frames before the current frame, the temporal features of the total length of 1 / 4 input frames before the current frame, and the spatial features of the current frame.

[0028] 3.6 Temporal features (e1), (e2), (e3), and spatial features (e4) are concatenated and fused into the final temporal feature (f). Temporal feature (f) combines the temporal features of the entire time series, the temporal features of the current frame with a total length of 1 / 2 the number of input frames, the temporal features of the current frame with a total length of 1 / 4 the number of input frames, and the spatial features of the current frame.

[0029] 3.7 The temporal feature (f) passes through four independent linear layers to obtain the final surgical stage label, surgical step label, surgical triplet label and surgical scene segmentation label.

[0030] Since the present invention outputs surgical scene information of four different granularities simultaneously through a network, it is difficult to converge the training at the same time, and the construction of the four task data sets is difficult to achieve. Therefore, the present invention constructs a proprietary model training method, specifically: (1) First train the surgical stage recognition task The surgical phase recognition task is the simplest, and the dataset is the easiest to create, so it is trained first. We first assign large initial learning rates to the multi-frame spatial feature extraction module and the multi-frame temporal feature fusion module. We use the surgical phase loss as the overall model loss function and begin training the model until convergence is achieved for the surgical phase recognition task.

[0031] (2) Secondly, training the surgical step recognition task The surgical step recognition task is relatively modest, and the dataset preparation is also relatively modest. Therefore, after the surgical stage recognition task has converged, training for the surgical step recognition task begins. First, a small initial learning rate is assigned to the multi-frame spatial feature extraction module and the multi-frame temporal feature fusion module, while a larger initial learning rate is assigned to the linear layer of the surgical step task. The sum of the surgical stage and surgical step losses is used as the overall model loss function, and model training begins until convergence is achieved for the surgical step recognition task.

[0032] (3) Retraining the surgical triplet recognition task The surgical step recognition task is relatively difficult, and dataset preparation is also challenging. Therefore, after achieving convergence in both the surgical stage and step recognition tasks, training for the surgical triplet recognition task begins. First, a small initial learning rate is assigned to the multi-frame spatial feature extraction module and the multi-frame temporal feature fusion module, while a larger initial learning rate is assigned to the linear layer for the surgical triplet task. The model training begins with the sum of the losses for the surgical stage, surgical step, and surgical triplet as the overall loss function, until convergence in the surgical triplet recognition task is achieved.

[0033] (4) Finally, train the surgical scene segmentation task Surgical scene segmentation is the most challenging task, and dataset preparation is also the most challenging. Therefore, after achieving convergence in the surgical stage recognition, surgical step recognition, and surgical triad tasks, training for the surgical scene segmentation task begins. Initially, the multi-frame spatial feature extraction module and the multi-frame temporal feature fusion module are given a small initial learning rate, while the single-frame feature restoration module is given a larger initial learning rate. The sum of the losses for surgical stage, surgical step, surgical triad, and surgical scene segmentation is used as the overall model loss function. Model training begins until convergence in the surgical scene segmentation task.

Claims

1. A fully granular recognition system for laparoscopic surgery based on feature extraction and task segmentation, characterized by: The system consists of a multi-frame spatial feature extraction module, a single-frame feature restoration module, and a multi-frame temporal feature fusion module; wherein the multi-frame spatial feature extraction module is used to extract the spatial features of each frame image from the input endoscopic video stream composed of multiple frames of images; the single-frame feature restoration module is used to restore the spatial features of a single frame to the original image size through upsampling and generate surgical scene segmentation labels; the multi-frame temporal feature fusion module is used to extract temporal features of different time lengths from the multi-frame spatial features, and fuse them to output labels of surgical stages, surgical steps, and surgical triads; the multi-frame spatial feature extraction module and the single-frame feature restoration module adopt an image segmentation deep learning model in the form of an encoder-decoder.

2. The full-granularity recognition system for laparoscopic surgery based on feature extraction and task segmentation according to claim 1 is characterized by: The training process of the image segmentation deep learning model is carried out according to the following steps: training of the surgical stage recognition task, using the surgical stage loss as the total loss function of the model to train the model until the surgical stage recognition task converges; For the surgical step recognition task, the model is trained with the sum of the losses of the surgical stage and surgical steps as the total loss function of the model until the surgical step recognition task converges. For the surgical triplet recognition task, the model is trained with the sum of the losses of the surgical stage, surgical steps, and surgical triplet as the total loss function of the model until the surgical triplet recognition task converges. For the training of the surgical scene segmentation task, the model is trained with the losses of surgical stage, surgical steps, surgical triples, and surgical scene segmentation as the total loss function of the model until the surgical scene segmentation task converges.

3. The full-granularity recognition system for laparoscopic surgery based on feature extraction and task segmentation according to claim 1 is characterized by: The multi-frame spatial feature extraction module is selected from any one of the encoders of Unet, Swin-Unet, SAM, and SAM2.

4. The full-granularity recognition system for laparoscopic surgery based on feature extraction and task segmentation according to claim 1 is characterized by: The single-frame feature restoration module is selected from any one of the decoders of Unet, Swin-Unet, SAM, and SAM2.

5. A full-granularity recognition method for laparoscopic surgery based on feature extraction and task segmentation, using the system according to any one of claims 1 to 4, characterized in that: The following steps are involved: (1) Spatial feature extraction The endoscopic video stream composed of multiple frames of images is input into the multi-frame spatial feature extraction module to extract the spatial features of each frame of image; (2) Surgical scene segmentation label extraction Extract the spatial features of the current frame from all the frame spatial features output by the multi-frame spatial feature extraction module, and input them into the single-frame feature restoration module to output the surgical scene segmentation label; (3) Temporal feature extraction and fusion The multi-frame spatial features output from the multi-frame spatial feature extraction module are input to the multi-frame temporal feature fusion module to extract and fuse temporal features.

6. The full-granularity recognition method for laparoscopic surgery based on feature extraction and task segmentation according to claim 5 is characterized in that: The temporal feature extraction and fusion specifically include: (1) The multi-frame spatial features are first passed through the average pooling layer to obtain the pooled multi-frame spatial features; (2) The pooled multi-frame spatial features are repeated 4 times, where the first 3 repeated multi-frame spatial features are combined with 3 different class labels respectively, and the last repeated multi-frame spatial features are not processed; the class label is used to represent the temporal features later; (3) The multi-frame spatial features obtained in step (2) are respectively passed through different mask layers to obtain masked multi-frame spatial features; the mask layers include an empty mask layer, a mask layer of 1 / 2 input frame length, a mask layer of 3 / 4 input frame length, and a mask layer of other input frame lengths except the current frame; (4) The masked multi-frame spatial features are respectively subjected to the self-attention layer to extract temporal features; the class label is extracted as the temporal feature extracted from the multi-frame spatial features; the temporal features include the temporal features of the complete time series length, the temporal features of the total length of the previous frame being 1 / 2 of the input frames, the temporal features of the total length of the previous frame being 1 / 4 of the input frames, and the spatial features of the current frame; (5) The extracted time features are spliced together through the connection operation and fused into the final time features; (6) The final temporal features are processed through four independent linear layers to obtain the final surgical stage label, surgical step label, surgical triplet label, and surgical scene segmentation label.

Citation Information

Cited By

  • Structured task modeling method and system for surgical procedures

    CN122677069A