A real-time effectiveness editing method and system for laparoscopic videos
By applying a lightweight model based on Transformer in laparoscopic surgical videos, combining parameter sharing and knowledge distillation technology, real-time effective editing of laparoscopic surgical videos is achieved, solving the problem of video editing in resource-limited equipment, and improving the efficiency and accuracy of surgical information acquisition.
Patent Information
- Application Number
- CN202310541482.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-15
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-05-15
AI Technical Summary
There is a large amount of video content in existing laparoscopic surgery videos that are not related to surgical operations, which affects the efficiency of doctors in finding key surgical steps, and deep neural networks are difficult to deploy and promote in hospital equipment with limited resources.
A network model based on Transformer is adopted, combining parameter sharing and knowledge distillation methods, a lightweight student model is designed, and real-time effective editing of laparoscopic surgical videos is achieved through video framing, annotation and preprocessing.
Real-time classification and editing of laparoscopic surgical videos in hospital equipment with limited resources is realized, the inference speed and classification accuracy are improved, invalid video clips are removed, and the consistency and efficiency of surgical information acquisition are improved.
Smart Images

Figure CN116630848B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video understanding, and particularly to a method and system for real-time effective editing of laparoscopic videos. Background Art
[0002] Laparoscopic surgery is a brand-new surgical technique different from traditional surgery. Besides the differences in concepts, equipment, and surgical instruments, it also has the particularity and complexity of technical operations. With the development and progress of laparoscopic surgical techniques in China over the past 30 years, laparoscopic techniques in China have developed from a difficult start to a booming development with the concerted efforts of many predecessors and successors. Currently, it has reached a relatively high level and has been popularized in hospitals at all levels. Laparoscopic surgery has become the future development trend of surgery, and is gradually forming a specialized discipline with distinct characteristics, which is regarded as another milestone in the history of surgical development. Its surgical indications include hepatobiliary system surgeries, gastrointestinal surgeries, urinary system surgeries, and gynecological disease surgeries, etc.
[0003] In terms of hepatobiliary system surgeries, with the continuous progress of equipment and technology, the application of laparoscopy in complex hepatobiliary surgeries has become increasingly widespread. Currently, laparoscopic liver segment and combined liver segment resection, and resection of huge liver cancers have become routine surgeries. Currently, all types of liver resections except liver transplantation can be completed, including laparoscopic resection of huge liver cancers, liver segment resection, combined sub-hepatic segment resection, liver resection combined with vascular reconstruction, and living donor liver resection, etc. Laparoscopy in cholecystectomy has also been rapidly promoted and popularized globally with its outstanding advantages such as small trauma and fast recovery, and has become the gold standard for the treatment of benign gallbladder diseases. Moreover, it has also received attention in the diagnosis and treatment of various gynecological diseases. As early as in the 1980s and 1990s of the 20th century, laparoscopic hysterectomy and laparoscopic radical hysterectomy and lymph node dissection were carried out in China. With the development of medical technology and industrial technology, laparoscopic techniques are still widely used in the diagnosis and treatment of gynecological diseases today. In terms of gastrointestinal surgeries, laparoscopic radical gastrectomy is currently the most widely used minimally invasive treatment technique for gastric cancer. Gastric cancer is one of the major diseases seriously threatening the lives and health of people around the world. Laparoscopic radical gastrectomy is a commonly used method for minimally invasive surgery of middle and late gastric cancer. The CLASS-02 study confirmed that for middle and late gastric cancer patients undergoing laparoscopic total gastrectomy, compared with traditional open surgery, the laparoscopic surgery time is longer, but the intraoperative bleeding is less; the surgical-related complications of laparoscopy and open surgery are similar. Endoscopic resection has become the preferred treatment method for early gastric cancer patients without the risk of lymph node metastasis. And the video of endoscopic surgery can be said to be a perfect reproduction of a surgery, containing rich surgical process information. With the rapid development of modern medical technology, effectively using endoscopic equipment and computer-aided technology can greatly improve the success rate of surgery.
[0004] The parsing of surgical video content is the basis of intelligent surgery. In a complete surgical video, video segments that record some cleaning actions on the camera during the operation or when the endoscope is placed aside during some operations outside the patient's body will be retained. However, after the operation, doctors actually only want to see effective surgical images. For example, in laparoscopic gastric cancer resection surgery, the operation usually lasts for 3-4 hours, with about 15 ineffective surgical images outside the body, each lasting about 20 seconds. The situation where the camera is placed aside for other reasons during the operation occurs about 2-4 times, with the duration of each time varying, ranging from less than 5 minutes to more than 20 minutes. In addition, there is a large amount of video content unrelated to actual surgical operations, which makes the subsequent parsing and reuse process more cumbersome, and at the same time makes the acquisition of surgical information incoherent, affecting the efficiency of doctors in finding key surgical steps. Therefore, the need to remove surgical segments unrelated to surgical operations from surgical videos after the operation is very urgent. Summary of the Invention
[0005] The technical problem to be solved by the present invention is:
[0006] In order to remove surgical segments unrelated to surgical operations from surgical videos and extract effective videos, the present invention provides a method and system for real-time effectiveness editing of laparoscopic videos. The present invention designs a network model based on Transformer. Transformer is a deep neural network with a core structure of self-attention mechanism. It was first applied to the machine translation task in the NLP field. Inspired by the amazing expressive ability of Transformer, researchers extended it to computer vision tasks. And it has greatly refreshed the previous records in object detection and semantic segmentation tasks, becoming the new mainstream of visual modeling. However, in hospital workstations or imaging devices, there is a certain degree of limited resources, which cannot meet the huge computing power and memory overhead required when using deep neural networks, which hinders their application and promotion, and overly complex models may exhibit overfitting phenomena.
[0007] For the above problems, the present invention adopts the methods of parameter sharing and knowledge distillation for the network model, compressing the number of parameters of the algorithm model by nearly 50% of the original, to reduce the difficulty of deploying the system in embedded devices in hospitals, and at the same time accelerating the inference speed to achieve real-time classification performance. At the same time, aiming at the accuracy degradation caused by parameter sharing, the method of knowledge distillation is used to transfer the knowledge learned in the pre-trained teacher model to the lightweight student model with a small loss of model accuracy, realizing the real-time effectiveness editing of laparoscopic videos.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0009] A real-time effectiveness editing method for laparoscopic videos, characterized by the following steps:
[0010] Step 1: Frame, label the category, and crop the videos in the laparoscopic surgery video dataset; divide the dataset into a training set, a validation set, and a test set;
[0011] Step 2: Perform supervised pre-training on the teacher network with training samples;
[0012] The selected teacher network is the network model DeiT based on Transformer. Compared with the original DeiT model structure that uses an additional cls_token for classification, in the present invention, the cls_token is removed, and instead, a global average pooling layer GAP and a fully connected layer are used at the end of the network to complete the classification of clear intra-cavity images, blurred intra-cavity images, and extra-cavity images. At the same time, the pos_token of the absolute position encoding in ViT is discarded and replaced with a relative position encoding method that is more flexible for changes in the input image size;
[0013] Step 3: Train the lightweight student model using knowledge distillation;
[0014] Transfer the knowledge learned in the pre-trained teacher model in Step 2 to the student model; calculate the cross-entropy loss for the predicted soft labels, attention feature maps, and hidden layer feature maps of the teacher model and the student model in pairs as the distillation loss, and optimize the distillation loss to make the student model closer to the prediction distribution of the teacher model;
[0015] The student model is implemented by sharing parameters in the feature extraction layer of the teacher network; two linear transformation matrices that do not share parameters are added to each layer;
[0016] Step 4: Predict the category of the preprocessed frame image using the distilled student model, and use the prediction result as the condition for whether to retain the frame, and save the video as video segments of multiple effective surgeries to achieve automatic editing of laparoscopic surgery videos.
[0017] A further technical solution of the present invention: The teacher network consists of three modules: an Embedding layer, a Transformer Encoder layer, and an MLP Head layer; the Embedding layer is used to transform the data; the Transformer Encoder layer is used to extract more abstract features in the intra- and extra-corporeal cavity views by repeatedly stacking Encoder Blocks multiple times; the MLP Head layer is used to achieve the classification of effectiveness in laparoscopic surgery videos.
[0018] A further technical solution of the present invention: The MLP Head layer consists of a global average pooling layer GAP and a fully connected layer.
[0019] A laparoscopic video real-time effectiveness editing system, characterized by including a video processing module and an intelligent classifier module.
[0020] The video processing module includes preprocessing and postprocessing of the video.
[0021] The intelligent classifier module predicts the category of the preprocessed video, and uses this prediction result as the condition for whether to retain this frame. Through the video processing module, the video is postprocessed and saved as video segments of multiple effective surgeries, realizing automatic editing of laparoscopic surgery videos.
[0022] A further technical solution of the present invention: The preprocessing of the video includes frame splitting, annotation, and cropping. After the video is preprocessed, it is divided into training samples, validation samples, and test samples. Each sample contains two categories: effective videos and ineffective videos.
[0023] A further technical solution of the present invention: The intelligent classifier module contains a lightweight Transformer network based on parameter sharing, and the performance of the network is improved through knowledge distillation training.
[0024] A further technical solution of the present invention: The category prediction values of each frame are obtained from multiple frame images passing through the intelligent classifier module, and it is determined whether to retain this frame according to this prediction value.
[0025] An application of a method for real-time effectiveness editing of laparoscopic videos, characterized in that after the operation, surgical segments irrelevant to the surgical operation are removed from the surgical video, and effective videos are extracted.
[0026] A computer system, characterized by including: one or more processors, and a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.
[0027] A computer-readable storage medium, characterized in that it stores computer-executable instructions, and the instructions are used to implement the above method when executed.
[0028] The beneficial effects of the present invention are as follows:
[0029] A method and system for real-time effectiveness editing of laparoscopic videos provided by the present invention can reach an inference speed of 88 FPS, solving the problems that existing laparoscopic surgery videos can only rely on manual analysis and cannot meet clinical needs, etc.
[0030] In view of the need to classify the effectiveness of laparoscopic surgery videos in real time, the present invention ingeniously introduces a neural network based on the self-attention mechanism to achieve an ideal intelligent editing effect for laparoscopic surgery videos. This is mainly for the scenario where embedded devices with limited resources and insufficient computing power are commonly used in hospitals. The present invention introduces a method of parameter sharing, which reduces the number of network parameters and improves the inference speed of the model. At the same time, the present invention introduces a knowledge distillation method, and the student model is trained by distillation through a pre-trained teacher network, so that the accuracy of the student model is significantly improved. Thanks to the above measures, the present invention can obtain real-time and accurate classification results in various laparoscopic scenarios, and finally achieve a very good editing effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The drawings are only for the purpose of illustrating specific embodiments and are not considered to be a limitation of the present invention. Throughout the drawings, the same reference numerals represent the same components.
[0032] Figure 1 Schematic diagram of the process of the present invention;
[0033] Figure 2 Schematic diagram of the data annotation process;
[0034] Figure 3 Network structure diagram of the present invention;
[0035] Figure 4 Schematic diagram of the model training process based on the self-attention mechanism network. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0037] In the present invention, the part of the video that is only in the patient's cavity and has a clear picture can be defined as a valid video. Other parts such as blurred pictures or pictures outside the body cavity are invalid videos, including multiple random removals of the endoscope for cleaning operations, bloodstain contamination of the lens, jitter blurred lens, mucosal tissue reflection lens, etc. The present invention designs a system to implement a lightweight and real-time high-quality classifier for distinguishing valid surgical videos (Valid) and invalid surgical videos (Invalid) by re-designing the parameter sharing and knowledge distillation algorithms for the Transformer-based network, so as to delete the invalid segments in the video and only leave the valid surgical segments to complete the editing of the video.
[0038] The present invention provides a preliminary intelligent editing method for laparoscopic surgery videos based on knowledge distillation. This method includes two modules, a video processing module and an intelligent classifier module. Video processing includes preprocessing and postprocessing of the video; the intelligent classifier module contains a lightweight Transformer network based on parameter sharing, and the performance of this network is improved through knowledge distillation training. First, the video processing module is used to frame the surgical video, and the framing frame rate is 1 frame per second. Secondly, the obtained multiple frame images are respectively input into the intelligent classifier module to obtain the predicted class values of each frame. Finally, the postprocessing part in the video processing module decides whether to retain this frame according to this predicted value. In this method, the knowledge distillation process of the classifier module for identifying the inside and outside views of the body cavity requires the use of a pre-trained teacher model, and the student model does not need pre-training and is updated online during the knowledge distillation process. The specific process is as follows:
[0039] (1) Data preparation: Data preprocessing is respectively performed on the laparoscopic surgery video dataset. After operations such as framing, annotating, and cropping the dataset, training samples, validation samples, and test samples are separated. Each sample contains two categories: valid videos and invalid videos.
[0040] (2) Use the training samples to perform supervised pre-training on the teacher network to obtain a classifier for identifying the inside and outside views of the body cavity with higher accuracy and complexity, which is used for subsequent knowledge distillation of the student network.
[0041] (3) The student model is implemented by sharing parameters of the feature extraction layer of the teacher network. The present invention additionally adds two linear transformation matrices that do not share parameters to each layer to increase complexity and improve the classification accuracy of valid and invalid videos in laparoscopic surgery videos.
[0042] (4) Use the method of knowledge distillation to train the lightweight student model. Transfer the knowledge learned in the pre-trained and more complex teacher model to the student model. Calculate the cross-entropy loss for the predicted soft labels, attention feature maps, and hidden layer feature maps of the teacher model and the student model pairwise as the distillation loss, and optimize the distillation loss to make the student model closer to the prediction distribution of the teacher model.
[0043] (5) During the inference process, first preprocess the video through the video processing module, predict the class of the processed frame images using the distilled student model, and use this prediction result as the condition for whether to retain this frame. Then, postprocess the video through the video processing module again, and save the video as video segments of multiple effective surgeries, finally realizing the automatic editing of laparoscopic surgery videos.
[0044] As Figure 1 shown, a preferred embodiment of the present invention is as follows:
[0045] Step 1: Label the categories of the videos in the laparoscopic surgery video dataset, as Figure 2 shown. First, frame the videos in the dataset, and downsample the frame images from the original resolution to 640×360. Second-by-second annotation is adopted for both datasets, that is, for a video with a frame rate of 25 FPS, one image is annotated every 25 frames, and for a video with a frame rate of 60 FPS, one image is annotated every 60 frames. Among them, the positive samples (valid videos) are the clear images inside the patient's cavity. The negative samples (invalid videos) are blurred images or images outside the cavity. This includes multiple random removals of the endoscope for cleaning operations, bloodstain-contaminated lenses, jitter-blurred lenses, mucosal tissue reflective lenses, etc. During training, the sample images used are uniformly scaled to a size of 224×224 by the video processing module. For each test frame image, no downsampling is required.
[0046] Step 2: Use the training samples to perform supervised pre-training on the teacher network to obtain a laparoscopic effective surgery video classifier with high accuracy and a relatively complex model, which is used for subsequent knowledge distillation of the student network. The present invention selects the network model DeiT based on Transformer as the teacher network. Compared with the structure of the original DeiT model, which uses an additional cls_token for classification, in the present invention, the cls_token is removed, and instead, a global average pooling layer (GAP) and a fully connected layer are used at the end of the network to complete the classification of clear images inside the cavity, blurred images inside the cavity, and images outside the cavity. At the same time, the pos_token of the absolute position encoding in ViT is discarded and replaced with a relative position encoding method that is more flexible for changes in the input image size. The teacher network is generally composed of three modules: the Embedding layer, the Transformer Encoder, and the MLP Head (the layer structure finally used for classification).
[0047] First, the picture is divided into slices of a fixed size, and the slice size is 16×16. Then, each image with a resolution of 224×224 will generate 196 slices. The slice length is 16×16×3 = 768. After passing through the linear mapping layer, the dimension is 196×768, which is used as the input of the network.
[0048] (1) Embedding layer. Since the attention module in the model requires the input to be a sequence of tokens (vectors), i.e., a two-dimensional matrix [num_token, token_dim]. First, a transformation is performed on the data through an Embedding layer. First, the frame image is divided into multiple slices according to the specified size. In the present invention, the scaled laparoscopic image (with a resolution of 224×224) is divided into slices of size 16×16, and after division, (224 / 16) 2 = 196 slices will be obtained. Then, each slice is mapped into a one-dimensional vector through a linear mapping. The size of each slice is 16×16×3, and a vector of length 768 is obtained through the mapping. In a specific implementation, it is achieved through a convolutional layer with a convolutional kernel size of 16×16, a stride of 16, and 768 convolutional kernels. Through convolution, the frame image dimension changes from [224, 224, 3] to [14, 14, 768], and then by flattening the two dimensions of H and W, it can be changed from [14, 14, 768] to [196, 768]. At this time, it exactly becomes a two-dimensional matrix, which is the input required by the attention layer. Then, a relative position encoding is added. For the slice x p after linear mapping to the embedding vector
[0049]
[0050] where represents the learnable vector added for the relative position of each slice during the encoding process. For the embedding vector z p generated on p = 1,…, N, it will be used as the input for the subsequent attention mechanism.
[0051] (2) Transformer Encoder layer. The implementation of this layer extracts more abstract features in the internal and external fields of view of the body cavity by repeatedly stacking Encoder Blocks multiple times, and it mainly consists of the following parts:
[0052] Regularization layer. Here, regularization processing is performed on each token. The calculation formula is as follows, where LN represents the batch normalization calculation output, x represents the input data, E[x] represents the mean of the x tensor, Var[x] represents the variance of the x tensor, ∈ represents a very small parameter to ensure that the denominator is not zero, and γ and β are learnable coefficients.
[0053]
[0054] The multi-head attention layer. The vector Q represents the query, which will subsequently be matched with each vector K. K represents the key, and V represents the value. In the subsequent process of matching Q and K, the correlation between the two is calculated. The greater the correlation, the greater the weight corresponding to V. The formula is as follows:
[0055]
[0056] MultiHead(Q, K, V) = Concat(head1, …, head h )W O
[0057]
[0058] The feed-forward network layer consists of a fully connected layer, an activation function, and Dropout. The first fully connected layer quadruples the number of input nodes, and the second fully connected layer restores the number of nodes to the original.
[0059] (3) MLP Head layer. The MLP Head layer here consists of a global average pooling layer (GAP) and a fully connected layer to achieve the classification of effectiveness in laparoscopic surgery videos.
[0060] Step 3: The student model in the present invention is implemented by sharing parameters in the feature extraction layer of the teacher network, as Figure 3 shown. First, the parameters of every two adjacent attention layers in the teacher network are shared, including the parameters of the multi-head self-attention layer and the feed-forward network layer:
[0061] Z i+1 = f(Z i ; θ), i = 0, …, L - 1
[0062] The present invention additionally adds two linear transformation matrices that do not share parameters to each layer to increase the complexity and improve the classification accuracy of the effectiveness of laparoscopic videos.
[0063]
[0064]
[0065] where M is the number of multi-head attentions, are the linear transformation matrices before and after softmax respectively, increasing the variance of the parameters between each attention module.
[0066] Step 4: The present invention uses the method of knowledge distillation, as Figure 3As shown, the knowledge information embedded in the pre-trained model is transferred to the lightweight model with shared weights in the form of soft labels. The method of knowledge distillation is specifically divided into three parts: output distillation, attention distillation, and hidden layer distillation.
[0067] Output distillation only transfers the classification results of the final laparoscopic video validity. The output form is the probability distribution of soft labels or the logical units of hard labels, classifying the laparoscopic video frames in difficult-to-distinguish scenarios. The objective function here uses the cross-entropy loss between the two prediction results, as shown in the formula:
[0068]
[0069] where z s and z t are the logical predictions of the student and teacher models respectively, and T is the temperature value controlling smoothness. In the present invention, T = 1 is set. CE represents the cross-entropy loss.
[0070] Attention distillation uses the attention feature maps in the Transformer model to guide the training of the student model, transferring the knowledge of the teacher model's attention to a certain part of the laparoscopic video frame to the student model. Specifically, in the multi-head attention mechanism, the cross-entropy loss is applied to the relationship between the query vector, key vector, and value vector. The same type of vectors in all attention heads are transformed into a matrix. For example, Q = [Q1,..., Q M , and the same applies to K and V. For the sake of symbol simplification, the symbols S1, S2, and S3 are used to represent Q, K, and V respectively. Then, 9 different relationship matrices defined by are generated. The self-attention distillation loss is expressed as follows.
[0071]
[0072] where R ij,n represents the nth row of R ij .
[0073] Output features through the MLP layer, generate the relationship matrix of the hidden layer in the MLP, and transfer the process of how the teacher model classifies the effectiveness of the intra- and extra-cavity images in the classification layer to the student model. Use to represent the hidden state of the Transformer layer. The hidden state distillation loss based on the relationship matrix is defined as:
[0074]
[0075] where R H,n represents the nth row of R H . Among them
[0076] Therefore, for the knowledge distillation training of the student model, it includes three dimensions: the distillation of the output results of the effectiveness of laparoscopic videos, the distillation of the attention positions of different pictures inside and outside the body cavity, and the distillation of the effectiveness classification process. The formula of the final distillation loss function is as follows:
[0077]
[0078] Among them, β and γ are hyperparameters, and the default values are 1 and 0.1 respectively.
[0079] Step 5: In the inference process, first preprocess the video through the video processing module, which is mainly implemented by the OpenCV library. Load the input video, read it frame by frame, and perform class prediction on each frame through the trained student network model. Use the prediction result as the condition for whether to retain the frame, and then post-process the video through the video processing module again. In the post-processing stage, enable the video writer every time a valid video is judged, start writing, and end writing until an invalid video is encountered. In this way, the writing of a valid video is completed. And so on, perform this operation on the entire video to achieve the editing of the entire laparoscopic video. Finally, the automatic editing of the surgical video is realized.
[0080] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the scope of the technology disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention.
Claims
1. A real-time validity clip method for laparoscopic videos, characterized in that The steps are as follows: Step 1: Frame the videos in the laparoscopic surgery video dataset, label the categories, and crop them; divide the dataset into a training set, a validation set, and a test set; Step 2: Conduct supervised pre-training on the teacher network with the training samples; The described teacher network selects the network model DeiT based on Transformer. Compared with the structure of the original DeiT model that classifies through an additional cls_token, the cls_token is removed, and instead, a global average pooling layer GAP and a fully connected layer are used at the end of the network to complete the classification of clear intra-cavity images, blurred intra-cavity images, and extra-cavity images. At the same time, the pos_token of the absolute position encoding in ViT is discarded and replaced with a relative position encoding method that is more flexible for changes in the input image size; Step 3: Train the lightweight student model using knowledge distillation; Transfer the knowledge learned in the pre-trained teacher model in Step 2 to the student model; calculate the cross-entropy loss for the prediction soft labels, attention feature maps, and hidden layer feature maps of the teacher model and the student model in pairs respectively as the distillation loss, and optimize the distillation loss to make the student model closer to the prediction distribution of the teacher model; The described student model is implemented by sharing parameters in the feature extraction layer of the teacher network; two linear transformation matrices that do not share parameters are added to each layer; Step 4: Predict the categories of the preprocessed frame images using the distilled student model, and use this prediction result as the condition for whether to retain the frame, and save the video as video segments of multiple effective surgeries to achieve automatic editing of laparoscopic surgery videos.
2. The laparoscopic video real-time effectiveness editing method according to claim 1, wherein: The described teacher network consists of three modules: Embedding layer, Transformer Encoder layer, and MLP Head layer; the Embedding layer is used to transform the data; the Transformer Encoder layer is used to extract more abstract features in the intra- and extra-corporeal cavity views by repeating the stacking of Encoder Blocks multiple times; the MLP Head layer is used to achieve the classification of the effectiveness in laparoscopic surgery videos.
3. The laparoscopic video real-time effectiveness editing method according to claim 2, wherein: The MLP Head layer consists of a global average pooling layer GAP and a fully connected layer.
4. A system for implementing the laparoscopic video real-time effectiveness editing method according to claim 1, characterized in that It includes a video processing module and an intelligent classifier module, The described video processing module includes preprocessing and postprocessing of the video; The described intelligent classifier module predicts the categories of the preprocessed video, and uses this prediction result as the condition for whether to retain the frame. Through the video processing module, postprocess the video and save the video as video segments of multiple effective surgeries to achieve automatic editing of laparoscopic surgery videos.
5. The laparoscopic video real-time effectiveness editing system according to claim 4, characterized in that The preprocessing of the video includes framing, labeling, and cropping. After preprocessing, the video is divided into training samples, validation samples, and test samples, and each sample contains two categories: effective videos and ineffective videos.
6. The laparoscopic video real-time effectiveness editing system according to claim 4, characterized in that The described intelligent classifier module includes a lightweight Transformer network based on parameter sharing, and the performance of the network is improved through knowledge distillation training.
7. The laparoscopic video real-time effectiveness editing system according to claim 4, characterized in that The post-processing of the video is as follows: the predicted class values of each frame are obtained from multiple frame images of the intelligent classifier module, and it is determined whether the frame needs to be retained according to the predicted value.
8. Application of the laparoscopic video real-time effectiveness editing method according to claim 1, characterized in that, After the operation, surgical segments unrelated to the surgical operation are removed from the surgical video, and the effective video is extracted.
9. A computer system, characterized in that Including: One or more processors, a computer-readable storage medium for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in claim 1.
10. A computer-readable storage medium, characterized in that Stored with computer-executable instructions that are used to implement the method described in claim 1 when executed.
Citation Information
Patent Citations
Dynamic behavior identification system based on key clip discrimination
CN114743263A
Compression method for laparoscope video based on ViT-Slim class
CN115526943A