Depth generation model training method and device for operation video data compression

Through the deep generation model training method, combined with multi-task joint training of spatiotemporal feature extractor, stage classifier and decoder, the problem of quality loss of surgical video at high compression rates is solved, efficient compression and reconstruction is achieved, decoding efficiency and detail restoration capabilities are improved, and high-standard application needs are adapted to the medical field.

CN120014506APending Publication Date: 2025-05-16FOURTH MILITARY MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510006502.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing surgical video compression technology is difficult to balance the compression rate and video quality at high compression rates, resulting in the loss of key details, low decoding efficiency, limited detail restoration capabilities, and unable to meet the high-standard application needs in the medical field.

Method used

The deep generation model training method is adopted, and the multi-task joint training of the spatiotemporal feature extractor, stage classifier and decoder is obtained to obtain the trained surgical video compression model to achieve key video details and efficient decoding at high compression rates.

Benefits of technology

The balance between high compression rate and high quality is achieved, the decoding efficiency is improved, the detailed restoration capability is significantly improved, and the optimal balance between visual quality and detailed performance is achieved under extremely small storage volumes, which is suitable for high-standard application needs in the medical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014506A_ABST
    Figure CN120014506A_ABST
Patent Text Reader

Abstract

The invention discloses a depth generation model training method and device for operation video data compression. The depth generation model training method for operation video data compression comprises the following steps: acquiring a video data set; an operation video compression model is obtained, the operation video compression model comprises a spatial-temporal feature extractor, a stage classifier and a decoder, the spatial-temporal feature extractor is used for extracting spatial features of operation video data frames and fusing the spatial-temporal features, the stage classifier is used for receiving video spatial features and video spatial-temporal features as input, and the decoder is used for decoding the video spatial features and the video spatial-temporal features; a stage classification result is obtained by further fusing the spatial features and the time features, and the decoder is used for receiving the video spatio-temporal features as input and decoding the video spatio-temporal features back to the video frame data; and training the operation video compression model through the video data set to obtain a trained operation video compression model. According to the invention, the balance between high compression ratio and high quality can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video compression technology, and in particular to a deep generative model training method for surgical video data compression and a deep generative model training device for surgical video data compression. Background Art

[0002] With the development of medical imaging technology, surgical video data plays an important role in medical teaching, remote surgical guidance, and postoperative analysis. Especially in high-precision operations such as minimally invasive surgery and endoscopic surgery, video recording can provide valuable support for postoperative analysis and medical quality improvement. However, surgical video data usually has high resolution, high frame rate, and long recording time, and the amount of data is very large. Although traditional video compression storage technology can reduce the storage volume to a certain extent, it will lead to the loss of video details under high compression rates, thus affecting the effectiveness and accuracy of surgical videos in the medical field.

[0003] Traditional compression algorithms (such as H.264, H.265, etc.) rely on technologies such as motion estimation, inter-frame prediction, and transform coding to effectively reduce redundant data. However, when these algorithms are applied to surgical videos at extremely high compression rates, important image details such as tiny blood vessels and tissue textures may be lost, which is extremely unfavorable for the fidelity presentation of key details in surgical videos. In order to meet the visual quality requirements of surgical videos after compression, current research has begun to introduce deep generative models to enhance the ability to restore surgical videos during the compression process. By learning subtle features in the video, deep generative models can reconstruct realistic images at high compression rates, meeting the strict requirements of the medical field for video details. Deep generative models such as generative adversarial networks (GANs) and variational autoencoders (VAEs) have potential in the field of video reconstruction and compression. Such models can generate images close to the original quality at high compression rates by learning the features of a large amount of surgical video data. For example, GAN achieves image detail completion and reconstruction through a generator-discriminator architecture. The generator is responsible for generating high-quality images, and the discriminator is used to evaluate the authenticity of the generated images, thereby achieving detail restoration after compression. VAE distributes the key features of surgical videos in a low-dimensional latent space through the probability model of encoding and decoding, and then reconstructs high-quality images through the decoder. These methods can improve the quality loss problem of traditional compression technology at extremely high compression rates, but GAN and VAE have certain limitations in terms of decoding efficiency and detail fidelity, especially when processing large-resolution surgical videos in real time, the computational cost is high. Some solutions combine traditional compression algorithms with deep generative models, using the efficiency of traditional compression algorithms to process most of the data in surgical videos, while using deep generative models to reconstruct and enhance key detail areas. This type of method usually applies deep generative models to key frames or important areas of the video to ensure that the core details of the surgical video are preserved while reducing the overall storage volume. For example, after the video is initially compressed using H.265 encoding, GAN is used to generate high-quality key frame images, thereby improving the compression effect without significantly increasing the computational complexity. However, this solution still faces the problem of high resource consumption when real-time decoding and detail reconstruction are required, and it is difficult to meet the high efficiency requirements of surgical videos in fast recovery and real-time processing. Self-supervised learning uses the temporal and spatial information of the video to learn features in surgical videos and generate adaptive compression models without the need for labels. This type of technology can adaptively and selectively compress unimportant areas during video compression and retain key areas. This type of method is expected to improve the compression rate without affecting key details of the video, but the applicability of this technology in complex surgical scenarios is still under verification, and the training and optimization process of the model is complex, which faces challenges in real-time applications.

[0004] The current surgical video compression scheme has certain limitations in terms of compression rate, detail restoration, real-time processing capabilities, etc., and has not yet fully met the high requirements of surgical videos in medical scenarios. The present invention aims to use a deep generative model to store surgical video data at a high compression rate, while restoring key video details at the decoding end, ensuring the optimal balance between visual quality and detail performance in a very small storage volume. In this way, the present invention can achieve efficient compression and reconstruction of surgical videos, solve the quality loss problem of traditional methods, and adapt to the high-standard application requirements of digital storage and transmission of surgical videos in the medical field.

[0005] The disadvantages of the prior art are as follows:

[0006] 1. Difficulty in balancing compression rate and quality: Traditional compression algorithms often lead to significant degradation of video quality at high compression rates, especially in surgical videos, where the loss of key details may affect medical judgment. Although existing methods based on generative models have improved the fidelity of some details, it is still difficult to balance compression rate and restoration quality at extremely high compression rates, and cannot fully preserve the key visual details of surgical videos.

[0007] 2. Low decoding and processing efficiency: Existing compression methods based on deep generative models require a lot of computing resources when decoding and reconstructing, especially in the real-time processing scenario of high-resolution surgical videos, which may lead to unsatisfactory processing speed and make it difficult to meet the needs of real-time decoding and restoration. This problem of low decoding efficiency is particularly disadvantageous in actual surgical guidance or telemedicine environments.

[0008] 3. Limited ability to restore details: Existing generative models are still insufficient in restoring details in surgical videos, especially for key parts such as tiny blood vessels and tissue textures, which may affect the visual quality and medical effectiveness of surgical videos. These methods may cause blurring or artifacts at extremely high compression rates, resulting in substandard video quality.

[0009] Therefore, it is hoped that a technical solution can be provided to solve or at least alleviate the above-mentioned deficiencies of the prior art. Summary of the invention

[0010] The purpose of the present invention is to provide a deep generative model training method for surgical video data compression to solve at least one of the above-mentioned technical problems.

[0011] The present invention provides the following scheme:

[0012] According to one aspect of the present invention, a deep generative model training method for surgical video data compression is provided, and the deep generative model training method for surgical video data compression comprises:

[0013] Step 1: Get the video dataset;

[0014] Step 2: Obtain a surgical video compression model, which includes a spatiotemporal feature extractor, a stage classifier, and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and fuse spatiotemporal features of surgical video data frames, the stage classifier is used to receive video spatial features and video spatiotemporal features as input, and obtain a stage classification result by further fusing spatial features and temporal features, and the decoder is used to receive video spatiotemporal features as input, and decode the video spatiotemporal features back into video frame data;

[0015] Step 3: Train the surgical video compression model using the video data set to obtain a trained surgical video compression model.

[0016] Optionally, step 3: training the surgical video compression model using the video data set to obtain the trained surgical video compression model includes:

[0017] Step 31: Initialize surgical video compression model parameters and training parameters;

[0018] Step 32: inputting the video data set into the spatiotemporal feature extractor to obtain video spatial features and video spatiotemporal features;

[0019] Step 33: inputting the video spatial features and the video spatiotemporal features into a stage classifier to obtain a predicted video frame category result;

[0020] Step 34: obtaining a first classification loss function according to the predicted video frame category result;

[0021] Step 35: Input the video spatiotemporal features into the decoder to obtain a reconstructed video frame;

[0022] Step 36: Obtain a second classification loss function according to the reconstructed video frame;

[0023] Step 37: Obtain a multi-task joint loss function according to the first classification loss function and the second classification loss function;

[0024] Step 38: Repeat the above steps until the iteration is completed or the loss value tends to be stable, then the training is completed, and the model parameters when the training is completed are obtained.

[0025] Optionally, the step 31: initializing the surgical video compression model parameters and training parameters includes:

[0026] The spatiotemporal feature extractor includes a ConvNext network and an LSTM module. The ConvNext network is initialized using pre-trained parameters trained on the ImageNet dataset, and the LSTM module parameters are initialized to a standard Gaussian distribution.

[0027] The stage classifier includes a Trans layer, which is initialized with Transformer parameters pre-trained on a public video dataset to ensure that the model has good generalization ability for the fusion of spatiotemporal features;

[0028] The decoder includes a Ground Decoding module, and the parameters of the Ground Decoding module are initialized to pre-trained parameters under the generative adversarial network framework. Optionally, the step 32: inputting the video data set into the spatiotemporal feature extractor to obtain video spatial features and video spatiotemporal features includes:

[0029] The ConvNext network extracts spatial features from each frame of the image and generates video spatial features represented by high-dimensional features;

[0030] The video spatial features are input into the LSTM module, and the temporal features of the video are learned using the temporal modeling capability of the long short-term memory network to obtain the video spatiotemporal features.

[0031] Optionally, the step 34: obtaining a first classification loss function according to the predicted video frame category result comprises:

[0032] The video spatial features and video spatiotemporal features are input into the stage classifier. The Trans layer in the stage classifier uses a multi-head self-attention mechanism to further fuse the spatial and temporal features, and classifies the fused features through a multi-layer perceptron MLP to obtain the predicted video frame category results, and calculate the loss with the real labels of the video frames in the surgical video dataset to minimize the difference between the classification results of the video frames and the real label values. The obtained classification loss function is as follows:

[0033]

[0034] Among them, L class represents the classification loss value, y is the true label value of the video frame, is the predicted video frame classification result, n is the number of categories, the loss value is calculated by the loss function, and the network parameters are updated in the direction of reducing the loss value. The network parameters that need to be updated in this part are the stage classifier and the spatiotemporal feature extractor.

[0035] Optionally, the step 35: inputting the video spatiotemporal features into a decoder to obtain a reconstructed video frame comprises:

[0036] The video spatiotemporal features are input into the decoder, and the Ground Decoding module decodes the spatiotemporal features and restores them to surgical video frame data. The decoding process combines a multi-layer self-attention mechanism with a generative adversarial network to obtain reconstructed video frames.

[0037] Optionally, the step 36: obtaining a second classification loss function according to the reconstructed video frame includes:

[0038] The loss of the reconstructed video frame and the real surgical video frame in the surgical video dataset is calculated to minimize the difference between the reconstructed video frame and the real surgical video frame. The classification loss function is as follows:

[0039]

[0040] Among them, L recon represents the reconstruction loss value, y is the real surgical video frame in the surgical video dataset, is the reconstructed surgical video frame generated by the model, C is the number of surgical video frame categories, the loss value is calculated by the loss function, and the network parameters are updated in the direction of reducing the loss value. The network parameters that need to be updated in this part include the spatiotemporal feature extractor and decoder.

[0041] Optionally, the step 37: obtaining a multi-task joint loss function according to the first classification loss function and the second classification loss function includes:

[0042] The classification loss and reconstruction loss of the integrated stage are combined to form a multi-task joint loss function:

[0043] L total =αL class +βL recon

[0044] Among them, α and β are weight parameters used to balance the two losses, and the entire model parameters, including the spatiotemporal feature extractor, stage classifier, and decoder, are optimized through back propagation.

[0045] Optionally, the deep generative model training method for surgical video data compression further comprises:

[0046] The trained surgical video compression model is verified, and the verification process includes:

[0047] Generate validation set video data;

[0048] The verification set video data is input into the spatiotemporal feature extractor in the model frame by frame. The ConvNext network extracts the spatial features of each frame of the video to obtain the verification video spatial features represented by high-dimensional features. The LSTM module performs time series modeling on the verification video spatial features of the video frames to generate the verification video spatiotemporal features.

[0049] Input the verification video spatial features and verification video spatiotemporal features into the stage classifier, use the Trans layer to classify the surgical stages, and evaluate the accuracy of the classification results based on the real label data of the surgical stages;

[0050] The video spatiotemporal features are input into the decoder, and the video frames are reconstructed using the Ground Decoding module. The decoder generates video frames through a multi-layer self-attention mechanism, calculates the similarity between the reconstructed video and the original video, and evaluates it using image quality assessment indicators;

[0051] Output model validation results.

[0052] The present application also provides a deep generative model training device for surgical video data compression, the deep generative model training device for surgical video data compression comprising:

[0053] A video data set acquisition module, wherein the video data set acquisition module is used to acquire a video data set;

[0054] A surgical video compression model acquisition module, the surgical video compression model acquisition module is used to acquire a surgical video compression model, the surgical video compression model includes a spatiotemporal feature extractor, a stage classifier and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and spatiotemporal features of surgical video data frames, the stage classifier is used to receive video spatial features and video spatiotemporal features as input, and obtain a stage classification result by further fusing spatial features and temporal features, and the decoder is used to receive video spatiotemporal features as input, and decode the video spatiotemporal features back into video frame data;

[0055] A training module, wherein the training module is used to train the surgical video compression model using the video data set, thereby obtaining a trained surgical video compression model.

[0056] The deep generative model training method for surgical video data compression in this application has the following advantages:

[0057] 1. Achieve a balance between high compression rate and high quality: The present invention uses a deep generative model to efficiently compress the surgical video data with a large storage volume, compressing the original video frame into a feature vector of each frame, greatly reducing the volume of stored data and achieving a high compression rate. At the same time, by combining the joint training of the decoder module and the classification module, the model parameters are optimized so that the generated feature vector can effectively retain the key visual details of the original video frame, achieving a balance between extremely high compression rate and high restoration quality, which is better than the contradiction between compression rate and restoration quality in the prior art that is difficult to reconcile.

[0058] 2. Improve decoding efficiency: The present invention uses a decoder module based on a deep neural network to decode the feature vector of the compressed video frame. It designs a lightweight decoder structure for efficient decoding and can use the good parallel computing capability of the deep neural network to further speed up the decoding speed of the feature vector of the compressed video frame, thereby improving the decoding efficiency compared to the existing methods.

[0059] 3. Enhanced detail restoration capability: Aiming at the high-fidelity requirement for important details in surgical videos, the present invention proposes a strategy for multi-task joint training of decoder modules and classifier modules. By optimizing the structure and training method of the deep generative model, the reconstruction effect of key areas (such as micro-blood vessels, tissue texture, etc.) is significantly improved, thereby ensuring a more realistic detail restoration effect under high compression rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a flowchart of a deep generative model training method for surgical video data compression in one embodiment of the present application.

[0061] Figure 2 4 is a schematic diagram of a multimodal adnexal mass malignancy risk assessment device in one embodiment of the present application.

[0062] Figure 3 It is a detailed flowchart of a deep generative model training method for surgical video data compression in one embodiment of the present application.

[0063] Figure 4 It is a detailed flowchart of the verification process in the deep generative model training method for surgical video data compression in one embodiment of the present application.

[0064] Figure 5 It is a schematic diagram of the Trans layer module structure in one embodiment of the present application.

[0065] Figure 6 It is a schematic diagram of the structure of the Ground Decoding module in one embodiment of the present application. DETAILED DESCRIPTION

[0066] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0067] like Figure 1 as well as Figure 2 The deep generative model training method for surgical video data compression shown includes:

[0068] Step 1: Get the video dataset;

[0069] Step 2: Obtain a surgical video compression model, which includes a spatiotemporal feature extractor, a stage classifier, and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and fuse spatiotemporal features of surgical video data frames, the stage classifier is used to receive video spatial features and video spatiotemporal features as input, and obtain a stage classification result by further fusing spatial features and temporal features, and the decoder is used to receive video spatiotemporal features as input, and decode the video spatiotemporal features back into video frame data;

[0070] Step 3: Train the surgical video compression model using the video data set to obtain a trained surgical video compression model.

[0071] The deep generative model training method for surgical video data compression in this application has the following advantages:

[0072] 1. Achieve a balance between high compression rate and high quality: The present invention uses a deep generative model to efficiently compress the surgical video data with a large storage volume, compressing the original video frame into a feature vector of each frame, greatly reducing the volume of stored data and achieving a high compression rate. At the same time, by combining the joint training of the decoder module and the classification module, the model parameters are optimized so that the generated feature vector can effectively retain the key visual details of the original video frame, achieving a balance between extremely high compression rate and high restoration quality, which is better than the contradiction between compression rate and restoration quality in the prior art that is difficult to reconcile.

[0073] 2. Improve decoding efficiency: The present invention uses a decoder module based on a deep neural network to decode the feature vector of the compressed video frame. It designs a lightweight decoder structure for efficient decoding and can use the good parallel computing capability of the deep neural network to further speed up the decoding speed of the feature vector of the compressed video frame, thereby improving the decoding efficiency compared to the existing methods.

[0074] 3. Enhanced detail restoration capability: Aiming at the high-fidelity requirement for important details in surgical videos, the present invention proposes a strategy for multi-task joint training of decoder modules and classifier modules. By optimizing the structure and training method of the deep generative model, the reconstruction effect of key areas (such as micro-blood vessels, tissue texture, etc.) is significantly improved, thereby ensuring a more realistic detail restoration effect under high compression rate.

[0075] In this embodiment, the video data set can be obtained in the following manner:

[0076] To ensure effective training of the model, it is necessary to determine the consistency of the surgery type and video content and collect qualified surgical video data. In the embodiment of the present invention, endoscopic tumor resection surgery videos are used, and the key stages of the surgery are selected and intercepted to ensure that the collected videos have a consistent format and high-quality content to meet the model training requirements.

[0077] Confirm the type of surgery and stage division. When collecting surgical video data, first confirm whether the type of surgery involved meets the standards for endoscopic tumor resection surgery. At the same time, identify and select the key stages in the surgical video, including but not limited to tumor localization, resection, hemostasis and other operation steps, so as to provide accurate stage annotations for subsequent model training.

[0078] Validity of surgical video content. In order to ensure the validity of surgical video content, the embodiment intercepts the key stages of surgical video data collected from the hospital, intercepts the complete surgical video according to the surgical process, and only retains the key stages of the surgery that need to be paid attention to, reducing redundant data and concentrating resources on the part that the model needs to learn most.

[0079] The collected endoscopic tumor resection surgery video data are preprocessed and labeled to form a surgical video dataset that is convenient for model training.

[0080] In order to improve the model training effect, the surgical video data is comprehensively preprocessed, including the following steps: ① Invalid area removal: Endoscopic surgical videos may contain invalid areas due to shooting angles or obstructions, such as blurred vision, black screens, or non-surgical related scenes. Use image processing algorithms to automatically detect and remove these invalid areas to retain as much effective surgical operation content as possible. ② Video format standardization: The processed surgical videos are formatted in a standardized manner, including unifying the resolution, setting a frame rate of 25 frames per second, adjusting the video encoding format, etc., to ensure the consistency of all videos, which is convenient for subsequent feature extraction and model training.

[0081] In order to achieve high-precision surgical stage recognition tasks, the processed surgical videos are annotated in detail. According to the specific operation process of endoscopic tumor resection surgery, the surgical videos are annotated into different stage categories, such as tumor localization, resection, hemostasis, etc. In seconds, the surgical videos used for dataset production are downsampled to 1 frame per second. Each frame of video data is accurately matched with the surgical stage category to which it belongs, ensuring the synchronization and accuracy of the stage information and video content, and providing a clear and specific learning goal for the model.

[0082] The annotated video dataset was divided to ensure the generalization ability of the model. All processed surgical video data were divided into training set and test set in a ratio of 8:2. The training set was used to train the deep learning model, and the test set was used to evaluate the performance of the model. When dividing, it was ensured that each surgical stage category had a relatively balanced distribution in the training set and the test set to avoid data bias affecting the learning effect of the model.

[0083] In this embodiment, step 3: training the surgical video compression model using the video data set to obtain the trained surgical video compression model includes:

[0084] Step 31: Initialize surgical video compression model parameters and training parameters;

[0085] Step 32: inputting the video data set into the spatiotemporal feature extractor to obtain video spatial features and video spatiotemporal features;

[0086] Step 33: inputting the video spatial features and the video spatiotemporal features into a stage classifier to obtain a predicted video frame category result;

[0087] Step 34: obtaining a first classification loss function according to the predicted video frame category result;

[0088] Step 35: Input the video spatiotemporal features into the decoder to obtain a reconstructed video frame;

[0089] Step 36: Obtain a second classification loss function according to the reconstructed video frame;

[0090] Step 37: Obtain a multi-task joint loss function according to the first classification loss function and the second classification loss function;

[0091] Step 38: Repeat the above steps until the iteration is completed or the loss value tends to be stable, then the training is completed, and the model parameters when the training is completed are obtained.

[0092] In this embodiment, the deep generative model used in the embodiment of the present invention includes three components: a spatiotemporal feature extractor, a stage classifier, and a decoder. The model realizes surgical stage recognition and video data reconstruction simultaneously through multi-task learning, thereby retaining key video details while compressing and storing. The spatiotemporal feature extractor consists of a ConvNext network and an LSTM module, which respectively completes the spatial feature extraction and spatiotemporal feature fusion of surgical video data frames. The stage classifier is composed of a Trans layer, which receives video spatial features and video spatiotemporal features as input, and further fuses spatial features and temporal features through an attention mechanism similar to Transformer to obtain a stage classification result. The decoder is composed of a Ground Decoding module, which receives video spatiotemporal features as input, and decodes the video spatiotemporal features back to video frame data in a manner similar to Transformer Decoder.

[0093] In this embodiment, the step 31: initializing the surgical video compression model parameters and training parameters includes:

[0094] The spatiotemporal feature extractor includes a ConvNext network and an LSTM module. The ConvNext network is initialized using pre-trained parameters trained on the ImageNet dataset, and the LSTM module parameters are initialized to a standard Gaussian distribution.

[0095] The stage classifier includes a Trans layer, which is initialized with Transformer parameters pre-trained on a public video dataset to ensure that the model has good generalization ability for the fusion of spatiotemporal features; in this embodiment, the Trans layer module is a model structure based on multi-head self-attention and feedforward network, which is usually used as the basic component of the Transformer model. The Trans layer structure used for classification in the present invention is as follows Figure 5 As shown, x i Represents the feature vector of the video frame at the i-th moment, Linear represents the linear layer, Softmax represents the activation function, Feedback represents the feedforward network, and Cls represents the classification vector.

[0096] The decoder includes a Ground Decoding module, and the parameters of the Ground Decoding module are initialized to the pre-trained parameters under the generative adversarial network framework. In this embodiment, the Ground Decoding module is a method for high-resolution feature map decoding in deep learning, which achieves high-quality output by gradually restoring the spatial resolution of the feature map while retaining key semantic information. The Trans layer structure used for classification in the present invention is as follows Figure 6As shown, where x t Represents the video frame feature vector at the tth moment, Normalize represents the normalization layer, Fully-connected represents the fully connected layer, PixelNorm represents pixel normalization, Conv represents the convolutional layer, and Video frame represents the decoded video frame.

[0097] Specifically, in the spatiotemporal feature extractor ConNext+LSTM, the ConvNext network is initialized with pre-trained parameters trained on the ImageNet dataset to ensure strong extraction capabilities for the spatial features of video frames. The LSTM module parameters are initialized to a standard Gaussian distribution for efficient fusion of temporal features. In the stage classifier, Translayer is initialized with Transformer parameters pre-trained on a public video dataset to ensure that the model has good generalization capabilities for the fusion of spatiotemporal features. In the decoder, the Ground Decoding module parameters are initialized to pre-trained parameters under the Generative Adversarial Network (GAN) framework to enhance the video reconstruction effect. In addition, the training parameters of the model also need to be set, including the number of training rounds, the learning rate, and the number of video batch frames input to the model.

[0098] In this embodiment, video frame input data and corresponding stage label data are extracted from the video data set according to the number of batch frames, and the input frames are used as the real data of the video frames for reconstruction. The timing sequence of the data in each batch is ensured to be correct to stabilize the training process.

[0099] In this embodiment, the step 32: inputting the video data set into the spatiotemporal feature extractor to obtain video spatial features and video spatiotemporal features includes:

[0100] The ConvNext network extracts spatial features from each frame of the video dataset and generates video spatial features represented by high-dimensional features;

[0101] The video spatial features are input into the LSTM module, and the temporal features of the video are learned using the temporal modeling capability of the long short-term memory network to obtain the video spatiotemporal features.

[0102] In this embodiment, the step 34: obtaining a first classification loss function according to the predicted video frame category result includes:

[0103] The video spatial features and video spatiotemporal features are input into the stage classifier. The Trans layer in the stage classifier uses a multi-head self-attention mechanism to further fuse the spatial and temporal features, and classifies the fused features through a multi-layer perceptron MLP to obtain the predicted video frame category results, and calculate the loss with the real labels of the video frames in the surgical video dataset to minimize the difference between the classification results of the video frames and the real label values. The obtained classification loss function is as follows:

[0104]

[0105] Among them, L class represents the classification loss value, y is the true label value of the video frame, is the predicted video frame classification result, n is the number of categories, the loss value is calculated by the loss function, and the network parameters are updated in the direction of reducing the loss value. The network parameters that need to be updated in this part are the stage classifier and the spatiotemporal feature extractor.

[0106] In this embodiment, the step 35: inputting the video spatiotemporal features into the decoder to obtain the reconstructed video frame includes:

[0107] The spatiotemporal features of the video are input into the decoder. The Ground Decoding module decodes the spatiotemporal features and restores them to surgical video frame data through a structure similar to the Transformer decoder. The decoding process combines a multi-layer self-attention mechanism with a generative adversarial network to obtain reconstructed video frames.

[0108] In this embodiment, the step 36: obtaining the second classification loss function according to the reconstructed video frame includes:

[0109] The loss of the reconstructed video frame and the real surgical video frame in the surgical video dataset is calculated to minimize the difference between the reconstructed video frame and the real surgical video frame. The classification loss function is as follows:

[0110]

[0111] Among them, L recon represents the reconstruction loss value, y is the real surgical video frame in the surgical video dataset, is the reconstructed surgical video frame generated by the model, C is the number of surgical video frame categories, the loss value is calculated by the loss function, and the network parameters are updated in the direction of reducing the loss value. The network parameters that need to be updated in this part include the spatiotemporal feature extractor and decoder.

[0112] In this embodiment, the step 37: obtaining the multi-task joint loss function according to the first classification loss function and the second classification loss function includes:

[0113] The classification loss and reconstruction loss of the integrated stage are combined to form a multi-task joint loss function:

[0114] L total =αL class +βL recon

[0115] Among them, α and β are weight parameters used to balance the two losses, and the entire model parameters, including the spatiotemporal feature extractor, stage classifier, and decoder, are optimized through back propagation.

[0116] In this embodiment, the deep generative model training method for surgical video data compression further includes:

[0117] See also Figure 3 , verifying the trained surgical video compression model, the verification process includes:

[0118] Generate verification set video data (the verification set video data can directly use the above verification set);

[0119] In this embodiment, if the above verification set is used, the verification set is processed as follows:

[0120] First, invalid areas are removed. Edge detection, chromaticity analysis and other algorithms are used to identify invalid areas (such as black edges or blurred parts) in the video frame caused by endoscopic shooting, and these areas are removed. Then, the video frame is normalized. Each frame of the image is normalized, and the resolution, frame rate and other parameters of the video are adjusted to keep it consistent with the original video in the training data. The video is not downsampled here.

[0121] Use the trained surgical video compression model to load the trained model parameters, ensure that the loaded parameters are compatible with the structure of the surgical video compression model, and perform model initialization and parameter loading operations.

[0122] The verification set video data is input into the spatiotemporal feature extractor in the model frame by frame. The ConvNext network extracts the spatial features of each frame of the video to obtain the verification video spatial features represented by high-dimensional features. The LSTM module performs time series modeling on the verification video spatial features of the video frames to generate the verification video spatiotemporal features.

[0123] The verification video spatial features and verification video spatiotemporal features are input into the stage classifier, and the Trans layer is used to classify the surgical stages. The accuracy of the classification results is evaluated based on the real label data of the surgical stages. The evaluation indicators of the stage classification include accuracy, recall rate and F1 score.

[0124] The spatiotemporal features of the video are input into the decoder, and the video frames are reconstructed using the Ground Decoding module. The decoder generates video frames through a multi-layer self-attention mechanism, calculates the similarity between the reconstructed video and the original video, and uses image quality assessment indicators (such as structural similarity SSIM) to evaluate and quantify the reconstruction effect of the compression model;

[0125] The model verification results are output, including stage classification accuracy, video frame reconstruction quality indicators, etc. According to the compression and reconstruction effects of the surgical video, the fidelity of the key operation steps in the video is further analyzed to determine whether the model is suitable for use in actual surgical video compression applications.

[0126] The surgical video compression model after the output model verification result is fine-tuned according to the test effect on the actual surgical video compression task to complete the construction of the method of the present invention.

[0127] The present application also provides a deep generative model training device for surgical video data compression, the deep generative model training device for surgical video data compression comprises a video data set acquisition module, a surgical video compression model acquisition module and a training module, wherein:

[0128] The video data set acquisition module is used to acquire the video data set;

[0129] The surgical video compression model acquisition module is used to acquire the surgical video compression model, and the surgical video compression model includes a spatiotemporal feature extractor, a stage classifier and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and spatiotemporal features of surgical video data frames, the stage classifier is used to receive video spatial features and video spatiotemporal features as input, and obtain the stage classification result by further fusing the spatial features and the time features, and the decoder is used to receive video spatiotemporal features as input, and decode the video spatiotemporal features back to video frame data;

[0130] The training module is used to train the surgical video compression model using the video data set, thereby obtaining a trained surgical video compression model.

[0131] The above explanation of the method also applies to the explanation of the device.

[0132] Figure 2 It is a structural block diagram of an electronic device provided by one or more embodiments of the present invention.

[0133] like Figure 2As shown, the present application also discloses an electronic device (i.e., the master controller in the present application), comprising: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the deep generative model training method for surgical video data compression.

[0134] The present application also provides a computer-readable storage medium, which stores a computer program that can be executed by an electronic device. When the computer program runs on the electronic device, it can implement the steps of the deep generative model training method for surgical video data compression.

[0135] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0136] The electronic device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and a memory. The operating system can be any one or more computer operating systems that implement electronic device control through a process, such as a Linux operating system, a Unix operating system, an Android operating system, an iOS operating system, or a Windows operating system. In addition, in an embodiment of the present invention, the electronic device can be a handheld device such as a smart phone or a tablet computer, or can be an electronic device such as a desktop computer or a portable computer, which is not particularly limited in the embodiment of the present invention.

[0137] The execution subject of the electronic device control in the embodiment of the present invention may be an electronic device, or a functional module in the electronic device that can call and execute a program. The electronic device may obtain the firmware corresponding to the storage medium. The firmware corresponding to the storage medium is provided by the supplier. The firmware corresponding to different storage media may be the same or different, which is not limited here. After the electronic device obtains the firmware corresponding to the storage medium, the firmware corresponding to the storage medium may be written into the storage medium, specifically, the firmware corresponding to the storage medium may be burned into the storage medium. The process of burning the firmware into the storage medium may be implemented using existing technology, which will not be described in detail in the embodiment of the present invention.

[0138] The electronic device may also obtain a reset command corresponding to the storage medium. The reset command corresponding to the storage medium is provided by the supplier. The reset commands corresponding to different storage media may be the same or different, and are not limited here.

[0139] At this time, the storage medium of the electronic device is a storage medium in which the corresponding firmware is written, and the electronic device can respond to the reset command corresponding to the storage medium in the storage medium in which the corresponding firmware is written, so that the electronic device resets the storage medium in which the corresponding firmware is written according to the reset command corresponding to the storage medium. The process of resetting the storage medium according to the reset command can be implemented by the existing technology and will not be described in detail in the embodiments of the present invention.

[0140] For the convenience of description, the above devices are described in terms of functions and are divided into various units and modules. Of course, when implementing the present application, the functions of each unit and module can be implemented in the same or multiple software and / or hardware.

[0141] Those skilled in the art will appreciate that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined.

[0142] For the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0143] It can be known from the description of the above implementation modes that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute the methods described in the various implementation modes of the present application or certain parts of the implementation modes.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A deep generative model training method for surgical video data compression, characterized in that: The deep generative model training method for surgical video data compression includes: Step 1: Get the video dataset; Step 2: Obtain a surgical video compression model, which includes a spatiotemporal feature extractor, a stage classifier, and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and fuse spatiotemporal features of surgical video data frames, the stage classifier is used to receive video spatial features and video spatiotemporal features as input, and obtain a stage classification result by further fusing spatial features and temporal features, and the decoder is used to receive video spatiotemporal features as input, and decode the video spatiotemporal features back into video frame data; Step 3: Train the surgical video compression model using the video data set to obtain a trained surgical video compression model.

2. The deep generative model training method for surgical video data compression according to claim 1, characterized in that: Step 3: Training the surgical video compression model using the video data set to obtain the trained surgical video compression model includes: Step 31: Initialize surgical video compression model parameters and training parameters; Step 32: inputting the video data set into the spatiotemporal feature extractor to obtain video spatial features and video spatiotemporal features; Step 33: inputting the video spatial features and the video spatiotemporal features into a stage classifier to obtain a predicted video frame category result; Step 34: obtaining a first classification loss function according to the predicted video frame category result; Step 35: Input the video spatiotemporal features into the decoder to obtain a reconstructed video frame; Step 36: Obtain a second classification loss function according to the reconstructed video frame; Step 37: Obtain a multi-task joint loss function according to the first classification loss function and the second classification loss function; Step 38: Repeat the above steps until the iteration is completed or the loss value tends to be stable, then the training is completed, and the model parameters when the training is completed are obtained.

3. The deep generative model training method for surgical video data compression according to claim 2, characterized in that: The step 31: initializing the surgical video compression model parameters and training parameters includes: The spatiotemporal feature extractor includes a ConvNext network and an LSTM module. The ConvNext network is initialized using pre-trained parameters trained on the ImageNet dataset, and the LSTM module parameters are initialized to a standard Gaussian distribution. The stage classifier includes a Trans layer, which is initialized with Transformer parameters pre-trained on a public video dataset to ensure that the model has good generalization ability for the fusion of spatiotemporal features; The decoder includes a Ground Decoding module, and the parameters of the Ground Decoding module are initialized to the pre-trained parameters under the generative adversarial network framework.

4. The deep generative model training method for surgical video data compression according to claim 3, characterized in that: The step 32: inputting the video data set into the spatiotemporal feature extractor to obtain video spatial features and video spatiotemporal features includes: The ConvNext network extracts spatial features from each frame of the image and generates video spatial features represented by high-dimensional features; The video spatial features are input into the LSTM module, and the temporal features of the video are learned using the temporal modeling capability of the long short-term memory network to obtain the video spatiotemporal features.

5. The deep generative model training method for surgical video data compression according to claim 4, characterized in that: The step 34: obtaining a first classification loss function according to the predicted video frame category result comprises: The video spatial features and video spatiotemporal features are input into the stage classifier. The Trans layer in the stage classifier uses a multi-head self-attention mechanism to further fuse the spatial and temporal features, and classifies the fused features through a multi-layer perceptron MLP to obtain the predicted video frame category results, and calculate the loss with the real labels of the video frames in the surgical video dataset to minimize the difference between the classification results of the video frames and the real label values. The obtained classification loss function is as follows: Among them, L class represents the classification loss value, y is the true label value of the video frame, is the predicted video frame classification result, n is the number of categories, the loss value is calculated by the loss function, and the network parameters are updated in the direction of reducing the loss value. The network parameters that need to be updated in this part are the stage classifier and the spatiotemporal feature extractor.

6. The deep generative model training method for surgical video data compression according to claim 5, characterized in that: The step 35: inputting the video spatiotemporal features into the decoder to obtain the reconstructed video frame comprises: The video spatiotemporal features are input into the decoder, and the Ground Decoding module decodes the spatiotemporal features and restores them to surgical video frame data. The decoding process combines a multi-layer self-attention mechanism with a generative adversarial network to obtain reconstructed video frames.

7. The deep generative model training method for surgical video data compression according to claim 6, characterized in that: The step 36: obtaining a second classification loss function according to the reconstructed video frame comprises: The loss of the reconstructed video frame and the real surgical video frame in the surgical video dataset is calculated to minimize the difference between the reconstructed video frame and the real surgical video frame. The classification loss function is as follows: Among them, L recon represents the reconstruction loss value, y is the real surgical video frame in the surgical video dataset, is the reconstructed surgical video frame generated by the model, C is the number of surgical video frame categories, the loss value is calculated by the loss function, and the network parameters are updated in the direction of reducing the loss value. The network parameters that need to be updated in this part include the spatiotemporal feature extractor and decoder.

8. The deep generative model training method for surgical video data compression according to claim 7, characterized in that: The step 37: obtaining a multi-task joint loss function according to the first classification loss function and the second classification loss function comprises: The classification loss and reconstruction loss of the integrated stage are combined to form a multi-task joint loss function: L total =αL class +βL recon Among them, α and β are weight parameters used to balance the two losses, and the entire model parameters, including the spatiotemporal feature extractor, stage classifier, and decoder, are optimized through back propagation.

9. The deep generative model training method for surgical video data compression according to claim 8, characterized in that: The deep generative model training method for surgical video data compression further comprises: The trained surgical video compression model is verified, and the verification process includes: Generate validation set video data; The verification set video data is input into the spatiotemporal feature extractor in the model frame by frame. The ConvNext network extracts the spatial features of each frame of the video to obtain the verification video spatial features represented by high-dimensional features. The LSTM module performs time series modeling on the verification video spatial features of the video frames to generate the verification video spatiotemporal features. Input the verification video spatial features and verification video spatiotemporal features into the stage classifier, use the Trans layer to classify the surgical stages, and evaluate the accuracy of the classification results based on the real label data of the surgical stages; The video spatiotemporal features are input into the decoder, and the video frames are reconstructed using the Ground Decoding module. The decoder generates video frames through a multi-layer self-attention mechanism, calculates the similarity between the reconstructed video and the original video, and evaluates it using image quality assessment indicators; Output model validation results.

10. A deep generative model training device for surgical video data compression, characterized in that: The deep generative model training device for surgical video data compression comprises: A video data set acquisition module, wherein the video data set acquisition module is used to acquire a video data set; A surgical video compression model acquisition module, the surgical video compression model acquisition module is used to acquire a surgical video compression model, the surgical video compression model includes a spatiotemporal feature extractor, a stage classifier and a decoder, wherein the spatiotemporal feature extractor is used to extract spatial features and spatiotemporal features of surgical video data frames, the stage classifier is used to receive video spatial features and video spatiotemporal features as input, and obtain a stage classification result by further fusing spatial features and temporal features, and the decoder is used to receive video spatiotemporal features as input, and decode the video spatiotemporal features back into video frame data; A training module, wherein the training module is used to train the surgical video compression model using the video data set, thereby obtaining a trained surgical video compression model.