Operation stage identification method based on dynamic data balance strategy
By constructing the Xception-dual-stream LSTM model, combining dynamic sampling rate and dual-stream LSTM module, the category imbalance and timing information loss in the surgical stage recognition are solved, the recognition accuracy and robustness of the surgical stage are improved, and the development of intelligent surgical assistance systems is supported.
Patent Information
- Application Number
- CN202510379876.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing surgical stage identification model faces the problems of category imbalance and local timing information loss, and it is difficult to effectively identify complex surgical timing data.
Using a surgical stage recognition method based on dynamic data balance strategy, the Xception-dual stream LSTM recognition model is constructed, combining dynamic sampling rate and dual stream LSTM modules, the spatiotemporal information and dynamic relationships in surgical videos are captured, the data imbalance problem is alleviated and timing feature extraction is enhanced.
It improves the recognition accuracy and robustness of the surgical stage, can capture timing feature information more comprehensively, significantly improves the classification performance of the model in the surgical stage identification task, and supports surgical training and clinical decision-making.
Smart Images

Figure CN120236232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visual recognition, and in particular to a surgical stage recognition method based on a dynamic data balancing strategy. Background Art
[0002] Analysis of surgical workflows is one of the core research topics in the field of Computer Assisted Surgery (CAS). Its goal is to improve patient safety, reduce surgical errors, and optimize communication efficiency in the operating room. This task mainly focuses on identifying and accurately dividing different stages of surgery from recorded laparoscopic surgery videos through automation techniques, and then supporting applications such as surgical process monitoring, process optimization, safety assessment, surgical plan extraction, and decision support. Automated surgical stage recognition is not only crucial for surgical training and intraoperative assistance, but also promotes the optimization of surgical processes, thereby improving overall efficiency and safety.
[0003] Laparoscopic cholecystectomy (LC) is a common minimally invasive surgical procedure. An endoscope camera and special surgical tools are inserted through small incisions, and the whole process highly depends on the perspective provided by the endoscope for precise operation. In recent years, due to the high dependence of this surgery on endoscope visualization technology, surgical data has been widely publicized and accumulated. The growth of this data has laid a foundation for the automated analysis and evaluation of surgical processes based on machine vision and deep learning, and provided the potential for further improving surgical safety and efficiency.
[0004] In the research of applying deep learning (DL) to surgical workflow analysis, the surgical phase recognition (SPR) task often uses the following standard datasets, such as: M2CAI16 (cholecystectomy videos), Cholec80 and CholecT45 (cholecystectomy videos), HeiChole (cholecystectomy videos), Cataract-101 (cataract surgery videos), CATARACTS (cataract surgery videos), and LapSig300 (colorectal surgery videos), etc. However, the class imbalance of surgical phases is a common challenge in these datasets. Taking the Cholec80 dataset as an example, this dataset contains a total of 7 surgical phase categories, among which the 2nd and 4th phases account for more than 70% of the training set frames, while the less frequent 5th and 7th phases only account for about 4% and 3% respectively. This indicates significant heterogeneity and uneven class frequencies in the datasets for the surgical phase recognition task. Such performance will have an adverse impact on the prediction performance of deep learning models, easily causing the model to be biased towards predicting the phases with larger sample sizes, thereby reducing the recognition accuracy of few-sample classes and potentially leading to overfitting problems. Although relevant research in the field of cholecystectomy phase recognition has shown this problem, in-depth analysis of the class imbalance problem is still lacking, resulting in most models still being biased towards predicting more frequent surgical phases.
[0005] Existing surgical phase recognition models usually adopt architectures based on the combination of convolutional neural networks (CNNs) and recurrent neural networks (RNNs), or architectures based on attention mechanisms. In the CNN-RNN hybrid architecture, the RNN module usually utilizes its recursive characteristics to capture the full-time scale dependencies of the input sequence. However, this structure may cause local information loss during the training process, resulting in the model's weakened response to long-term dependencies (manifested as lower weights), while being overly sensitive to short-term dependencies (manifested as higher weights), thus affecting the overall prediction performance of the model. This bias towards short-term dependencies may cause the model to be insensitive to changes on longer time scales, and thus perform poorly when dealing with complex time series data.
[0006] To address the above problems of uneven class frequencies and local time series information loss, a surgical phase recognition method based on a dynamic data balancing strategy is proposed. Summary of the Invention
[0007] To solve the technical problems of uneven class frequencies and local time series information loss, the present invention provides a surgical phase recognition method based on a dynamic data balancing strategy.
[0008] The present invention is implemented by the following technical solutions: A surgical stage recognition method based on a dynamic data balancing strategy, comprising the following steps:
[0009] Step S1: Collect the original surgical data and annotate the collected original data;
[0010] Step S2: Preprocess the annotated original data by using a dynamic data balancing strategy to form a data set;
[0011] Step S3: Construct an Xception-two-stream LSTM recognition model;
[0012] Step S4: Use the data set to train and optimize the Xception-two-stream LSTM recognition model;
[0013] Step S5: Select the Xception-two-stream LSTM recognition model with the highest training accuracy to perform recognition scoring on the surgical video.
[0014] As a further improvement of the above solution, the original surgical data in step S1 is the original data in the Cholec80 data set, including video segments of different cholecystectomy surgeries. The video segments are all labeled with different surgical operation processes and detailed time series annotations are made for the video segments; the video segments record the complete process from the start to the end of the surgery, and the video segments include a preparation stage, a triangle dissection stage, a blood vessel shearing stage, a gallbladder dissection stage, a gallbladder bagging stage, a blood clot or coagulated tissue cleaning stage, and a gallbladder recovery stage.
[0015] As a further improvement of the above solution, in step S2, the video segments are undersampled into short segments composed of a set number of frames, that is, the sampling rate of the surgical stages of the video segments is calculated according to the stride of the surgical stages, and the image frames of the surgical stages are undersampled according to the stride. The stride calculation uses the following formula:
[0016]
[0017] where S is the stride, representing the sampling interval, rounded to the nearest integer; R f is the frame rate of the original video, F represents the number of image frames of the current surgical stage, and N represents the maximum number of image frames in all surgical stages.
[0018] As a further improvement of the above solution, in step S3, the Xception-two-stream LSTM recognition model includes an Xception module and a two-stream LSTM module. The Xception module is used to extract the visual features of the dataset and form feature vectors. The two-stream LSTM module extracts the spatio-temporal information in the surgical actions from the feature vectors, captures the overall motion features within the video clip, calculates the dynamic relationships between frames, and fuses the predicted scores output by matrix addition to finally determine the predicted score of the surgical stage.
[0019] As a further improvement of the above solution, in step S3, the Xception module includes multiple stacked depthwise separable convolutional layers, a TimeDistributed layer, and a max pooling layer. The depthwise separable convolutional layer performs depthwise and pointwise convolution operations on the dataset data to extract the visual feature data in the dataset. The TimeDistributed layer performs normalization processing on the visual feature data extracted by the depthwise separable convolutional layer. The max pooling layer implements pooling processing on the visual feature data.
[0020] As a further improvement of the above solution, in step S3, the two-stream LSTM module includes a bidirectional LSTM structure with temporal mapping and an LSTM structure with sequence embedding;
[0021] Among them, the bidirectional LSTM structure with temporal mapping is used to extract the temporal dependencies of the context from short segments, and the LSTM structure with sequence embedding is used to extract the semantic information of each frame and capture the short-term temporal relationships between frames.
[0022] As a further improvement of the above solution, in step S4, the dataset is divided into a training set, a validation set, and a test set. The training set is used for training and tuning the parameters of the Xception-two-stream LSTM recognition model, and the test set is used for accuracy testing of the Xception-two-stream LSTM recognition model. During this period, the parameters of the Xception-two-stream LSTM recognition model are continuously adjusted, and finally an Xception-two-stream LSTM recognition model with the best accuracy and its parameters are selected as the initialized Xception-two-stream LSTM recognition model.
[0023] As a further improvement of the above solution, in step S4, when training the Xception-two-stream LSTM recognition model, a data loading repetition mechanism is adopted, and the data is loaded in batches within each training cycle.
[0024] As a further improvement of the above solution, during the training of the Xception-two-stream LSTM recognition model in step S4, the Adam optimizer is used to initialize the recognition model, and the initial learning rate, training dropout rate, training epochs, and training batch size for the training of the Xception-two-stream LSTM recognition model are set. The Xception-two-stream LSTM recognition model uses the dropout rate method for random dropout training, and the sigmoid and tanh activation functions are used. The categorical cross-entropy loss function is used to evaluate the progress of the Xception-two-stream LSTM recognition model.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0026] 1. By means of short-time segment division and dynamic sampling rate allocation, the present invention effectively alleviates the problem of biased learning, effectively reduces the influence of data volume differences on the model learning process, reduces bias, and improves the overall recognition accuracy of the model for surgical stages.
[0027] 2. The present invention can simultaneously fuse temporal dynamic information and static features, enhancing the accuracy and robustness of surgical stage recognition; by parallel processing of short-term and long-term temporal dependence information, it significantly enhances the understanding of the surgical process; it can not only effectively solve the problem of data imbalance, capture temporal feature information more comprehensively, and thus exhibits higher classification performance in the surgical stage recognition task.
[0028] 3. The present invention effectively improves the overall recognition accuracy of surgical stages, and can also provide more structured support for surgical training, quality control, and clinical decision-making, further laying a foundation for the research and development of intelligent surgical assistance systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 It is a flowchart of a surgical stage recognition method based on a dynamic data balance strategy provided by the present invention;
[0030] Figure 2 It is a structural schematic diagram of a two-stream LSTM module provided by the present invention;
[0031] Figure 3 It is a graph showing the F1-score trend of the fixed undersampling rate strategy and the dynamic data balance strategy in Embodiment 3 provided by the present invention;
[0032] Figure 4 It is a graph showing the AUC trend of the fixed undersampling rate strategy and the dynamic data balance strategy in Embodiment 3 provided by the present invention;
[0033] Figure 5 It is a comparison graph of confusion matrices when applying the fixed undersampling rate strategy and the dynamic data balance strategy provided by the present invention;
[0034] Figure 6 Flow chart of preprocessing the original data provided by the present invention using a dynamic data balance strategy. Specific implementation manners
[0035] Next, in combination with the accompanying drawings and specific implementation manners, the present invention will be further described. It should be noted that, on the premise of no conflict, the following described embodiments or technical features can be arbitrarily combined to form new embodiments.
[0036] Embodiment 1:
[0037] Please combine Figure 1 , a surgical stage recognition method based on a dynamic data balance strategy in this embodiment includes the following steps:
[0038] Step S1: Collect the original surgical data and annotate the collected original data;
[0039] The original surgical data is the original data in the Cholec80 dataset, including video segments of different cholecystectomy surgeries. The video segments are all labeled with different surgical operation processes and detailed time series annotations are made for the video segments; the video segments record the complete process from the start to the end of the surgery, and the video segments include a preparation stage, a triangle dissection stage, a blood vessel shearing stage, a gallbladder dissection stage, a gallbladder bagging stage, a blood clot or coagulated tissue cleaning stage, and a gallbladder recovery stage;
[0040] Step S2: Preprocess the annotated original data using a dynamic data balance strategy to form a dataset;
[0041] The video segments are undersampled into short segments composed of a set number of frames, that is, the sampling rate of the surgical stages of the video segments is calculated according to the stride of the surgical stages, and the image frames of the surgical stages are undersampled according to the stride. The stride calculation uses the following formula:
[0042]
[0043] where S is the stride, representing the sampling interval, rounded to the nearest integer; R f is the frame rate of the original video, F represents the number of image frames of the current surgical stage, and N represents the maximum number of image frames in all surgical stages;
[0044] Step S3: Construct an Xception - two - stream LSTM recognition model;
[0045] The Xception-two-stream LSTM recognition model includes an Xception module and a two-stream LSTM module. The Xception module is used to extract the visual features of the dataset and form feature vectors. The two-stream LSTM module extracts the spatio-temporal information in the surgical actions from the feature vectors, captures the overall motion features within the video clips, calculates the dynamic relationships between frames, and fuses them through the predicted scores output by matrix addition to finally determine the predicted scores of the surgical stages.
[0046] The Xception module includes multiple stacked depthwise separable convolutional layers, a TimeDistributed layer, and a max pooling layer. The depthwise separable convolutional layer performs depthwise and pointwise convolution operations on the dataset data to extract the visual feature data in the dataset. The TimeDistributed layer performs uniform processing on the visual feature data extracted by the depthwise separable convolutional layer. The max pooling layer implements the pooling processing of the visual feature data.
[0047] The two-stream LSTM module includes a bidirectional LSTM structure with temporal mapping and an LSTM structure with sequence embedding.
[0048] Among them, the bidirectional LSTM structure with temporal mapping is used to extract the temporal dependencies of the context from short segments, and the LSTM structure with sequence embedding is used to extract the semantic information of each frame and capture the short-term temporal relationships between frames.
[0049] Step S4: Use the dataset to train and optimize the Xception-two-stream LSTM recognition model.
[0050] The dataset is divided into a training set, a validation set, and a test set. The training set is used for training and tuning the parameters of the Xception-two-stream LSTM recognition model, and the test set is used for accuracy testing of the Xception-two-stream LSTM recognition model. During this period, the parameters of the Xception-two-stream LSTM recognition model are continuously adjusted, and finally an Xception-two-stream LSTM recognition model with the best accuracy and its parameters are selected as the initialized Xception-two-stream LSTM recognition model.
[0051] When training the Xception-two-stream LSTM recognition model, a data loading repetition mechanism is adopted, and the data is loaded in batches within each training cycle.
[0052] When training the Xception-two-stream LSTM recognition model, the Adam optimizer is used to initialize the recognition model. Set the initial learning rate, training dropout rate, number of training epochs, and training batch size for the training of the Xception-two-stream LSTM recognition model. The Xception-two-stream LSTM recognition model uses dropout for random dropout training, and uses sigmoid and tanh activation functions. The categorical cross-entropy loss function is used to evaluate the progress of the Xception-two-stream LSTM recognition model.
[0053] Step S5: Select the Xception-two-stream LSTM recognition model with the highest training accuracy to perform recognition scoring on the surgical video.
[0054] Example 2:
[0055] A surgical stage recognition method based on a dynamic data balancing strategy, comprising the following steps:
[0056] Step S1: Collect the original surgical data and annotate the collected original data;
[0057] The original surgical data is the original data in the Cholec80 dataset, which contains video segments of 80 different cholecystectomy surgeries, as shown in Table 1;
[0058] Table 1 is the category statistics table of the Cholec80 dataset:
[0059] Phase Behavior Total number of frames Total duration (s) phase1 Preparation - Preparation 214301 8572 phase2 Calot’s Triangle Dissection - Calot Triangle Dissection 1870651 74826 phase3 Clipping Cutting - Clipping Blood Vessels 352000 14080 phase4 Gallbladder Dissection - Gallbladder Dissection 1460825 58433 phase5 Gallbladder Packaging - Gallbladder Bagging 190450 7618 phase6 Cleaning Coagulation - Cleaning Blood Clots or Coagulating Tissues 358303 14332 phase7 Gallbladder Retraction - Gallbladder Retrieval 166002 6640 Total 4612532 184501
[0060] The video segments are all annotated with different surgical procedures and detailed time series annotations are made for the video segments. All videos come from real medical surgical procedures, so the video quality is relatively high in medical applications; most videos are from a standard laparoscopic perspective, with high clarity and a relatively fixed perspective; the total length of each video is approximately 20 - 30 minutes; the video segments record the complete process from the start to the end of the surgery, and the video segments include the preparation stage, triangulation stage, blood vessel cutting stage, gallbladder dissection stage, gallbladder bagging stage, blood clot or coagulated tissue cleaning stage, and gallbladder retrieval stage;
[0061] The video frame rate is 25 frames per second (fps), which is sufficient to ensure that fast movements during the surgery can be accurately captured and provide high-time-resolution data. The video resolution is usually 480p (640x480) or 720p (1280x720). This dataset can provide key time series data for in-depth analysis of the surgical process, which helps to further study and optimize the automatic recognition and classification of surgical stages;
[0062] Step S2: Preprocess the original data with annotations using a dynamic data balancing strategy to form a dataset;
[0063] As shown in Table 2, for the sampling results of the Cholec80 dataset at the original sampling rate, when a unified 1 frame per second (fps) undersampling strategy was adopted for the Cholec80 dataset, the data distribution of the undersampled videos in all stages is presented. This undersampling strategy, due to its low computational cost, is a commonly used sampling method in the current surgical stage recognition task. However, due to the imbalance in the data distribution of each stage, it may lead to biased learning during the training process of the model, thus having an adverse impact on the classification performance.
[0064] Table 2 is the sampling table of the Cholec80 dataset at the original sampling rate:
[0065]
[0066] The dynamic data balancing strategy automatically adjusts the sampling rate to undersample video segments into short segments composed of a set number of frames, which are used as inputs for model training. That is, the sampling rate of the surgical stage of the video segment is calculated according to the stride of the surgical stage, and the image frames of the surgical stage are undersampled according to the stride. The stride is calculated using the following formula:
[0067]
[0068] where S is the stride, representing the sampling interval, rounded to the nearest integer; R f is the frame rate of the original video, F represents the number of image frames in the current surgical stage, and N represents the maximum number of image frames in all surgical stages; therefore, the stride of each stage is determined based on the stage with the largest number of image frames;
[0069] As Figure 6 shown, taking video01 in the Cholec80 dataset as an example, the number of frames in the second stage is the largest. Therefore, the stride of this stage is set to 25 (equivalent to 1 frame per second). At the same time, the strides of other stages will be calculated and adjusted according to their frame numbers to ensure that the number of short segments in each stage of each video is approximately equal; specifically, in the first stage, 15 consecutive frames are extracted from 525 frames, with a stride of 1, meaning that data is extracted from each frame, generating 52 short segments; for other stages, the stride calculation method is the same; adopting the dynamic sampling strategy helps to perform more detailed feature extraction for the time-dependent relationships of actions in each stage;
[0070] Table 3 details the sampling results of each video in the Cholec80 dataset under the dynamic data balancing strategy; through this strategy, a total of 38,344 short segments were obtained from 80 videos, and the total number of segments corresponding to the 7 stages were: 4,854, 5,828, 5,799, 5,859, 5,440, 5,351, and 5,213 segments respectively; this dynamic data balancing strategy can ensure a relatively balanced number of short segments in each stage while maximizing the retention of the temporal information of the video data, effectively addressing the data imbalance problem, thereby enhancing the model's recognition ability and classification accuracy for each surgical stage;
[0071] Table 3 shows the sampling results under the dynamic data balancing strategy:
[0072]
[0073] Step S3: Construct an Xception-two-stream LSTM recognition model;
[0074] The Xception-two-stream LSTM recognition model includes an Xception module and a two-stream LSTM module. The Xception module is used to extract the visual features of the dataset and form feature vectors. The two-stream LSTM module extracts the spatio-temporal information in the surgical actions from the feature vectors, captures the overall motion features within the video segments, calculates the dynamic relationships between frames, and fuses them through the predicted scores output by matrix addition to finally determine the predicted scores of the surgical stages;
[0075] The Xception module includes multiple stacked depthwise separable convolutional layers, a TimeDistributed layer, and a max pooling layer; the depthwise separable convolutional layers perform depth and pointwise convolution operations on the dataset data to extract the visual feature data in the dataset, and the TimeDistributed layer performs uniform processing on the visual feature data extracted by the depthwise separable convolutional layers; the max pooling layer implements pooling processing on the visual feature data;
[0076] The depthwise separable convolutional layer includes depthwise convolution and pointwise convolution;
[0077] Depthwise convolution applies a separate convolutional kernel to each input feature dimension for convolution operation; if the input feature map size is H×W×C in , and the convolutional kernel size is K d ×K d ×1, then the convolution operation for each feature dimension can be expressed as:
[0078]
[0079] Among them, x(i, j, c) is the pixel value at position (i, j) and dimension c in the input feature, and w(m, n, c) is the parameter of the convolutional kernel in this feature dimension; the output feature map size of depthwise convolution is H×W×C in Pointwise convolution is to perform convolution on each feature dimension of the feature map through a convolutional kernel with a size of 1×1 to increase the feature dimension or perform feature combination;
[0080] The calculation formula of pointwise convolution is:
[0081]
[0082] Among them, w(1, 1, c, k) is the parameter of the 1×1 convolutional kernel, and x(i, j, c) is the feature map output by depthwise convolution;
[0083] The depthwise separable convolutional layer combines depthwise convolution and pointwise convolution to effectively reduce the computational amount; the overall formula is expressed as:
[0084] y = DepthwiseConv(x) then y′ = PointwiseConv(y)
[0085] Among them, x represents the input feature map with a dimension of H×W×C in ; y is the output of depthwise convolution with a dimension of H×W×C in ; y' is the output after pointwise convolution with a dimension of H×W×C out ;
[0086] The feature dimension output by the depthwise separable convolutional layer of the last layer is (None, 15, 7, 7, 512), where None represents the batch size, 15 is the number of video frames, 7×7 is the spatial feature map size of each frame, and 512 is the number of channels at each position; in order to perform consistent processing on the feature maps of each frame, the TimeDistributed layer is used so that the feature maps of each frame undergo the same operation, thus ensuring that f t = Xception(x t ) has an output shape of (None, 15, 7, 7, 512); next, a global max pooling layer (GlobalMaxPooling2D) is used to perform pooling operations on the 7×7 feature maps of each frame to compress them into a 512-dimensional feature vector; after passing through the GlobalMaxPooling2D layer, the output feature dimension is (None, 15, 512), that is, the 15 frames of each video segment are compressed into 15 512-dimensional feature vectors; subsequently, the feature vectors are respectively input into the two-stream LSTM module to effectively capture temporal information;
[0087] The dual-stream LSTM module includes a bidirectional LSTM structure for time series mapping and a LSTM structure for sequence embedding;
[0088] The bidirectional LSTM structure of temporal mapping is used to extract the temporal dependency of context from short clips, and the LSTM structure of sequence embedding is used to extract the semantic information of each frame and capture the short-term temporal relationship between each frame.
[0089] The dual-stream LSTM module uses a bidirectional LSTM structure with time-series mapping in one stream, such as Figure 2 As shown in the figure, the characteristic is that the forward LSTM and the backward LSTM are used to process the input data at the same time, thereby enhancing the model's modeling ability for temporal information; it is used to extract the temporal dependency of the context from the short clip. The structure focuses on processing the dynamic information between each time step, capturing the temporal series relationship, combining the input of each time step, gradually updating the hidden state, and compressing the feature sequence of each 15 frames into a 128-dimensional vector; the other stream adopts the sequence embedding LSTM structure (Sequence Embedding LSTM). On the basis of the traditional LSTM, each frame of input data is embedded into a high-dimensional space, and the semantic information of each time step (i.e., each frame) is extracted as the embedding representation of the sequence, aiming to capture the short-term temporal relationship between each frame. The structure focuses on building a high-level semantic representation for the features of each time step, and finally retains the 128-dimensional representation of the feature map of each frame. The two LSTM streams work together to comprehensively extract the spatiotemporal information in the surgical action, capture the overall motion characteristics within the clip, and model the dynamic relationship between each frame. Finally, the model fuses the prediction scores output by the two LSTM streams through matrix addition to obtain the final prediction score of the surgical stage. The specific parameters and implementation details of the network architecture are shown in Table 4. It aims to improve the accuracy and robustness of surgical video analysis through sophisticated temporal modeling and multi-level feature learning;
[0090] Table 4 shows the relevant parameters of the Xception-two-stream LSTM recognition model:
[0091]
[0092] Step S4: Use the data set to train and optimize the Xception-two-stream LSTM recognition model;
[0093] The data set is divided into a training set, a validation set, and a test set. The training set is used for training and adjusting the parameters of the Xception-two-stream LSTM recognition model, and the test set is used for the accuracy test of the Xception-two-stream LSTM recognition model. During this period, the parameters of the Xception-two-stream LSTM recognition model are continuously adjusted, and finally an Xception-two-stream LSTM recognition model with the best accuracy and parameters are selected as the initialization Xception-two-stream LSTM recognition model.
[0094] The Xception-two-stream LSTM recognition model uses a repeated mechanism for data loading during training, loading data in batches in each training cycle;
[0095] The Adam optimizer is used to initialize the recognition model during the training of the Xception-two-stream LSTM recognition model. The initial learning rate, training dropout rate, training cycle and training batch of the Xception-two-stream LSTM recognition model are set. The Xception-two-stream LSTM recognition model uses the dropout rate method to perform random neglect training. The sigmoid and tanh activation functions are used. The classification cross entropy loss function is used to evaluate the progress of the Xception-two-stream LSTM recognition model.
[0096] Due to the huge scale of video frame data, a single training cannot carry a large video data load of about 130GB. Therefore, in order to efficiently process the data set, a repeated mechanism for data loading is adopted, and the data is loaded in batches in each training cycle to ensure that a valid data window is always maintained in the memory, and the multi-threaded loading technology is used to accelerate the reading and preprocessing of data, thereby improving the training efficiency; it is implemented in the Python3.8 environment and relies on TensorFlow-GPU 2.3.4 for training and optimization of deep learning models. The hardware platform uses a four-card NVIDIA GeForce RTX 3090 graphics card and adopts data parallel distributed training. The sample data of each batch is evenly distributed to multiple graphics cards to ensure computing power and accelerate the training process.
[0097] The following hyperparameter settings are mainly used to optimize the model training process:
[0098] Optimizer: Adam optimizer, which combines the advantages of momentum method and RMSProp algorithm. It can dynamically adjust the learning rate of each parameter during training, thereby accelerating convergence and improving the stability of the training process.
[0099] LSTM dropout rate: 0.25, which means that 25% of neurons will be randomly ignored in each training iteration. This prevents overfitting and improves the generalization ability of the model.
[0100] Initial learning rate: 0.0001. Using a smaller initial learning rate helps achieve more stable convergence in the initial stage of training, avoid overly large weight updates, and thus improve the stability and performance of the model. To further optimize the training process, a learning rate decay strategy is adopted. Every 10 epochs of training, the learning rate is reduced to 0.1 of the original value to enable more refined parameter adjustment in the later stage of training.
[0101] Number of training epochs (Epochs): 100. The model will perform 100 iterations on the entire training dataset to ensure that the model fully learns the features of the data and can adapt to complex task requirements.
[0102] LSTM activation functions: sigmoid and tanh activation functions. The sigmoid activation function is applied to the gating mechanisms (including the forget gate, input gate, and output gate) to determine whether to retain or discard the extracted information and smoothly control the flow of information; the tanh activation function is used for updating the candidate memory cells. It has symmetry and can effectively promote the propagation of gradients, alleviate the vanishing gradient problem, and thus improve the efficiency and performance of training.
[0103] Loss function: Categorical Cross - Entropy. This loss function can effectively reflect the prediction accuracy of the model for each category by calculating the difference between the predicted probability distribution and the actual labels, and promote the optimization and performance improvement of the model in multi - category problems. The formula is as follows:
[0104]
[0105] where C is the number of categories; y c is the true label; is the predicted probability obtained by classifying the output of the last time step of the LSTM; is the natural logarithm of the predicted probability to ensure that the loss function imposes a large penalty on incorrect predictions.
[0106] Batch size: 256. When updating the model parameters each time, 256 samples are used as a batch for forward propagation and backpropagation calculations. The selection of this batch size helps achieve a good balance between computational efficiency and memory resources.
[0107] Through the above hyperparameter settings, the training process of the model is optimized, ensuring good convergence performance and strong generalization ability.
[0108] Step S5: Select the Xception - two - stream LSTM recognition model with the highest training accuracy to perform recognition scoring on the surgical video.
[0109] Example 3:
[0110] Experiment and Result Analysis:
[0111] To comprehensively evaluate the effectiveness of the Xception-two-stream LSTM recognition model in gallbladder resection stage recognition, especially in the case of class imbalance, four experiments were conducted, covering model architecture, the number of input video frames, data balancing processing, and comparison with existing models. Through these experiments, the effects of various factors on the model performance were explored, and the advantages of the model were also evaluated.
[0112] 1.1 Model Performance Evaluation Metrics:
[0113] Due to the temporal dependence among surgical stage videos, the Cholec80 dataset was divided into a training set, a validation set, and a test set in a 7:2:1 ratio based on the time sequence. All video frames were adjusted to 224×224 pixels and normalized. To evaluate the model's performance in the stage recognition task, detailed statistical analysis was performed on the data in the test set, and multiple evaluation metrics were adopted to address the class frequency imbalance problem in the gallbladder resection stage recognition task.
[0114] Specifically, they include: Accuracy: representing the proportion of correctly predicted samples in the total number of samples; Precision: representing the proportion of truly positive samples among all samples predicted as positive; Recall: representing the proportion of truly positive samples that are successfully predicted as positive; F1-score: the harmonic mean of Precision and Recall, comprehensively considering the model's precision and recall ability; Specificity: representing the proportion of correctly predicted negative samples among all truly negative samples; ROC curve and Area Under Curve (AUC): By plotting the ROC curve and calculating the area under the curve (AUC), the closer the AUC value is to 1, the stronger the model's class discrimination ability. These metrics can provide a more comprehensive model evaluation in the case of class imbalance. Especially when the dataset has class frequency imbalance, metrics such as F1-score and AUC can better reflect the actual performance of the model than simple accuracy.
[0115] 1.2 Comparison Results of Stage Recognition Performance among Different CNN Models:
[0116] To verify the advantages of the Xception model in feature extraction for this task, it was compared with several other common CNN architectures, including: VGG16, AlexNet, ResNet50, and InceptionV3. These networks were all used as feature extraction modules, and their output results were respectively passed as inputs to the two-stream LSTM module for subsequent stage recognition verification. The specific experimental results are shown in Table 5. The results indicate that the Xception module achieved the best performance in feature extraction for the stage recognition task, with an F1 score of 77.31%, higher than other CNNs. In addition, the Xception module also outperformed other CNNs in other evaluation metrics. These results show that the Xception module can more effectively extract visual features suitable for subsequent temporal modeling by the two-stream LSTM network from single-frame images of each stage, thereby improving the overall performance of the model in the stage recognition task.
[0117] Table 5 shows the comparison results of the overall stage recognition performance of different CNN architectures:
[0118] CNN F1 - score (%) AUC (%) ACC (%) PR (%) RE (%) SP (%) VGG16 76.51 96.32 84.06 76.46 75.61 96.14 AlexNet 76.22 95.86 83.35 75.21 76.50 96.79 ResNet50 76.98 95.33 83.95 77.37 76.04 96.33 InceptionV3 77.05 95.68 84.27 76.32 75.88 95.91 Xception 77.31 96.52 84.74 78.04 76.69 96.49
[0119] In summary, the Xception module outperforms VGG16, AlexNet, ResNet50, and InceptionV3 in all evaluation metrics. Therefore, the Xception module was selected as the basic architecture of the middle convolutional neural network (CNN) module. The depthwise separable convolution in Xception significantly reduces the computational amount and the number of parameters compared with the traditional fully connected convolution operations in VGG16 and AlexNet, while improving the efficiency and expressiveness of feature extraction. Since it abandons the complex multi-branch parallel convolution operations in InceptionV3 and adopts a more efficient modular structure, it reduces redundant calculations and further enhances the ability to extract features. At the same time, Xception draws on the residual connection design in ResNet, allowing the network to skip some layers during training and directly pass the input information to subsequent layers, solving the gradient vanishing problem in deep networks and enhancing the stability of training. These designs enable the Xception network to demonstrate stronger generalization ability and more excellent recognition performance compared with other network architectures in the current surgical stage recognition task.
[0120] 1.3 Determination of the number of frames in a single video clip:
[0121] For short - segment - based stage recognition, the number of frames in a video directly affects the model's ability to capture temporal information. If the number of frames is small, the model may not be able to capture long - term temporal information in the video; if the number of frames is too large, it may lead to information redundancy and increase computational complexity. Moreover, when the total amount of data is fixed, a smaller number of frame settings can obtain more training samples, which helps to increase sample diversity and improve the generalization ability of the model. However, if the number of frames is too small, the information of each video is likely to be overly simplified; with a larger number of frame settings, the computational resources and storage space required for the video are larger, the total number of training samples decreases, and the model is prone to overfitting. Considering that the in - cavity operations of the surgeon during laparoscopic cholecystectomy are all stable and small - amplitude movements, and the feature changes between consecutive frames are few, choosing too long a video frame may result in feature redundancy.
[0122] Considering the above factors, the experiment was conducted in a relatively low frame - number range. The accuracy, F1 - score, AUC, and running time per batch were compared when the number of frames in a single segment was set to 5, 10, 15, 20, and 25. All comparisons were trained and calculated using the same dataset and the sampling method with the original sampling rate (Sampling Rate = 1), keeping the LSTM module architecture of the Xception - two - stream LSTM model unchanged and only modifying the input number of frames. The comparison results are shown in Table 6.
[0123] The experimental results show that as the input number of frames increases, the overall accuracy of the model shows an upward trend, and the training cycle shows a shortening trend. However, after exceeding 15 frames, the performance improvement tends to be stable; its F1 - score and AUC value show a trend of rising first and then falling, indicating that when the number of frames increases and the overall training data decreases, the model's ability to recognize low - frequency categories decreases. Considering computational resources comprehensively, it is recommended to use 15 frames as the best input configuration.
[0124] Table 6 shows the model comparison results for different numbers of frames in a single segment:
[0125] Number of frames ACC (%) F1 - score (%) AUC (%) Running time per batch (minutes) 5 83.15 75.56 95.43 25 10 84.34 76.88 96.08 22 15 84.74 77.31 96.52 20 20 84.78 77.21 96.41 19 25 84.83 77.26 96.32 19
[0126] 1.4 Effectiveness of the dynamic data balancing strategy for the problem of uneven class frequencies:
[0127] To verify the effectiveness of the dynamic data balancing strategy in the stage recognition of cholecystectomy, the F1 - score and AUC value of the seven - stage recognition task of cholecystectomy were compared when applying uniform fixed 1 frame per second (fps) undersampling and applying the dynamic data balancing strategy.
[0128] To ensure the fairness of the experiment, the two groups of experiments were conducted under the same conditions: each video segment consisted of 15 frames, the Xception - two - stream LSTM model architecture was used, and the above - mentioned network parameters and hyperparameters were used for model training. The experimental results of the two groups are shown in Table 7.
[0129] Table 7 shows the F1 scores and AUC values for applying the fixed undersampling rate strategy and the dynamic data balancing strategy. Among them, (1) indicates the use of the fixed undersampling rate strategy; (2) indicates the use of the dynamic data balancing strategy:
[0130]
[0131] The experimental results show that by calculating the F1 scores of the seven stages when using the fixed undersampling rate strategy, it is verified that there is a certain degree of data imbalance problem in this sampling method, and the recognition effect of the model in the few-sample stage is not good. Compared with using the fixed undersampling rate strategy, after introducing the dynamic data balancing strategy, the model shows significant improvements in both the overall F1 score and the AUC value. The comprehensive F1 score has increased by 10.75%, and the AUC value has increased by 1.67%. To a certain extent, it alleviates the problem of data bias.
[0132] Both the F1 score and the AUC value are important indicators for evaluating the performance of the model when dealing with imbalanced datasets. However, they are single numerical metrics and are affected by multiple factors, including the selection of test data, the distribution of data, and the randomness during the experiment (such as the division of the training set and the test set). Therefore, the results of a single experiment may fluctuate and it is difficult to accurately reflect the potential impact of these factors on the model's performance. To evaluate whether the F1 score and the AUC value are statistically significant after applying the dynamic data balancing strategy and to determine the effectiveness of this strategy for the problem of uneven class frequencies, the recognition results of the 1st, 3rd, 5th, 6th, and 7th stages (the total number of frames in these five stages is relatively small and can better reflect whether the dynamic data balancing strategy is effective, so they are taken as the key objects of investigation) are selected for analysis. The two-sample t-test method is used for verification. This method can be used to test whether there is a significant difference between the means of two independent samples (groups). By calculating the t-statistic and combining the degrees of freedom, and then looking up the t-distribution table to obtain the p-value, the p-value reflects the probability of observing the current difference under the premise that the null hypothesis (that is, assuming that the means of the two groups are not different) holds. If the p-value is less than the set significance level (usually 0.05), it is considered that the difference is statistically significant, thus verifying whether the effect of the dynamic data balancing strategy is significant. The formula for the t-statistic is as follows:
[0133]
[0134] Among them, and are the sample means of the two sets of data, S1 and S2 are the sample variances of the two sets of data, and n1 and n2 are the sample sizes of the two sets of samples.
[0135] The formula for calculating the degrees of freedom is as follows:
[0136] df = n1 + n2 - 2
[0137] As Figure 3 and Figure 4 shown, the trends of F1 scores and AUC values in stages 1, 3, 5, 6, and 7 and the t-test results (* indicates the significance level of statistical differences) are presented. The results show that after adopting this strategy, significant statistical differences in F1 scores in the 3rd, 5th, 6th, and 7th stages are shown compared with the fixed undersampling strategy (p < 0.05), and the significant differences in the 5th and 7th stages are more prominent (p < 0.01). Meanwhile, after adopting this strategy, significant statistical differences in AUC values in the 3rd, 5th, and 6th stages are also shown compared with the fixed undersampling strategy (p < 0.05). These results indicate that the F1 scores and AUC values obtained from this experiment are statistically significant.
[0138] As Figure 5 shown, where Figure 5 (a) is the fixed undersampling rate strategy; Figure 5 (b) is the dynamic data balancing strategy, showing a comparison of confusion matrices when applying the fixed undersampling rate strategy and the dynamic data balancing strategy. The brighter the color of the area in the matrix, the higher the correct classification proportion of the corresponding category. According to Figure 5 the results, compared with before applying the dynamic data balancing strategy, the classification results of the model in each stage have been improved. This change indicates that by dynamically balancing the sample numbers in each stage of each video, the model's biased learning on certain categories is effectively avoided, thereby improving the overall classification performance. The above improvements are not only reflected in the increase of F1 scores and AUC values but also further verified by the smaller p-values and the reduction of the misclassification proportion in the confusion matrix. The above results show that the proposed dynamic data balancing strategy plays a positive role in improving the accuracy and stability of the model.
[0139] In summary, a dynamic data balancing strategy is proposed, which makes the training sample numbers in each stage tend to be balanced by applying different sampling rates in each surgical stage of each video. This strategy dynamically sets the sampling rates in each stage of each video according to the total sample amount, increasing the sample quantity and overall data complexity for the relatively data-scarce stages 1, 3, 5, 6, and 7. The experimental results prove that applying the dynamic data balancing strategy can effectively improve the accuracy of surgical stage recognition and the overall performance of the model.
[0140] 1.5 Comparison with previous models:
[0141] The Xception-two-stream LSTM model was compared with several existing classical models to evaluate its performance in the task of cholecystectomy stage recognition. The models compared included PhaseLSTM, EndoLSTM, ResNetLSTM, and EndoN2N, all of which adopted a structure that cascaded a convolutional neural network (CNN) with a recurrent neural network (RNN).
[0142] The proposed model structure adopted the Xception module as the basic feature extraction model, and two parallel long short-term memory networks (Long Short-Term Memory, LSTM) were used to further extract more temporal feature information to a greater extent. The performance comparison results of different models in the surgical stage recognition task are shown in Table 8. The experimental results show that the F1 score of the proposed model is 88.06%, which is about 6% higher than that of the previously better-performing EndoN2N model. In addition, the model reached 98.19% and 91.51% in terms of AUC and accuracy, respectively. These results fully demonstrate the advantages of the model in the surgical stage recognition task.
[0143] Table 8 shows the comparison results of the existing models and the model based on the Cholec80 dataset:
[0144]
[0145]
[0146] In summary, when the proposed Xception-two-stream LSTM model was compared with other cholecystectomy video stage recognition models based on the CNN-LSTM architecture, it showed significant stage recognition performance. This improvement can be attributed to the ability of the Xception module to extract basic visual features, and the two-stream LSTM network constructed by the parallel temporal mapping bidirectional LSTM structure and the sequence embedding LSTM structure, which can effectively reduce the loss of time information and enhance the model's learning ability of temporal features. Compared with a single LSTM architecture, this model can not only capture the motion changes between image frames, but also understand and extract the surgical operation patterns in short video segments as a whole, thus providing more accurate results when recognizing surgical stages.
[0147] Through short-time segment division and dynamic sampling rate allocation, the present invention effectively alleviates the problem of biased learning, effectively reduces the impact of data volume differences on the model learning process, reduces bias, and improves the overall recognition accuracy of the model for surgical stages; it can simultaneously fuse temporal dynamic information and static features, enhancing the accuracy and robustness of surgical stage recognition; by parallel processing of short-term and long-term temporal dependence information, it significantly enhances the understanding of the surgical process; it can not only effectively solve the problem of data imbalance and more comprehensively capture temporal feature information, thus showing higher classification performance in the surgical stage recognition task; it can provide more structured support for surgical training, quality control and clinical decision-making, and further lay a foundation for the research and development of intelligent surgical assistance systems.
[0148] The above embodiments are only the preferred embodiments of the present invention, and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art on the basis of the present invention fall within the scope of protection required by the present invention.
Claims
1. A surgical stage identification method based on a dynamic data balancing strategy, characterized in that: The following steps are involved: Step S1: Collecting raw surgical data and annotating the collected raw data; Step S2: preprocessing the annotated raw data using a dynamic data balancing strategy to form a data set; Step S3: Build the Xception-two-stream LSTM recognition model; Step S4: Use the data set to train and optimize the Xception-two-stream LSTM recognition model; Step S5: Select the Xception-two-stream LSTM recognition model with the highest training accuracy to perform recognition and scoring on the surgical video.
2. A surgical stage identification method based on a dynamic data balancing strategy as claimed in claim 1, characterized in that: The surgical raw data in step S1 is the raw data in the Cholec80 dataset, including video clips of different cholecystectomy operations, each of which is marked with different surgical operation processes, and the video clips are annotated in detail in time series; the video clips record the complete process from the beginning to the end of the operation, and the video clips include the preparation stage, triangulation stage, blood vessel cutting stage, gallbladder dissection stage, gallbladder bagging stage, blood clot or coagulation tissue cleaning stage and gallbladder recovery stage.
3. A surgical stage identification method based on a dynamic data balancing strategy as claimed in claim 1, characterized in that: In step S2, the video clip is undersampled into a short clip consisting of a set number of frames, that is, the sampling rate of the surgical stage of the video clip is calculated according to the stride of the surgical stage, and the image frames of the surgical stage are undersampled according to the stride. The stride calculation uses the following formula: Where S is the stride, which indicates the sampling interval, rounded to the nearest integer; R f is the frame rate of the original video, F is the number of image frames in the current surgical stage, and N is the maximum number of image frames in all surgical stages.
4. A surgical stage identification method based on a dynamic data balancing strategy as claimed in claim 1, characterized in that: The Xception-two-stream LSTM recognition model in step S3 includes an Xception module and a two-stream LSTM module. The Xception module is used to extract the visual features of the data set and form a feature vector. The two-stream LSTM module extracts the spatiotemporal information in the surgical action from the feature vector, captures the overall motion features in the video clip, calculates the dynamic relationship between each frame, and fuses the prediction scores output by matrix addition to finally determine the surgical stage prediction score.
5. A surgical stage identification method based on a dynamic data balancing strategy as claimed in claim 4, characterized in that: In step S3, the Xception module includes multiple stacked depth-separable convolutional layers, TimeDistributed layers and maximum pooling layers; the depth-separable convolutional layer performs depth and point-by-point convolution operations on the data set data to extract visual feature data in the data set, and the TimeDistributed layer performs consistency processing on the visual feature data extracted by the depth-separable convolutional layer; the maximum pooling layer realizes pooling processing of the visual feature data.
6. A surgical stage identification method based on a dynamic data balancing strategy as claimed in claim 4, characterized in that: The dual-stream LSTM module in step S3 includes a bidirectional LSTM structure for time series mapping and a LSTM structure for sequence embedding; Among them, the bidirectional LSTM structure of temporal mapping is used to extract the temporal dependency of context from short clips, and the LSTM structure of sequence embedding is used to extract the semantic information of each frame and capture the short-term temporal relationship between each frame.
7. A surgical stage identification method based on a dynamic data balancing strategy as claimed in claim 1, characterized in that: In step S4, the data set is divided into a training set, a validation set and a test set. The training set is used for training and adjusting parameters of the Xception-two-stream LSTM recognition model, and the test set is used for accuracy testing of the Xception-two-stream LSTM recognition model. During this period, the parameters of the Xception-two-stream LSTM recognition model are continuously adjusted, and finally an Xception-two-stream LSTM recognition model and parameters with the best accuracy are selected as the initialization Xception-two-stream LSTM recognition model.
8. A surgical stage identification method based on a dynamic data balancing strategy as claimed in claim 7, characterized in that: In step S4, a repetitive mechanism of data loading is adopted during the training of the Xception-two-stream LSTM recognition model, and data is loaded in batches in each training cycle.
9. A surgical stage identification method based on a dynamic data balancing strategy as claimed in claim 7, characterized in that: In step S4, the Adam optimizer is used to initialize the recognition model during the training of the Xception-two-stream LSTM recognition model, and the initial learning rate, training discard rate, training cycle and training batch of the Xception-two-stream LSTM recognition model training are set. The Xception-two-stream LSTM recognition model adopts the discard rate method to perform random neglect training, and uses sigmoid and tanh activation functions. The progress of the Xception-two-stream LSTM recognition model is evaluated using the classification cross entropy loss function.
Citation Information
Patent Citations
ConvNeXt-based multi-stage adaptive class balance operation stage identification method
CN119516429A
Cited By
Operation stage identification method based on time sequence modeling
CN121839029A
A surgical phase recognition method based on time series modeling
CN121839029B