Laparoscopic surgery stage identification system based on dual-granularity time convolution
By designing a laparoscopic surgical stage recognition system based on dual-particle time convolution, the problems of the model lacking flexibility and scalability, high risk of overfitting, and insufficient feature fusion ability in the prior art are solved, and higher recognition accuracy and reliability are achieved, and dynamic monitoring capabilities of the surgical process are enhanced.
Patent Information
- Application Number
- CN202510292057.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art realizes laparoscopic surgical phase recognition by using a two-particle time convolution network, but cannot effectively deal with the features of specific tasks, lacks real-time processing capabilities for new data, lacks flexibility and scalability of the model, increases the risk of overfitting, and cannot fuse and classify global and local features in the time series, resulting in inaccurate identification results.
A laparoscopic surgical stage recognition system based on two-particle size time convolution is designed, including a supervision platform, data acquisition module, data preprocessing module, feature extraction module, model training module, stage identification module, result evaluation module and integrated application module. The system extracts multi-scale features through a two-particle time convolution model and recognizes them through the fusion of global and local features, and trains the model in real time to improve adaptability and accuracy.
It significantly improves the accuracy and reliability of identification in the surgical stage, reduces the risk of overfitting, improves the adaptability and dynamic optimization capabilities of the model, and enhances the dynamic monitoring capabilities of the surgical process and the systematic evaluation capabilities of the identification results.
Smart Images

Figure CN120196992A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of laparoscopic surgery, and more specifically, to a laparoscopic surgery stage recognition system based on dual-granularity temporal convolution. Background Art
[0002] The laparoscopic surgery stage recognition system analyzes and processes the images and videos during the surgery by using computer vision and deep learning technologies, so as to realize the automatic recognition of the surgery stage. This system can track the surgery process in real time, provide accurate surgery stage information for doctors, help doctors better master the surgery rhythm, and improve the surgery quality.
[0003] The patent application with the publication number CN114372962A discloses a laparoscopic surgery stage recognition method and system based on dual-granularity temporal convolution, including: constructing a laparoscopic surgery data set; using the dual-granularity temporal convolution module of the dual-granularity temporal convolution network to perform preliminary feature extraction on the picture sequence and output the initial prediction result for each frame of image; using the single-granularity temporal convolution module of the dual-granularity temporal convolution network to correct the initial prediction result output by the dual-granularity temporal convolution module; mapping the prediction result to the interval (0, 1) to obtain the final surgery stage recognition result; the present invention uses the dual-granularity temporal convolution network to realize laparoscopic surgery stage recognition, has higher accuracy and better generalization ability in different backgrounds, can accurately detect different types of surgery stages, and can solve the problem in the field of deep learning that it can identify the surgery stage category but is difficult to accurately distinguish the stage transition frames by using the visual and temporal information of the surgery video;
[0004] However, the existing technology realizes laparoscopic surgery stage recognition by using the existing dual-granularity temporal convolution network, cannot effectively handle the characteristics of specific tasks, has insufficient real-time processing ability for new data, the existing dual-granularity temporal convolution model lacks flexibility and scalability, increases the risk of overfitting, and lacks transparency; at the same time, although the existing technology can extract the global information and local information of the temporal features through the dual-granularity temporal convolution module, it cannot fuse and classify the global features and local features in the time series, cannot accurately identify the stage of the surgery, cannot systematically evaluate the recognition result, cannot provide targeted feedback for the analysis of the recognition result, cannot promote the dynamic optimization and continuous learning of the stage recognition model, and reduces the accuracy and reliability of the recognition of the stage of the surgery.
[0005] Therefore, we propose a laparoscopic surgery stage recognition system based on dual-granularity temporal convolution for the above problems. Summary of the Invention
[0006] The object of the present invention is to provide a laparoscopic surgery stage recognition system based on dual-granularity temporal convolution, which solves the problems in the prior art that when using the existing dual-granularity temporal convolution network to realize laparoscopic surgery stage recognition, it is unable to effectively handle the characteristics of specific tasks, has insufficient real-time processing ability for new data, the existing dual-granularity temporal convolution model lacks flexibility and scalability, increases the risk of overfitting, and lacks transparency; at the same time, although the prior art can extract the global information and local information of temporal features through the dual-granularity temporal convolution module, it is unable to fuse and classify the global features and local features in the time series, cannot accurately identify the stage of the surgery, cannot systematically evaluate the recognition results, cannot provide targeted feedback for the analysis of the recognition results, cannot promote the dynamic optimization and continuous learning of the stage recognition model, and reduces the accuracy and reliability of the recognition of the stage of the surgery.
[0007] The object of the present invention is achieved through the following technical solutions:
[0008] A laparoscopic surgery stage recognition system based on dual-granularity temporal convolution includes a supervision platform, a data acquisition module, a data preprocessing module, a feature extraction module, a model training module, a stage recognition module, a result evaluation module, and an integrated application module;
[0009] The data acquisition module is used to collect the surgical stage data during the laparoscopic surgery in real time, and the surgical stage data includes intra-abdominal pressure data, surgical instrument usage data, vital sign data, and video data;
[0010] The data preprocessing module is used to preprocess the collected surgical stage data;
[0011] The feature extraction module applies a dual-granularity temporal convolution model to extract key features from the preprocessed data;
[0012] The model training module trains the dual-granularity temporal convolution model according to the surgical stage data with real labels;
[0013] The stage recognition module uses the trained dual-granularity temporal convolution model to perform real-time recognition on new data, judge the current surgical stage, and output the corresponding recognition result;
[0014] The result evaluation module is used to systematically evaluate the results of laparoscopic surgery stage recognition;
[0015] The integrated application module is used to integrate the recognition results with the clinical decision support unit and the surgical robot.
[0016] As a preferred embodiment of the present invention, the specific steps for the data preprocessing module to preprocess the collected surgical stage data are as follows:
[0017] Data cleaning: Process the signal using a filter to remove electromagnetic interference and background noise, fill in missing values using interpolation, and delete records with a large number of missing values;
[0018] Data standardization: Convert the data to a distribution with a mean of 0 and a standard deviation of 1 using the following formula:
[0019] where X is the numerical sequence of the signal, μ is the mean, and σ is the standard deviation;
[0020] Data transformation: Extract time features and perform categorical data encoding;
[0021] Feature selection: Select features that are closely related to the target variable based on correlation analysis and remove redundant features;
[0022] Data partitioning: Divide the data into a training set, a validation set, and a test set, with an allocation ratio of 70% for training, 15% for validation, and 15% for testing.
[0023] As a preferred embodiment of the present invention, the specific process of the feature extraction module extracting key features from the preprocessed data is as follows:
[0024] Input data preparation: The preprocessed data obtained from laparoscopic surgery is represented as a multi-dimensional time series: X = {x1, x2, …, x t};
[0025] where X is the input multi-dimensional time series data set, containing t time steps, t is the total number of time steps, representing the total length of the data, and x t is the feature vector at time step t, represented as x t ∈R n , where n is the dimension of the feature;
[0026] Convolutional neural network design: Feature extraction is implemented through a convolutional neural network to extract important features in the time series data;
[0027] Use a one-dimensional convolutional kernel to process the time series data. Assume we define a convolutional kernel W:
[0028] W = {w1, w2, …, w k};
[0029] where W is the convolutional kernel, with a size of K, used to capture local patterns in the time series;
[0030] The result of the convolution operation is:
[0031] where Z t is the convolution output at time step t, representing the current feature response, and wm is the m-th value in the convolutional kernel W, and b is the bias term, a scalar value used to adjust the convolutional output;
[0032] Apply a non-linear activation function to the convolutional output to introduce non-linearity:
[0033] Z t ' = f(Z t ) = max(0, Z t );
[0034] where Z t ' is the convolutional result after being processed by the activation function, representing the non-linear form of the current feature response.
[0035] As a preferred embodiment of the present invention, dual-granularity convolution uses two different granularities to extract features, namely coarse-grained and fine-grained;
[0036] Coarse-grained convolution uses a larger convolutional kernel and stride to extract long-term features, defined as follows: the convolutional kernel is W1, with size K1 and stride S1;
[0037] The convolutional formula for coarse-grained is as follows:
[0038] where Z large [t] is the output of the coarse-grained feature at time step t, w 1,m is the m-th weight in the convolutional kernel W1, and b1 is the bias term of the coarse-grained convolution;
[0039] Fine-grained convolution applies a smaller convolutional kernel and uses a smaller stride to capture short-term features, defined as follows: the convolutional kernel is W2, with size K2 and stride S2;
[0040] The convolutional formula for fine-grained is as follows:
[0041] where Z small [t] is the output of the fine-grained feature at time step t, w 2,m is the m-th weight in the convolutional kernel W2, and b2 is the bias term of the fine-grained convolution;
[0042] Feature fusion: Generate a fused feature through a concatenation operation: Z fused = [Z large , Z small ;
[0043] where Z fused is the fused feature vector, containing features from both coarse-grained and fine-grained convolutions;
[0044] Fuse features through weighted averaging: Z fused = αZlarge +(1 - α)Z small ;
[0045] Where α is a weight coefficient, 0 ≤ α ≤ 1, which controls the contribution degrees of the two features;
[0046] Fully connected layer: After feature fusion, these features are input into the fully connected layer to generate the final feature representation: F = σ(W f ·Z fused +b f );
[0047] Where F is the final feature vector, the output after being processed by the fully connected layer, W f is the weight matrix of the fully connected layer, each row corresponding to one class, and b f is the bias vector of the fully connected layer, with the length the same as the output feature dimension, and σ is the activation function;
[0048] Feature selection: For the extracted features, apply feature selection methods to remove redundant or useless features;
[0049] Output features: The finally output feature F will be transmitted to the subsequent stage recognition module through the supervision platform.
[0050] As a preferred embodiment of the present invention, the specific steps for the model training module to train the dual - granularity time - convolutional model are as follows:
[0051] Forward propagation: Input each batch of the training set into the dual - granularity time - convolutional model, and calculate the output layer by layer: Calculate the feature map, process the input data through convolution, adding bias, and activation function, combine the coarse - granularity and fine - granularity features into a vector, perform linear combination and activation on the feature vector, and predict the class;
[0052] Calculate the loss: Use the cross - entropy loss function to calculate the difference between the output of the model and the true label. The formula for cross - entropy is:
[0053] Where y m is the true label, is the probability predicted by the model;
[0054] Backward propagation: According to the partial derivative of the loss function with respect to the model parameters, calculate the gradient of each layer through the chain rule, and use the optimization algorithm to update the weights and biases of each layer. The update formula is:
[0055] θ = θ - η·▽L;
[0056] Where θ is the model parameter, η is the learning rate, and ▽L is the gradient of the loss function with respect to the parameter.
[0057] As a preferred embodiment of the present invention, the training data is divided into small batches, and each time a batch of data is input for forward propagation and backward propagation. Completing one traversal of the entire training set is called an epoch, and multiple epochs are required for training;
[0058] During the training process, the validation set is regularly input into the model for validation, and the following metrics are monitored: the loss of the model on the validation set, the accuracy of classification on the validation set. The performance of the model is monitored using the validation set. If various performance metrics no longer improve or overfitting occurs, the training is terminated early and the current best model is saved.
[0059] As a preferred embodiment of the present invention, the specific process of the stage recognition module for real-time recognition of new data is as follows:
[0060] Video frames and vital feature data are collected in real time using sensors and cameras. The video data is converted to a fixed size, and N frames are extracted from the video, with the timestamp of each frame being t k (k = 1, 2, …, N);
[0061] For recognition, a fixed number of consecutive frames are extracted as input data every T seconds, that is, the input is:
[0062] X = {I(t1), I(t2), …, I(t T )};
[0063] Among them, I(t k ) represents the image frame extracted at time t k ;
[0064] The local features of each frame are extracted through a fine-grained convolutional layer. For the input image I(t k ), it can be expressed as: V f (t k ) = W f *I(t k ) + b f ;
[0065] Among them, W f is the convolution kernel, b f is the bias term, * represents the convolution operation, and V f (t k ) is the local feature map;
[0066] The global features in the time series are extracted through a coarse-grained convolutional layer, and the data features of the entire time period are:
[0067] Among them, W c is a larger convolution kernel for combining time series features.
[0068] As a preferred embodiment of the present invention, fine-grained and coarse-grained features are combined into a feature representation:
[0069] F = f(V f , V c ), where f is a fusion function;
[0070] The fused feature F is classified through a fully connected layer:
[0071] Y = W y ·F + b y , where W y is the weight of the fully connected layer, and b y is the bias term;
[0072] For the output vector Y, calculate the probability of each stage:
[0073] where P(y m ) is the probability of stage y m ;
[0074] Select the stage y * with the highest probability: y * = arg max P(y m );
[0075] The recognition result is displayed in real time on the monitoring interface in the operating room;
[0076] A feedback mechanism is established to record the doctor's accuracy evaluation of the recognition result, and the dual-grained temporal convolutional model is continuously trained using the collected new data and feedback information.
[0077] As a preferred embodiment of the present invention, the specific process of the result evaluation module for systematically evaluating the result of laparoscopic surgery stage recognition is as follows:
[0078] Evaluate the performance of the model on the test set using the following metrics:
[0079] Accuracy: Calculate the proportion of correctly recognized stages in the total stages;
[0080] Recall: Measure the ability of the model to recognize a specific stage, that is, the proportion of the sum of true positives and false negatives;
[0081] Precision: Indicate the proportion of actually correct predictions among those recognized as a specific stage, that is, the proportion of the sum of true positives and false positives;
[0082] F1-score: Combine the harmonic mean of precision and recall to provide a single evaluation metric, suitable for cases of class imbalance;
[0083] Generate a confusion matrix to display the recognition results at each stage in detail, including the comparison between the true labels and the predicted labels, providing an intuitive view for model performance analysis;
[0084] Deeply analyze the misclassified cases, identify the reasons for the model's recognition failure, and feedback the evaluation results and analysis to the model training module.
[0085] As a preferred implementation manner of the present invention, the specific process of integrating the recognition results with the clinical decision support unit and the surgical robot by the integrated application module is as follows:
[0086] Combine the recognition results with the patient's historical data, generate complete patient situation information through a data fusion algorithm, and the clinical decision support unit generates implementation suggestions based on the latest recognition results and the patient's condition;
[0087] Convert the instructions from the decision support unit into specific control commands for the robot, and transmit the instructions to the surgical robot in real time through a low-latency network protocol, and the surgical robot executes actions according to the received instructions;
[0088] Continuously monitor the working status of the surgical robot, the operation of medical equipment, and the vital sign data of the patient;
[0089] Record all data during the operation to form a complete surgical data file, and review and analyze the collected data after the operation;
[0090] Provide an intuitive graphical user interface to display surgical information, stage recognition results, and decision support suggestions, and provide multiple interaction methods for surgeons to quickly obtain information and perform operations;
[0091] Collect the usage feedback of surgeons on the system, and continuously optimize and upgrade the recognition algorithm, decision support system, and user interface based on the feedback information.
[0092] Compared with the prior art, the advantages of the present invention are as follows:
[0093] (1) In the present invention, the dual-granularity temporal convolutional model is trained in real time by the model training module according to the existing real surgical data, which can accurately capture the multi-scale features during the operation, thereby significantly improving the recognition accuracy and precision of the model. Moreover, the training based on real data ensures the adaptability of the model in actual applications, reduces the risk of overfitting, and can be optimized for specific surgical types to improve the effect and reliability of surgical stage recognition;
[0094] (2) In the present invention, the stage recognition module recognizes the stage of the operation, fuses and classifies the global features and local features in the time series during the recognition process, can quickly judge the current surgical stage and output accurate results, and improves the dynamic monitoring ability of the operation process;
[0095] (3) In the present invention, the result evaluation module systematically evaluates the recognition results, can independently verify the recognition results, improve the accuracy and reliability of recognition. Through comprehensive analysis of the output results, targeted feedback can be provided to promote the dynamic optimization and continuous learning of the recognition model. Moreover, the accumulated evaluation data not only helps to improve the model performance, but also provides decision-making support for surgeons, enhancing the safety and effectiveness of the surgery. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] Figure 1 is the system block diagram of the present invention;
[0097] Figure 2 is the system block diagram of Embodiment 1 in the present invention;
[0098] Figure 3 is the method flowchart for the model training module to train the dual-granularity temporal convolutional model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0099] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0100] Embodiment 1: As Figure 1 shown, the laparoscopic surgery stage recognition system based on dual-granularity temporal convolution proposed by the present invention includes a supervision platform, a data acquisition module, a data preprocessing module, a feature extraction module, a model training module, a stage recognition module, a result evaluation module, and an integrated application module;
[0101] The supervision platform is bidirectionally communicatively connected to the data acquisition module, the data preprocessing module, and the feature extraction module. The supervision platform is unidirectionally communicatively connected to the model training module and the integrated application module. The model training module is bidirectionally communicatively connected to the stage recognition module. The stage recognition module is unidirectionally communicatively connected to the result evaluation module and the integrated application module. The result evaluation module is unidirectionally communicatively connected to the model training module;
[0102] The data acquisition module is used to collect the surgical stage data during the laparoscopic surgery in real time. The surgical stage data includes intra-abdominal pressure data, surgical instrument usage data, vital sign data, and video data;
[0103] The data acquisition module can collect comprehensive and accurate surgical environment information in real time. These multi-dimensional data not only provide a rich foundation for subsequent analysis and model training, but also can monitor the surgical status in real time, helping surgeons to identify potential problems in a timely manner, thereby improving the safety and effectiveness of the operation.
[0104] A data preprocessing module, used for preprocessing the collected surgical stage data;
[0105] The specific steps of preprocessing the collected surgical data by the data preprocessing module are as follows:
[0106] Data cleaning: Use filters to process signals, remove electromagnetic interference and background noise, use interpolation to fill missing values, and delete records with many missing values;
[0107] Data standardization: Convert the data to a distribution with a mean of 0 and a standard deviation of 1 using the following formula:
[0108] Where X is the numerical sequence of the signal, μ is the mean, and σ is the standard deviation;
[0109] Data conversion: extract time features and encode classified data;
[0110] Feature selection: Based on correlation analysis, select features that are closely related to the target variable and remove redundant features;
[0111] Data partitioning: The data is divided into training set, validation set and test set, with the allocation ratio of 70% training, 15% validation and 15% test;
[0112] The data preprocessing module effectively processes the collected surgical stage data through a series of specific steps, which can ensure the quality and consistency of the data, thereby improving the accuracy of subsequent analysis and model training. First, data cleaning removes interference and noise through filtering and missing value processing to ensure data reliability. Secondly, data standardization unifies data of various dimensions to the same standard to facilitate model learning. The data conversion and feature selection steps further extract and optimize key information, remove redundant features, and improve the learning efficiency of the model. Finally, data partitioning ensures the fairness and effectiveness of model evaluation, thereby enhancing the performance of the overall system in surgical stage identification.
[0113] Feature extraction module, which applies a dual-granularity temporal convolution model to extract key features from the preprocessed data;
[0114] The specific process of the feature extraction module extracting key features from the preprocessed data is as follows:
[0115] Input data preparation: The preprocessed data obtained from laparoscopic surgery is represented as a multi-dimensional time series: X = {x1, x2, …, x t};
[0116] where X is the input multi-dimensional time series dataset, containing t time steps, t is the total number of time steps, representing the total length of the data, and x t is the feature vector at time step t, denoted as x t ∈ R n , where n is the dimension of the feature;
[0117] Convolutional neural network design: Feature extraction is achieved through a convolutional neural network, which is used to extract important features from time series data;
[0118] One-dimensional convolutional kernels are used to process time series data. Suppose we define a convolutional kernel W:
[0119] W = {w1, w2, …, w k};
[0120] where W is the convolutional kernel, with size K, used to capture local patterns in the time series;
[0121] The result of the convolution operation is:
[0122] where Z t is the convolution output at time step t, representing the current feature response, w m is the m-th value in the convolutional kernel W, and b is the bias term, a scalar value used to adjust the convolution output;
[0123] A non-linear activation function is applied to the convolution output to introduce non-linearity:
[0124] Z t ' = f(Z t ) = max(0, Z t );
[0125] where Z t ' is the convolution result after being processed by the activation function, representing the non-linear form of the current feature response;
[0126] Dual-granularity convolution uses two different granularities to extract features, namely coarse-grained and fine-grained;
[0127] Coarse-grained convolution uses a larger convolutional kernel and stride to extract long-term features, defined as follows: the convolutional kernel is W1, with size K1 and stride S1;
[0128] The convolution formula for coarse-grained is as follows:
[0129] Among them, Z large [t] is the output of the coarse-grained feature at time step t, w 1,m is the m-th weight in the convolutional kernel W1, and b1 is the bias term of the coarse-grained convolution;
[0130] The fine-grained convolution applies a smaller convolutional kernel and uses a smaller stride to capture short-term features, defined as follows: the convolutional kernel is W2, the size is K2, and the stride is S2;
[0131] The convolutional formula for the fine-grained is as follows:
[0132] Among them, Z small [t] is the output of the fine-grained feature at time step t, w 2,m is the m-th weight in the convolutional kernel W2, and b2 is the bias term of the fine-grained convolution;
[0133] Feature fusion: Generate the fused feature through a concatenation operation: Z fused =[Z large , Z small ;
[0134] Among them, Z fused is the fused feature vector, containing features from the coarse-grained and fine-grained convolutions;
[0135] Fuse features through weighted averaging: Z fused =αZ large +(1 - α)Z small ;
[0136] Among them, α is the weight coefficient, 0 ≤ α ≤ 1, controlling the contribution degree of the two features;
[0137] Fully connected layer: After feature fusion, input these features into the fully connected layer to generate the final feature representation: F = σ(W f ·Z fused +b f );
[0138] Among them, F is the final feature vector, the output after being processed by the fully connected layer, W f is the weight matrix of the fully connected layer, each row corresponding to a class, b f is the bias vector of the fully connected layer, with the length the same as the output feature dimension, and σ is the activation function;
[0139] Feature selection: For the extracted features, apply feature selection methods to remove redundant or useless features;
[0140] Output feature: The finally output feature F will be passed to the subsequent stage recognition module through the supervision platform for the classification and recognition of the surgical stage;
[0141] The feature extraction module can efficiently capture important information in multi-scale time series data, thereby improving the performance and accuracy of the model. By automatically identifying deep features related to the surgical stage, this module reduces the need for manual intervention and feature engineering while retaining rich information in the data and enhancing the generalization ability of the model. This process not only optimizes subsequent recognition tasks but also provides a solid foundation for real-time monitoring and decision support in the surgical stage.
[0142] Example 2: The technical solution of this embodiment of the present invention is different from that of Example 1 in that:
[0143] As Figure 2 shown, the model training module trains the dual-granularity temporal convolutional model according to the surgical stage data containing true labels;
[0144] The specific steps for the model training module to train the dual-granularity temporal convolutional model are as follows:
[0145] Forward propagation: Input each batch of the training set into the dual-granularity temporal convolutional model and calculate the output layer by layer: Calculate the feature map, process the input data through convolution, adding bias, and activation functions, combine the coarse-grained and fine-grained features into a vector, perform a linear combination and activation on the feature vector, and predict the category;
[0146] Calculate the loss: Use the cross-entropy loss function to calculate the difference between the output of the model and the true label. The formula for cross-entropy is:
[0147] where y m is the true label, is the probability predicted by the model;
[0148] Backward propagation: Calculate the gradient of each layer according to the partial derivative of the loss function with respect to the model parameters through the chain rule, and use the optimization algorithm to update the weights and biases of each layer. The update formula is:
[0149] θ = θ - η·▽L;
[0150] where θ is the model parameter, η is the learning rate, and ▽L is the gradient of the loss function with respect to the parameter;
[0151] Dividing the training data into small batches, inputting one batch of data each time for forward propagation and backward propagation, and completing one traversal of the entire training set is called an epoch, and multiple epochs are required for training;
[0152] During the training process, the validation set is regularly input into the model for validation, and the following metrics are monitored: the loss of the model on the validation set, which helps to judge the learning effect of the model; the accuracy of classification on the validation set, which is used to evaluate the performance of the model. The validation set is used to monitor the performance of the model. If various performance metrics no longer improve or overfitting occurs, the training is terminated early and the current best model is saved.
[0153] The model training module can be combined with the feature extraction module, and such a combination has significant advantages. First, by training the dual-granularity temporal convolutional model based on the surgical stage data with true labels, the feature extraction module can efficiently extract key features highly relevant to the surgical stage. Such a combination not only ensures that the extracted features are highly representative and accurate but also enables the model to quickly adapt to different surgical scenarios in practical applications, improving the recognition performance. At the same time, the trained model can refine deeper information, optimize the recognition process, and ultimately provide strong support for the accurate monitoring and decision-making of the surgical stage.
[0154] The stage recognition module uses the trained dual-granularity temporal convolutional model to perform real-time recognition on new data, judge the current surgical stage, and output the corresponding recognition result.
[0155] The specific process of the stage recognition module for real-time recognition of new data is as follows:
[0156] Use sensors and cameras to collect video frames and vital feature data in real time, convert the video data into a fixed size, extract N frames from the video, and the timestamp of each frame is t k (k = 1, 2, …, N);
[0157] For recognition, a fixed number of consecutive frames are extracted as input data every T seconds, that is, the input is:
[0158] X = {I(t1), I(t2), …, I(t T )};
[0159] Among them, I(t k ) represents the image frame extracted at time t k ;
[0160] Extract the local features of each frame through the fine-grained convolutional layer. For the input image I(t k ), it can be expressed as: V f (t k ) = W f * I(t k ) + b f ;
[0161] Among them, W f is the convolutional kernel, b fis the bias term, * represents the convolution operation, V f (t k ) is the local feature map;
[0162] Extract the global features in the time series through the coarse-grained convolutional layer. The data features for the entire time period are:
[0163] Among them, W c is a larger convolutional kernel used to combine the time series features;
[0164] Combine the fine-grained and coarse-grained features into a feature representation:
[0165] F = f(V f , V c ), where f is a fusion function;
[0166] Classify the fused feature F through the fully connected layer:
[0167] Y = W y ·F + b y , where W y is the weight of the fully connected layer, and b y is the bias term;
[0168] For the output vector Y, calculate the probability for each stage:
[0169] Among them, P(y m ) is the probability of stage y m ;
[0170] Select the stage y * with the highest probability: y * = arg max P(y m );
[0171] Display the recognition result in real time on the monitoring interface in the operating room. The recognition result includes the current stage, the corresponding probability, and the timestamp;
[0172] Establish a feedback mechanism to record the doctor's evaluation of the accuracy of the recognition result, and continue to train the dual-grained time convolution model using the collected new data and feedback information to optimize the feature extraction and judgment accuracy;
[0173] Result evaluation module, used to systematically evaluate the results of laparoscopic surgery stage recognition;
[0174] The specific process of the result evaluation module systematically evaluating the results of laparoscopic surgery stage recognition is as follows:
[0175] Evaluate the performance of the model on the test set using the following metrics:
[0176] Accuracy: Calculate the proportion of correctly recognized stages in the total number of stages;
[0177] Recall: Measure the ability of the model to recognize specific stages, that is, the proportion of the sum of true positives and false negatives;
[0178] Precision: Indicate the proportion of actually correct predictions among those identified as a specific stage, that is, the proportion of the sum of true positives and false positives;
[0179] F1-score: Combine the harmonic mean of precision and recall to provide a single evaluation metric, which is suitable for cases of class imbalance;
[0180] Generate a confusion matrix to show in detail the recognition results of each stage, including the comparison between the true label and the predicted label, providing an intuitive view for model performance analysis;
[0181] Deeply analyze misclassified cases, identify the reasons for the model's recognition failure, feedback the evaluation results and analysis to the model training module, and display the evaluation results through visualization tools, including performance metric charts and confusion matrices;
[0182] Combining the stage recognition module and the result evaluation module can have the following advantages: real-time feedback and dynamic adjustment, improved recognition accuracy, and enhanced system reliability; through systematic evaluation, the recognition module can learn from the feedback and optimize its recognition strategy to improve adaptability specifically; in addition, the result evaluation provides multi-level verification of the recognition output to ensure the accuracy of decisions; combining the advantages of both can not only comprehensively evaluate the surgical process, but also accumulate data for subsequent analysis, providing strong support for surgeons' decisions and enhancing the safety and effectiveness of surgeries.
[0183] Integrated application module for integrating the recognition results with the clinical decision support unit and the surgical robot;
[0184] The specific process of the integrated application module for integrating the recognition results with the clinical decision support unit and the surgical robot is as follows:
[0185] Combine the recognition results with the patient's historical data, generate complete patient situation information through data fusion algorithms, and the clinical decision support unit generates implementation suggestions based on the latest recognition results and the patient's condition. The implementation suggestions include: automatically generating surgical operation steps, predicting potential complications and risks, and providing immediate coping strategies;
[0186] Convert the instructions from the decision support unit into specific control commands for the robot, and transmit the instructions to the surgical robot in real time through a low-latency network protocol. The surgical robot performs actions according to the received instructions;
[0187] Continuously monitor the working status of the surgical robot, the operation of medical equipment, and the vital sign data of the patient. If any abnormal situation is detected, the supervision platform quickly issues an alarm and provides emergency handling suggestions;
[0188] Record all data during the operation, including recognition results, decision support suggestions, and robot operation parameters, to form a complete surgical data file. After the operation, review and analyze the collected data to evaluate the surgical effect and the effectiveness of recognition and decision support;
[0189] Provide an intuitive graphical user interface to display surgical information, stage recognition results, and decision support suggestions, and provide multiple interaction methods for surgeons to quickly obtain information and perform operations;
[0190] Collect the usage feedback of surgeons on the system, and continuously optimize and upgrade the recognition algorithm, decision support system, and user interface based on the feedback information;
[0191] By integrating the application module to combine the recognition result with the clinical decision support unit and the surgical robot, the intelligence and safety of the surgical operation can be improved. This module can real-time feedback the recognition result to the clinical decision support system, provide accurate surgical stage analysis for surgeons, and help them make more effective decisions. In addition, the integration and cooperation with the surgical robot can achieve more efficient operation and precise task execution, reducing the risk of human error. This close integration not only optimizes the surgical process but also significantly enhances the overall effect and safety of the operation.
[0192] The working principle of the present invention: When in use, in the present invention, the dual-granularity temporal convolutional model is trained in real time by the model training module according to the existing real surgical data, which can accurately capture the multi-scale features during the operation, thereby significantly improving the recognition accuracy and precision of the model. And the training based on real data ensures the adaptability of the model in actual application, reduces the risk of overfitting, and can be optimized for specific surgical types to improve the effect and reliability of surgical stage recognition. And through the stage recognition module to identify the stage of the operation, the global features and local features in the time series are fused and classified during the recognition process, which can quickly judge the current surgical stage and output accurate results, improving the dynamic monitoring ability of the operation process. Through the result evaluation module to systematically evaluate the recognition result, it can independently verify the recognition result, improve the accuracy and reliability of the recognition. Through the comprehensive analysis of the output result, it can provide targeted feedback, promote the dynamic optimization and continuous learning of the recognition model, and the accumulated evaluation data not only helps to improve the model performance but also provides decision support for surgeons, improving the safety and effectiveness of the operation.
[0193] The above is only a preferred specific implementation manner of the present invention; however, the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, making equivalent substitutions or changes according to the technical solution and its improved concept of the present invention, shall be covered by the protection scope of the present invention.
Claims
1. Laparoscopic surgery stage recognition system based on dual-granularity temporal convolution, characterized by: It includes supervision platform, data collection module, data preprocessing module, feature extraction module, model training module, stage identification module, result evaluation module and integrated application module; A data acquisition module is used to collect real-time data during the laparoscopic surgery, including intra-abdominal pressure data, surgical instrument usage data, vital signs data, and video data; A data preprocessing module, used for preprocessing the collected surgical stage data; Feature extraction module, which applies a dual-granularity temporal convolution model to extract key features from the preprocessed data; Model training module, which trains the dual-granularity temporal convolution model based on surgical stage data with true labels; The stage recognition module uses the trained dual-granularity temporal convolution model to recognize new data in real time, determine the current surgical stage, and output the corresponding recognition results; an outcome assessment module for systematic evaluation of the outcomes identified during the laparoscopic surgery phase; Integrated application module for integrating recognition results with clinical decision support units and surgical robots.
2. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 1 is characterized in that: The specific steps of preprocessing the collected surgical stage data by the data preprocessing module are as follows: Data cleaning: Use filters to process signals, remove electromagnetic interference and background noise, use interpolation to fill missing values, and delete records with many missing values; Data standardization: Convert the data to a distribution with a mean of 0 and a standard deviation of 1 using the following formula: Where X is the numerical sequence of the signal, μ is the mean, and σ is the standard deviation; Data conversion: extract time features and encode classified data; Feature selection: Based on correlation analysis, select features that are closely related to the target variable and remove redundant features; Data partitioning: The data is divided into training set, validation set and test set with the allocation ratio of 70% training, 15% validation and 15% testing.
3. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 1 is characterized in that: The specific process of the feature extraction module extracting key features from the preprocessed data is as follows: Input data preparation: The preprocessed data obtained from laparoscopic surgery is represented as a multidimensional time series: X = {x1, x2, ..., x t }; Where X is the input multidimensional time series data set, which contains t time steps, t is the total number of time steps, indicating the total length of the data, and x t is the feature vector at time step t, denoted by x t ∈R n , where n is the dimension of the feature; Convolutional Neural Network Design: Feature extraction is achieved through convolutional neural networks to extract important features from time series data; Use a one-dimensional convolution kernel to process time series data. Suppose we define a convolution kernel W: In={in1,in2,…,in k }; Among them, W is the convolution kernel with size K, which is used to capture local patterns in the time series; The result of the convolution operation is: Among them, Z t is the convolution output at time step t, representing the current feature response, w m is the mth value in the convolution kernel W, b is the bias term, a scalar value used to adjust the convolution output; Apply a nonlinear activation function to the convolution output to introduce nonlinear characteristics: Z t '=f(Z t )=max(0,Z t ); Among them, Z t ' is the convolution result after being processed by the activation function, which represents the nonlinear form of the current feature response.
4. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 3 is characterized in that: Dual-granularity convolution uses two different granularities to extract features, namely coarse-grained and fine-grained; Coarse-grained convolution uses larger convolution kernels and strides to extract long-term features, which are defined as follows: the convolution kernel is W1, the size is K1, and the stride is S1; The coarse-grained convolution formula is as follows: Among them, Z large [t] is the output of the coarse-grained feature at time step t, w 1,m is the mth weight in the convolution kernel W1, and b1 is the bias term of the coarse-grained convolution; Fine-grained convolution applies a smaller convolution kernel and uses a smaller step size to capture short-term features, which is defined as follows: the convolution kernel is W2, the size is K2, and the step size is S2; The fine-grained convolution formula is as follows: Among them, Z small [t] is the output of the fine-grained feature at time step t, w 2,m is the mth weight in the convolution kernel W2, and b2 is the bias term of the fine-grained convolution; Feature fusion: Generate fusion features through connection operations: Z fused =[Z large ,Z small ]; Among them, Z fused is the fused feature vector, which contains features from coarse-grained and fine-grained convolutions; Fusion of features by weighted averaging: Z fused =αZ large +(1-α)Z small ; Among them, α is the weight coefficient, 0≤α≤1, which controls the contribution of the two features; Fully connected layer: After feature fusion, these features are input into the fully connected layer to generate the final feature representation: F = σ(W f ·Z fused +b f ); Among them, F is the final feature vector, the output after processing by the fully connected layer, W f is the weight matrix of the fully connected layer, each row corresponds to one class, b f is the bias vector of the fully connected layer, its length is the same as the output feature dimension, and σ is the activation function; Feature selection: For the extracted features, feature selection methods are applied to remove redundant or useless features; Output features: The final output feature F will be passed to the subsequent stage recognition module through the supervision platform.
5. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 1, characterized in that: The specific steps of the model training module to train the dual-granularity temporal convolution model are as follows: Forward propagation: Input each batch of the training set into the dual-granularity temporal convolutional model and calculate the output layer by layer: Calculate feature maps, process input data through convolution, bias and activation functions, combine coarse-grained and fine-grained features into a vector, perform linear combination and activation on the feature vector, and predict the category; Calculate the loss: Use the cross entropy loss function to calculate the difference between the model's output and the true label. The formula for cross entropy is: Among them, y m is the true label, The probability predicted by the model; Back propagation: According to the partial derivatives of the loss function with respect to the model parameters, the gradient of each layer is calculated by the chain rule, and the weights and biases of each layer are updated using the optimization algorithm. The update formula is: Among them, θ is the model parameter, η is the learning rate, is the gradient of the loss function with respect to the parameters.
6. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 5, characterized in that: Divide the training data into small batches, input a batch of data each time for forward propagation and backward propagation, and completing the traversal of the entire training set is called an epoch. Multiple epochs are required for training; During the training process, the validation set is regularly input into the model for verification, and the following indicators are monitored: the loss of the model on the validation set, the classification accuracy on the validation set, and the model performance is monitored using the validation set. If various performance indicators no longer improve or overfitting occurs, the training is terminated early and the current best model is saved.
7. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 1, characterized in that: The specific process of the stage recognition module for real-time recognition of new data is as follows: Use sensors and cameras to collect video frames and life feature data in real time, convert the video data into a fixed size, extract N frames from the video, and the timestamp of each frame is t k (k=1,2,…,N); For recognition, a fixed number of consecutive frames are extracted every T seconds as input data, that is, the input is: X={I(t1),I(t2),…,I(t T )}; Among them, I(t k ) indicates that at time t k The image frame extracted at The local features of each frame are extracted through the fine-grained convolution layer. k ), which can be expressed as: V f (t k )=W f *I(t k )+b f ; Among them, W f is the convolution kernel, b f is the bias term, * indicates the convolution operation, V f (t k ) is a local feature map; The global features in the time series are extracted through the coarse-grained convolution layer. The data features of the entire time period are: Among them, W c It is a larger convolution kernel used to combine time series features.
8. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 7, characterized in that: Combine fine-grained and coarse-grained features into one feature representation: F=f(V f ,V c ), where f is a fusion function; The fused feature F is classified through the fully connected layer: Y=W y ·F+b y , where W y is the weight of the fully connected layer, b y is the bias term; For the output vector Y, calculate the probability of each stage: Among them, P(y m ) is stage y m probability; Select the stage y with the highest probability * :y * = arg max P(y m ); The recognition results are displayed in real time on the monitoring interface of the operating room; A feedback mechanism is established to record doctors’ assessment of the accuracy of the recognition results, and the collected new data and feedback information are used to continue training the dual-granularity temporal convolution model.
9. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 8, characterized in that: The specific process of the result evaluation module systematically evaluating the results of laparoscopic surgery stage identification is as follows: Evaluate the performance of the model on the test set using the following metrics: Accuracy: calculate the proportion of correctly identified stages to the total stages; Recall: measures the ability of the model to identify a specific stage, that is, the ratio of true positives to false negatives; Precision: indicates the proportion of predictions identified as specific stages that are actually correct, that is, the ratio of the sum of true positives to false positives; F1-score: combines the harmonic mean of precision and recall to provide a single evaluation metric, which is suitable for cases with imbalanced categories. Generate a confusion matrix that shows the recognition results of each stage in detail, including the comparison between the true label and the predicted label, providing an intuitive view for model performance analysis; In-depth analysis of misclassified cases, identification of reasons for model recognition failure, and feedback of evaluation results and analysis to the model training module.
10. The laparoscopic surgery stage recognition system based on dual-granularity temporal convolution according to claim 9, characterized in that: The specific process of the integrated application module integrating the recognition results with the clinical decision support unit and the surgical robot is as follows: The recognition results are combined with the patient's historical data to generate complete patient status information through data fusion algorithms. The clinical decision support unit generates implementation recommendations based on the latest recognition results and patient conditions. The instructions from the decision support unit are converted into specific control commands for the robot, and the instructions are transmitted to the surgical robot in real time through a low-latency network protocol. The surgical robot performs actions according to the received instructions. Continuously monitor the working status of surgical robots, the operation of medical equipment, and the vital signs of patients; Record all data during the operation to form a complete surgical data file, and review and analyze the collected data after the operation; Provides an intuitive graphical user interface that displays surgical information, stage identification results, and decision support recommendations, and provides multiple interactive methods for surgeons to quickly obtain information and perform operations; Collect surgeons' feedback on the use of the system, and continuously optimize and upgrade the recognition algorithm, decision support system and user interface based on the feedback information.
Citation Information
Patent Citations
Laparoscopic surgery stage identification method and system based on dual-granularity time convolution
CN114372962A
Cited By
Endoscope operation evaluation system
CN121032747A