Layered discriminant method and system for key frame recognition in minimally invasive surgery
By extracting and analyzing features during laparoscopic surgery, a multi-layer discriminant model was established, which solved the problem of lack of quantitative assessment in laparoscopic manipulation and enabled the robot to make autonomous judgments and improve surgical efficiency.
Patent Information
- Application Number
- CN202210867796.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-07-22
AI Technical Summary
In current minimally invasive laparoscopic surgery, the operation of holding the scope relies on the experience of the assistant, lacks quantitative assessment indicators, resulting in low efficiency, long operation time, psychological and physiological burden on the assistant, and high cost of training excellent assistants.
By analyzing images, features in laparoscopic surgery, such as instrument type, location, observation distance, instrument tip condition, and energy instrument activation status, are extracted. A multi-layer discrimination model is established to autonomously determine key surgical tasks and optimize laparoscopic manipulation.
It enables the autonomous judgment and intermittent following of the instrument end-effector movement of the laparoscopic robot, optimizes the laparoscopic operation effect, provides optimized control of the laparoscopic autonomous laparoscopic robot and evaluation standards in the medical field, and improves surgical efficiency and safety.
Smart Images

Figure CN115634046B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of fastening equipment, and more particularly, to a hierarchical discrimination method and system for key frame recognition in minimally invasive surgery. BACKGROUND
[0002] Abdominal minimally invasive surgery has the advantages of small trauma to patients and short recovery time, and has been widely popularized in surgery. Abdominal minimally invasive surgery requires several small holes to be punched in the patient's abdomen, and a laparoscope and surgical instruments are inserted into the vicinity of the affected area from the small holes, and the surgeon operates outside the body to complete the surgery. During the operation, the holding assistant holds the laparoscope, and the main surgeon commands the holding assistant to move the laparoscope by means of oral instructions to provide the view that the main surgeon wants. However, there are several problems in the actual operation process: (1) the efficiency of oral instructions is low, and at the same time, the psychological burden of the main surgeon is increased; (2) the operation time is long, sometimes even up to 10 hours, which brings psychological and physiological burden to the holding assistant; (3) it takes a long time of actual operation experience to cultivate excellent holding assistants, and the training cost of the hospital is high. Therefore, many researchers have developed a robotic laparoscope autonomous mirror holding system to improve the efficiency of laparoscopic surgery. Among them, automatic skill assessment is of great significance in the training of laparoscopic autonomous mirror holding robots. A lot of research has been done in the evaluation of surgical skills, and various quantitative indicators have been proposed. However, in the field of surgery, there is no written quantitative indicator for laparoscopic mirror holding operation skills because of the variety of laparoscopic surgeries and the complexity of the scene, and the real mirror holding operation is very dependent on the experience of the holding assistant and the tacit understanding with the surgeon.
[0003] In the field of mirror holding robot development, a large number of studies have taken the following accuracy of the instrument tip as the only evaluation index. However, in real laparoscopic surgery, when performing a stage task, the holding assistant usually first moves the lens to a good viewing angle through the left process, and basically does not move again, that is, the laparoscope is not continuously moving all the time following the instrument tip, but is intermittently moving according to the key task in the surgery. Considering the dependence of mirror holding operation on experience, the European Cognitive-Guided Surgery team disassembles the surgical process task by task, allowing the doctor to intervene in the robot mirror holding operation for training, and evaluating the quality of the current picture (good / average / poor) by the doctor's monitoring, with the expert's evaluation representing the experience-dependent part. However, this evaluation scheme cannot evaluate a momentary view at a fixed time, and the mirror holding is carried out in the entire surgical process, and the model that guarantees the minimum prediction error is also the best model for the entire mirror holding operation. SUMMARY
[0004] In view of the above defects or improvement needs of the prior art, the present application provides a hierarchical discrimination method and system for key frame recognition in minimally invasive surgery, which extracts features in the image that can be applied to intraoperative evaluation, including the type, position, observation distance, instrument tip state and energy instrument excitation of the instruments in the image, and then sends the detected features into a multi-layer discrimination recognition system to complete the recognition of the key operation. This algorithm process enables the mirror holding robot to autonomously judge the key tasks in the surgical process, provides a triggering mechanism for the mirror holding robot to intermittently follow the instrument tip movement, optimizes the mirror holding operation effect, and determines the features and data set annotations for evaluation through task characteristic analysis of laparoscopic surgery.
[0005] In order to achieve the above-mentioned purpose, according to the first aspect of the present application, a hierarchical discrimination method for key frame recognition in minimally invasive surgery is provided, comprising the following steps:
[0006] S100: According to the task characteristics of laparoscopic surgery, the feature type and specific information used for evaluation are determined through image processing analysis, the key operation observation priority of the feature type is determined, and the detection process and key operation recognition process are optimized;
[0007] S200: Extract the feature type and specific information, establish the kinematic feature model and special feature model, and make corresponding data sets to develop model training to extract feature values for recognizing the key operation;
[0008] S300: Generate a discrimination model based on the feature values of the key operation and build a multi-layer recognition model, and send the extracted feature values into the multi-layer recognition model to obtain the key operation determination result.
[0009] Further, in step S100, the feature type includes explicit features and hidden features, the explicit features include one or more of instrument type, instrument tip position, instrument movement speed, instrument distance or instrument tip, and the hidden features are environmental features.
[0010] Further, in step S100, the observation priority of the key operation decreases in turn, and is respectively ligature clip, scissors, electric hook, right-angle separating forceps, separating forceps, bipolar electrocoagulation and intestinal forceps.
[0011] Further, in step S200, the establishment of the kinematic feature model includes:
[0012] S201: Given n images of all features x i , i = 1,...,n, the multivariate KDE of x is generally:
[0013]
[0014] where x∈Rd is the fitted probability density function, K(·) is a kernel function with a symmetric positive definite bandwidth matrix H ∈ R d×d , Gaussian is selected as the kernel function, i.e., K(·) ∝ Φ(·);
[0015] S202: Obtain the edge kernel density D of all features from the labeled key frames * :
[0016]
[0017] The subscript ∈ {p, v, d} corresponds to the features of the upper instrument, where p is the position feature, v is the velocity feature, and d is the distance feature to the laparoscope.
[0018] Further, in step S200, the establishment of the special feature model includes:
[0019] S203: Calculate the empirical smoke density threshold according to the average gray scale of the key frames in the thermal disconnection task as the special feature of the energy instrument;
[0020] S204: After video extraction, an image sequence with a length of T is obtained, and in the first step, the current model parameters are Based on the joint distribution of the conditional probability P(O, I | λ) The expected value of
[0021]
[0022] Where: λ = (A, B, Π), A, B, and Π are the sets of state transition probability distribution, observable value probability distribution, and initial state distribution, respectively; O = {o1, o2,..., o T} is the corresponding observation sequence, I = {i1, i2,..., i T} is the corresponding state sequence, where i T ∈ Q, o T ∈ V;
[0023] S205: Based on the continuous iteration of step S204, until converges, producing updated model parameters:
[0024]
[0025] According to the difference between the tip state transition between the key operation and other operations, two different random processes are modeled: the tip state and the result frame decision, as the specific features of the instrument.
[0026] Further, step S300 includes:
[0027] S301: Obtain six types of feature values of the instruments in the key frame, and generate instrument position kernel density model, instrument speed kernel density model, instrument distance kernel density model and instrument tip state hidden Markov model according to the instrument types;
[0028] S302: Each instrument generates a key operation model according to the feature values in the key operation, wherein the separation clamp feature information is processed into a separation clamp position density kernel density model, a separation clamp speed kernel density model, a separation clamp distance kernel density model and a separation clamp instrument tip state hidden Markov model;
[0029] S303: The above key operation models are sequentially connected to form a multi-layer recognition system. When the feature extraction of the main instrument in a frame is sent to the multi-layer recognition system of the instrument and simultaneously satisfies the four recognition models, it is determined that the frame is a key operation. The threshold values of each model are all one half of the minimum value of the data used to generate the key operation model sent to the model for prediction. The threshold values are x1, x2, x3,..., x n The data generates a kernel density model kde, and the threshold value of the model is:
[0030]
[0031] Further, step S300 includes:
[0032] S304: The environmental feature information, i.e., the smoke concentration, is judged. If the ratio of the smoke concentration to the highest concentration is greater than 0.6, it represents that the smoke concentration is high at this time and the energy instrument is in an excited state. At this time, the frame is a key operation;
[0033] S305: Otherwise, the multi-layer recognition system of the main instrument category is selected according to the obtained feature information, and then the instrument position, instrument speed, instrument distance and instrument tip state are sequentially sent to the corresponding recognition model. When the prediction results of each model to the input features are all greater than the set threshold value, it is determined that the frame at this time is a key operation. Otherwise, it is a non-key operation. Thus, the recognition process of whether the current frame is a key operation is obtained.
[0034] According to the second aspect of the present application, a hierarchical discrimination system for key frame recognition in minimally invasive surgery is provided, comprising:
[0035] A feature evaluation module is configured to determine the feature types and specific information used for evaluation, determine the key operation observation priority of the feature types, and optimize the detection process and key operation recognition process according to the task characteristics of the laparoscopic surgery through image processing analysis;
[0036] A feature extraction module is configured to extract the feature types and specific information, establish kinematic feature models and special feature models, make corresponding data sets, and develop model training to extract feature values used for recognizing the key operation.
[0037] The feature recognition module is configured to generate a discrimination model based on the feature values of the key operation and build a multi-layer recognition model, and send the extracted feature values into the multi-layer recognition model to obtain a key operation determination result.
[0038] According to a third aspect of the present application, a computer readable storage medium is provided, and a hierarchical discrimination method for key frame recognition in minimally invasive surgery is stored on the computer readable storage medium, and the hierarchical discrimination method for key frame recognition in minimally invasive surgery is configured to implement the steps of the hierarchical discrimination method for key frame recognition in minimally invasive surgery when executed by a processor.
[0039] According to a fourth aspect of the present application, a terminal device is provided, and the terminal device comprises a memory, a processor, and a hierarchical discrimination program for key frame recognition in minimally invasive surgery stored on the memory and executable on the processor, and the hierarchical discrimination program for key frame recognition in minimally invasive surgery is configured to implement the steps of the hierarchical discrimination method for key frame recognition in minimally invasive surgery.
[0040] Overall, compared with the prior art, the above technical solutions conceived by the present application can achieve the following beneficial effects:
[0041] 1. The method of the present application extracts features in the image that can be used for intraoperative evaluation, including the type, position, observation distance, instrument tip state, and energy instrument excitation of the instrument in the image, and then sends the detected features into a multi-layer discrimination recognition system to complete the recognition of the key operation, so that the mirror supporting robot can autonomously judge the key task in the surgical process, provide a trigger mechanism for the mirror supporting robot to follow the instrument tip movement intermittently, optimize the mirror supporting operation effect, determine the features and data set annotations for evaluation through the task characteristic analysis of the laparoscopic surgery, on the one hand, provide a target function for the optimization control of the laparoscopic autonomous mirror supporting robot, and on the other hand, provide a reference basis for formulating the mirror supporting evaluation standard in the medical field.
[0042] 2. The method of the present application introduces a feature extraction algorithm based on YOLO_v5, slices the whole video of the cholecystectomy surgery frame by frame, and equivalently obtains the types, positions, observation distances, instrument tip states, and other interpretable features of the instruments appearing in the intraoperative time sequence by calculating the position and size information of the Bounding Box. By extracting and analyzing the interpretable features in the laparoscopic surgery process, the interpretability of recognizing the key operation through these features is ensured.
[0043] 3. The method of the present application uses the smoke concentration in the laparoscopic image to indirectly represent the excitation state of the energy instrument in the surgical process, abstracts the problem of difficult detection of obvious features into a simple image classification problem, and thus efficiently obtains the key features for mirror supporting skill evaluation.
[0044] 4. The method of the present application, by quantifying the extracted features and performing statistical analysis to test the rationality of the extracted features, on the one hand, provides interpretability for laparoscopic optimization control, and on the other hand, the extracted distribution rule of surgical instruments lays a foundation for related research of autonomous surgery in hepatobiliary surgery. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 For the key operation recognition workflow of the laparoscopic surgery image of the embodiment of the present application;
[0046] Figure 2 For the method flow diagram of the six types of describable features acquired by the feature extraction system based on YOLO_v5 and Resnet of the embodiment of the present application;
[0047] Figure 3 For the process flow diagram of generating a discriminant model based on key operation feature values during a surgical procedure and building a multi-layer recognition model of the embodiment of the present application;
[0048] Figure 4 For the key operation recognition process diagram of the embodiment of the present application;
[0049] Figure 5 For the interpretable features during the use of a diathermy hook to disconnect a cystic catheter of the embodiment of the present application;
[0050] Figure 6 For the KDE model of the position and speed of the key frame with a ligature clip as the observation target of the embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.
[0052] As Figures 1-4 shown, the embodiment of the present application provides a hierarchical discriminant method for key frame recognition in minimally invasive surgery, comprising the following steps:
[0053] Step one: determine the features used for evaluation through analysis of the task characteristics of laparoscopic surgery
[0054] In the task analysis process, the present application traverses the whole video of six types of operations in hepatobiliary surgery (EndoGIA off-liver pedicle and hepatic vein, placement of hepatic portal blocking band, separation of perihilar ligament, separation of coronary ligament and liver splitting, intrathecal separation of liver pedicle, and resection of gall bladder), and it is found that the laparoscope basically does not move after finding a good view of the stage task in the process of holding the mirror. Existing research focuses on tracking the end of the instrument without analyzing the characteristics of the operation mirror, does not combine the experience of the doctor, introduces the interference factors caused by the complex operation in the operation, and does not have autonomy. After analyzing with hepatobiliary surgery experts, the characteristics used for evaluation are set as the describable characteristics of the type, position, observation distance, instrument tip state and energy instrument excitation of the instrument, and the indescribable characteristics that the doctor highly depends on experience during evaluation, i.e. hidden characteristics. Among them, the energy instrument in the excitation state has no obvious characteristics, but the high temperature burning of the tissue will produce a large amount of smoke, so the present application uses the smoke concentration in the picture as an environmental feature to indirectly reflect the excitation state of the energy instrument, abstracts it as an image classification problem, and processes the confidence of the image classification result as a smoke concentration feature. There are six types of describable characteristics, which are summarized in the following table:
[0055]
[0056] In addition, only one instrument usually performs key operations during the operation, which is referred to as the main instrument in the picture; and the other instruments in the picture only have clamping effect, and their feature information has no evaluation effect on the key operation. The seven instruments appearing in the operation are assigned different observation priorities according to the importance of evaluation, and only the feature information of the instrument with the highest priority in the picture is extracted, so as to optimize the detection process and key operation identification process. The priority judgment table of the feature to be extracted is as follows:
[0057]
[0058] Step two: feature extraction algorithm based on YOLO_v5 and ResNet (Residual Neural Network)
[0059] The kinematic characteristics of the key frame <position, velocity, distance> show probability distribution or random phenomena in the statistical results. Therefore, the non-parametric kernel density estimation (Kernel Density Estimation, KDE) method is used to model the internal probability structure of the empirical distribution. Given n images of all target features x i , i = 1,...,n, the multivariate KDE of x is generally defined as:
[0060]
[0061] x ∈ R d is the fitted probability density function (PDF). K(·) is a kernel function (or window) with a symmetric positive definite bandwidth matrix H ∈ R d×d Gaussian is chosen as the kernel function, i.e., K(·) ∝ Φ(·). Then, the edge kernel density D *
[0062]
[0063] The subscript ∈ {p, v, d} corresponds to the feature symbols in the above table, where p is the position feature, v is the velocity feature, and d is the distance to the laparoscope feature.
[0064] (2) Special feature modeling:
[0065] As a special feature of the energy instrument, D S is an empirical smoke density threshold calculated from the average gray level of the key frames in the thermal disconnection task, rather than a PDF. As a specific feature of the forceps, the tip state exhibits dynamic statistical properties with time dependence, rather than static statistical properties independent of time. Another mathematical tool is needed to discover this dynamic property and model the stochastic process according to the difference in tip state transitions between key operations and other operations. This operation involves two different stochastic processes: tip state and result frame decision (key frame or not). The tip state is a measurable stochastic process, while the frame decision is an unobservable stochastic process. This naturally leads the present invention to a Hidden Markov Model (HMM). HMM is a double stochastic process with an unobservable latent stochastic process, which can be obtained through another set of stochastic processes that generate observation sequences.
[0066] The observed state V = {0, 1, 2} is the number of tips of each instrument appearing in the laparoscope view. The hidden state Q = {True, False} is the property of the image, corresponding to the frame decision. The image sequence of length T obtained after video extraction, I = {i1, i2,..., iT} is the corresponding state sequence, O = {o1, o2,..., o T T} is the corresponding observation sequence, where i T ∈ Q, o T ∈ V. TV. It is impossible to label all key frames in a complete surgery video, so the hidden state of each image is unknown. To solve this problem, the present invention adopts the Baum Welch algorithm, which is a special implementation of the Expectation-Maximization (EM) algorithm in HMM learning. The HMM's λ = (A, B, Π), where A, B, Π are the sets of state transition probability distribution, observable value probability distribution, and initial state distribution, respectively. In the first step of expectation, the current model parameters are Joint distribution based on conditional probability P(O, I | λ) The expected value of
[0067]
[0068] Converge after continuous iterations, produce updated model parameters as follows
[0069] (3) Generate a discriminant model based on the feature values of key operations in the surgery process and build a multi-layer recognition model.
[0070] A professional doctor selects key operation pictures in multiple surgery videos, obtains the six types of feature values of the instruments in the key pictures through the feature extraction system in (2), and generates instrument position kernel density models, instrument speed kernel density models, instrument distance kernel density models, and instrument tip state hidden Markov models according to the instrument types. Each instrument generates a model according to the feature values in the key operation, such as the feature information of the separating forceps in the key operation, which is processed into a separating forceps position density kernel density model, a separating forceps speed kernel density model, a separating forceps distance kernel density model, and a separating forceps instrument tip state hidden Markov model. Connecting the above four models in turn is a multi-layer recognition system. When the feature extraction of the main instrument of a picture is sent into the multi-layer recognition system of this type of instrument and simultaneously satisfies the four recognition models, it is determined that this picture is a key operation. The recognition process of the multi-layer recognition system is as follows:
[0071] The threshold value of each model is half of the minimum value predicted by sending the data used to generate the model into the model, such as x1, x2, x3,..., x n The kernel density model kde generated by the data, and the threshold value of the model is
[0072]
[0073] (4) Send the feature information extracted in the surgery picture into the multi-layer recognition system built to obtain the discriminant result.
[0074]
[0075] First, the environmental feature information, i.e. smoke concentration, is judged. If the ratio of the smoke concentration to the highest concentration is greater than 0.6, it represents that the smoke concentration is high at this time and the energy device is in an excited state. At this time, the picture is the key operation. Otherwise, the multi-layer recognition system of the main device category is selected according to the feature information obtained by the feature extraction system. Then, the device position, device speed, device distance and device tip state are sent to the corresponding recognition model in turn. When the prediction results of each model to the input features are all greater than the set threshold, it is determined that the picture at this time is the key operation, otherwise it is the non-key operation. Thus, the recognition process of whether the current picture is the key operation is obtained.
[0076] The loss function of the target detection task is generally composed of two parts: Classificition Loss and Bounding Box Regeression Loss.
[0077]
[0078] where I is an indicator function, which judges whether the center of obj falls in the grid. When the center of obj falls in the grid, I = 1, otherwise I = 0, p(c) is the probability distribution of each category detected, Theoretical probability distribution of each category.
[0079] In YOLO_v5, Bounding Box Regeression Loss is CIOU_Loss which takes into account three important factors: overlapping area, center point distance and aspect ratio
[0080]
[0081] where IOU is the intersection over union of the Prediction box and the Ground truth box, Distance_2 2 is the Euclidean distance between the two centers of the Prediction box and the Ground truth box, Distance_C 2 is the diagonal distance of the minimum enclosing box of the Prediction box and the Ground truth box after taking the union. v is a parameter that measures the consistency of the aspect ratio, which can be defined as:
[0082]
[0083] where w gt , h gt are the width and height of the Ground truth box, w P , h PWidth and height of the Prediction box.
[0084] In summary, the interpretable features of keyframes cover the criteria of key operations by surgical experts, thus are expected to be detected. When multiple instruments in the scene perform suspicious key operations, the present application selects one priority observation target according to the priority column in the priority decision table to analyze its interpretable features, so as to improve the efficiency of feature extraction. Therefore, a set of interpretable features are analyzed based on the instrument category, which are divided into general instrument kinematic features <position, velocity, distance> and specific role attribute features <tip state, environmental features>. The feature description and corresponding extraction method are summarized as follows
[0085]
[0086]
[0087] As shown in Figure 5 and Figure 6 , in the embodiment of the present application, the experimental process and results are as follows:
[0088] (1) Modeling result analysis:
[0089] Extracting effective features is the main step of statistical modeling of interpretable features. In order to evaluate the effectiveness of the extracted features, the present application clips the video according to the start and end time of the subtask, extracts all the target features from these video segments, and compares these features with the action attributes. The correlation between the subtask and the feature is intuitively established based on the visualization of the extracted features. For example, Figure 5 shows the interpretable features during the disconnection of the cystic duct with a diathermy hook, where the red dots represent the keyframes at the moment when the diathermy hook is powered on (diathermy hook burn). The scatter plot shows that the diathermy hook is powered on multiple times during the entire subtask, indicating that there are keyframes in the thermal disconnection task. The keyframes of the energy instrument in the excited state have statistical regularity: the position, velocity and distribution are within the experience range, and the ratio of smoke concentration to the highest concentration in the keyframe with power-on state is greater than 0.6.
[0090] After intuitively confirming the reliability of the extracted features, a statistical model can be created to build the rule base of the subsystem. The present application classifies all keyframes according to the category of the observation target. The statistical model of the subsystem based on the keyframe criteria of the surgeon is trained according to the corresponding action. For example, the KDE model of the position and velocity of the keyframes with the ligature clip as the observation target is shown in Figure 6 .
[0091] For the tip state contained in the special feature, the present application trains HMM to detect keyframes in time series. Taking the ligature clip as an example, the final learned HMM of the tip state is shown in As learned, the tip state model parameters of the ligature clip are:
[0092]
[0093] Pi represents the sparsity of keyframes during subtasks. Pi represents the sparsity of keyframes during subtasks.
[0094] (2) Hierarchical discrimination results:
[0095] To evaluate the performance of the proposed keyframe detection system, the present application performs self-validation on the training set for rule base construction. The number of keyframes selected by the method of the present application is almost twice that of the original labeled results. Part of the results in the annotated video are shown in the following figure. The method of the present application can identify unlabeled keyframes according to the hierarchical discriminators of the explicit features extracted online. The detection results in the labeled video. The larger image at the top is the keyframe labeled by the surgeon, and the smaller image at the bottom is the unlabeled keyframe detected to perform the same key operation, and the detected keyframe is consistent with the predefined key operation.
[0096] The system of the embodiment of the present application is realized by relying on an electronic device, and therefore it is necessary to introduce the related electronic device. For this purpose, the embodiment of the present application provides an electronic device, which comprises at least one processor, a communications interface, at least one memory and a communications bus, wherein the at least one processor, the communications interface and the at least one memory complete the communication among each other through the communications bus. The at least one processor can call the logical instructions in the at least one memory to realize various systems provided in the system embodiment.
[0097] In addition, the logic instructions in the at least one memory described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the systems described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0098] The device embodiments described above are only illustrative, wherein the units illustrated as separate components can or can not be physically separated, and the components illustrated as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0099] From the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions essentially or the parts that contribute to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) realize the method or system described in each embodiment or some parts of the embodiment.
[0100] Those skilled in the art can easily understand that the above only describes the preferred embodiments of the present application and is not used to limit the present application, and any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a hierarchical discrimination program for key frame recognition in minimally invasive surgery, and the hierarchical discrimination program for key frame recognition in minimally invasive surgery is executed by a processor to implement a hierarchical discrimination method for key frame recognition in minimally invasive surgery. S100: According to the task characteristics of laparoscopic surgery, the feature types and specific information used for evaluation are determined through image processing analysis, the key operation observation priority of the feature types is determined, and the detection process and key operation recognition process are optimized; S200: Extract the feature types and specific information, establish the kinematic feature model and special feature model, and make corresponding data sets to develop model training to extract feature values for recognizing the key operation; S300: Generate a discrimination model based on the feature values of the key operation and build a multi-layer recognition model, and input the extracted feature values into the multi-layer recognition model to obtain a key operation determination result; In step S100, the feature types include explicit features and hidden features, the explicit features include one or more of instrument types, instrument tip positions, instrument moving speeds, instrument distances or instrument tips, and the hidden features are environmental features; In step S100, the observation priorities of the key operations are in descending order, and are respectively ligature clips, scissors, electric hooks, right-angle separating forceps, separating forceps, bipolar electrocoagulation and intestinal forceps; Step S300 includes: S301: Obtain six types of feature values of the instruments in the key frame, and generate instrument position kernel density models, instrument speed kernel density models, instrument distance kernel density models and instrument tip state hidden Markov models according to the instrument types; S302: Each instrument generates a key operation model according to the feature values in the key operation, wherein the separating forceps feature information is processed into a separating forceps position density kernel density model, a separating forceps speed kernel density model, a separating forceps distance kernel density model and a separating forceps instrument tip state hidden Markov model; S303: Connecting the above key operation models sequentially forms a multi-layer recognition system. When the features extracted from the main instrument in a certain frame are fed into the multi-layer recognition system of that instrument, and all four recognition models are satisfied simultaneously, this frame is determined to be a key operation. The threshold of each model is half of the minimum value predicted by the data used to generate the key operation model. The kernel density model kde generated from the data has the following threshold: (5)。 2. The computer-readable storage medium of claim 1, wherein, In step S200, the establishment of the kinematic feature model includes: S201: Given all features of images, in general a multivariate KDE is: (1) wherein , is a fitted probability density function, is a kernel function with a symmetric positive definite bandwidth matrix , a Gaussian is chosen as the kernel function, i.e. ; S202: Obtain edge kernel density of all features from the marked key frame : (2) Subscript Corresponding to the characteristic symbols of the above table, where p is the position feature, v is the velocity feature, and d is the distance feature to the laparoscope.
3. The computer-readable storage medium of claim 2, wherein, In step S200, the establishment of the special feature model includes: S203: According to the empirical smoke density threshold value calculated according to the average gray value of the key frame in the thermal disconnection task, the special feature of the energy instrument is obtained; S204: After video extraction, the image sequence with length is obtained, in the first step, the current model parameters are , and the expected value of the joint distribution based on the conditional probability is: (3) wherein: , , , are, respectively, a set of state transition probability distributions, a set of observable value probability distributions, and a set of initial state distributions; is a corresponding sequence of observations, is a corresponding sequence of states, wherein , ; S205: Based on the continuous iteration of step S204, until convergence, resulting in updated model parameters: (4) According to the difference between the tip state conversion of the key operation and other operations, a random process is modeled to obtain two different random processes: tip state and result frame decision, as the specific feature of the instrument.
4. The computer-readable storage medium of claim 1, wherein, Step S300 includes: S304: The environmental feature information, i.e. smoke concentration, is judged, and if the ratio of the smoke concentration to the highest concentration is greater than 0.6, it represents that the smoke concentration is high at this time and is in the excitation state of the energy instrument, and the frame at this time is the key operation; S305: Otherwise, select the multi-layer recognition system of the instrument category according to the obtained feature information, and then send the instrument position, instrument speed, instrument distance, and instrument tip state into the corresponding recognition model in turn. When the prediction results of each model for the input features are all greater than the set threshold, it is determined that the current frame is a key operation, otherwise it is a non-key operation. Thus, the recognition process of whether the current frame is a key operation is obtained.
5. A hierarchical discriminant system for key frame recognition in minimally invasive surgery, characterized in that, Comprise: a feature evaluation module for determining the feature type and specific information for evaluation, determining the key operation observation priority of the feature type, and optimizing the detection process and key operation recognition process through image processing analysis according to the task characteristics of laparoscopic surgery; a feature extraction module for extracting the feature type and specific information, establishing its kinematic feature model and special feature model, and making corresponding data sets to develop model training to extract feature values for recognizing the key operation; a feature recognition module for generating a discriminant model based on the feature values of the key operation and building a multi-layer recognition model, and sending the extracted feature values into the multi-layer recognition model to obtain the key operation determination result; In step S100, the feature type includes explicit features and hidden features, the explicit features include one or more of instrument type, instrument tip position, instrument movement speed, instrument distance, or instrument tip, and the hidden features are environmental features; In step S100, the observation priority of the key operation decreases in turn, and is respectively ligature clip, scissors, electric hook, right-angle separating forceps, separating forceps, bipolar electrocoagulation, and intestinal forceps; Step S300 comprises: S301: Obtain six types of feature values of the instruments in the key frame, and generate instrument position kernel density model, instrument speed kernel density model, instrument distance kernel density model, and instrument tip state hidden Markov model according to the instrument type; S302: Each instrument generates a key operation model according to the feature values in the key operation, wherein the separating forceps feature information is processed into separating forceps position density kernel density model, separating forceps speed kernel density model, separating forceps distance kernel density model, and separating forceps instrument tip state hidden Markov model; S303: Connecting the above key operation models sequentially forms a multi-layer recognition system. When the features extracted from the main instrument in a certain frame are fed into the multi-layer recognition system of that instrument, and all four recognition models are satisfied simultaneously, this frame is determined to be a key operation. The threshold of each model is half of the minimum value predicted by the data used to generate the key operation model. The kernel density model kde generated from the data has the following threshold: (5)。 6. A terminal device, characterized by comprising: The terminal device comprises a memory, a processor, and a layered discriminant program for key frame recognition in minimally invasive surgery stored on the memory and executable on the processor. The layered discriminant program for key frame recognition in minimally invasive surgery is configured to implement the layered discriminant method for key frame recognition in minimally invasive surgery according to any one of claims 1-4.
Citation Information
Patent Citations
Dynamic prediction method and system for use of surgical instrument
CN114005022A
Autonomous control method and system for visual field of endoscope and medium
CN114391793A