Multimodal Data Surgical Behavior Recognition Method Based on Multi-Level Supervision
Through the multimodal data surgical behavior recognition method based on multi-level supervision, using deep neural networks and multimodal data, the problem of insufficient accuracy and robustness of surgical behavior recognition in the prior art is solved, and more efficient and reliable surgical behavior recognition is achieved.
Patent Information
- Application Number
- CN202210448543.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-26
AI Technical Summary
The existing surgical behavior recognition methods lack a large amount of high-quality fine labeling data, resulting in poor recognition accuracy and robustness.
A multimodal data surgical behavior recognition method based on multi-level supervision is proposed. A multi-level surgical behavior category recognition model is constructed through deep neural networks, combining weak supervision and strong supervision annotation data, and using video and voice modal data for identification.
It improves the accuracy and robustness of surgical behavior recognition, saves the human resources of medical staff, achieves faster technological maturity, and provides reliable surgical behavior recognition support for clinical applications.
Smart Images

Figure CN114821784B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of intelligent medicine and video recognition, and particularly relates to a multi-modal data surgical behavior recognition method, system, and device based on multi-level supervision. Background Art
[0002] Surgical behavior recognition is an important development direction in the field of intelligent medicine. The quality of surgery is crucial to the life and health of patients, and there are significant differences in the quality of surgeries performed by doctors with different levels of experience. If an intelligent system can be used to recognize the surgeries of young doctors and provide corresponding guidance on the compliance of surgeries, it will be beneficial to accelerating the improvement of the surgical skills of young doctors. The key step among them is the recognition of surgical behaviors. The recording of surgical behaviors often involves data of multiple modalities, and how to make full use of the complementary relationship between data of each modality is an important guarantee for effectively recognizing behaviors. Among data of each modality, video data contains a huge amount of information. Currently, mainstream learning methods often require a large amount of high-quality and fine-grained annotations. However, the annotation of video data requires a very high labor cost, especially a large number of professional medical staff are needed for annotation, which brings many difficulties to video annotation. This patent proposes a surgical recognition method using multi-modal data under multi-level supervision for this demand and the problems it faces. Summary of the Invention
[0003] In order to solve the above problems in the prior art, that is, to solve the problem that the accuracy and robustness of existing surgical behavior recognition methods are poor due to the lack of a large amount of high-quality and fine-grained annotation data, in the first aspect of the present invention, a multi-modal data surgical behavior recognition method based on multi-level supervision is proposed. The method includes the following steps:
[0004] S100, obtaining a sequence of video frames to be recognized containing specific surgical behaviors and its corresponding speech sequence within a set time as input data;
[0005] S200, based on the input data, obtaining the category recognition results of each specific surgical behavior in the sequence of video frames to be recognized through a trained multi-level surgical behavior category recognition neural network model;
[0006] Wherein, the multi-level surgical behavior category recognition neural network model is constructed based on a deep neural network.
[0007] In some preferred embodiments, the training method of the multi-level surgical behavior category recognition neural network model is:
[0008] A100, obtaining training sample data and constructing a training set; the training sample data includes a sequence of video frames containing surgical behaviors, a speech sequence, and the ground truth labels of the category recognition results of each surgical behavior in the sequence of video frames;
[0009] A200. According to the training sample data, count the probability histograms corresponding to various types of surgical behaviors, and calculate the distribution entropy based on the probability histograms.
[0010] A300. Determine whether the distribution entropy of the surgical behavior of the current category is greater than the set lower threshold of the distribution entropy. If it is greater, jump to A500; otherwise, jump to A400.
[0011] A400. Select the two surgical behavior categories with the smallest corresponding probabilities in the probability histogram for merging. After merging, jump to step A300.
[0012] A500. After the merging is completed, input the surgical behaviors of each category into a pre-constructed multi-level surgical behavior category recognition neural network model to obtain the category recognition result of the surgical behavior as the prediction result.
[0013] Based on the prediction result and the true value label of the category recognition result of each surgical behavior, calculate the loss value through a pre-constructed loss function, and update the model parameters of the multi-level surgical behavior category recognition neural network model.
[0014] A600. Decompose each merged surgical behavior category obtained after the completion of A500, and use the decomposed surgical behaviors of each category as new training sample data, and repeat the execution of A200 - A500.
[0015] A700. Iteratively execute A600 until there are no new merged categories, and obtain a trained multi-level surgical behavior category recognition neural network model.
[0016] In some preferred embodiments, the method for decomposing each merged surgical behavior category obtained after the completion of A500 is: regarding each small surgical behavior category in the merged surgical behavior category as one category.
[0017] In some preferred embodiments, the pre-constructed loss function is constructed as follows:
[0018] Based on the prediction probabilities of various types of surgical behaviors output by the Softmax layer of the multi-level surgical behavior category recognition neural network model and the pre-annotated fine behavior labels, construct constraints on the discriminative accuracy of video fine annotation and the recognition accuracy of tools as the first loss function; the fine behavior label is the true value label of the surgical behavior category containing short actions.
[0019] By performing an infinity norm summation in the time domain on the recognition results of the behavior categories of each frame in the video frame sequence through weak supervision, a corresponding binary recognition vector is obtained, and combined with the video coarse-grained behavior classification label, a second loss function is constructed; the video coarse-grained behavior is a set of behaviors constructed from short actions included in each frame of a video frame sequence of a set duration.
[0020] Based on the foreground behavior features and background behavior features corresponding to each frame in the video frame sequence, a constraint function for the foreground norm length and a constraint function for the background norm length are respectively constructed as the third loss function and the fourth loss function.
[0021] The KL divergence is calculated between the predicted probabilities of various surgical behavior categories output by the Softmax layer of the multi-level surgical behavior category recognition neural network model and the distribution probabilities of the pre-obtained enhanced labels, and then a fifth loss function is constructed; the enhanced label is the pseudo-label of each short action in the video coarse-grained behavior obtained by optimizing through weak supervision constraints.
[0022] The prediction results corresponding to the video modality data and the speech modality data are obtained, and the cosine distance between the prediction results corresponding to the two modality data at the same time point is calculated, and then a sixth loss function is constructed; the video modality data is the video data of each frame in the video frame sequence; the speech modality data is the speech data of each frame in the speech sequence.
[0023] The surgical behavior recognition result is mapped to the surgical tool recognition result by linear mapping, and then a seventh loss function is constructed.
[0024] In some preferred embodiments, the pre-constructed loss function is:
[0025]
[0026]
[0027]
[0028]
[0029]
[0030]
[0031]
[0032]
[0033]
[0034]
[0035] Among them, represents the total loss, represents the first loss function, represents the second loss function, represents the sixth loss function, represents the seventh loss function, represents the fifth loss function, represents the fourth loss function, represents the third loss function, p t,c1 represents the predicted probabilities corresponding to various surgical behavior categories output by the Softmax layer of the multi-level surgical behavior category recognition neural network model. t represents the subscript of the traversed surgical time, c1 represents the set of surgical behavior categories, y t,c1 represents the fine behavior classification label, T represents the entire surgical behavior duration, Z t represents the binary recognition vector obtained by performing an infinity norm summation in the time domain on the results of recognizing the behavior categories of each frame of the video frame sequence through weak supervision, represents the predicted result corresponding to the video coarse-grained behavior, and the video coarse-grained behavior is a set of behaviors constructed from short actions included in each frame of a video frame sequence with a set duration, Y t represents the video coarse-grained behavior classification label, that is, the true value label of the video coarse-grained behavior category, V t represents the predicted result corresponding to the video modality data at time t, represents the transpose of the predicted result corresponding to the video modality data at time t, A t represents the predicted result corresponding to the speech modality data at time t, p t represents the predicted result corresponding to the surgical behavior category, represents the predicted probabilities corresponding to various surgical behavior categories at time t, represents the surgical tool recognition result, q t,c1 represents the distribution probability of the enhanced label. The enhanced label is the pseudo label of each short action in the video coarse-grained behavior obtained by optimizing through weak supervision constraints. s represents the subscript of the traversed background feature occurrence category, S represents the set of foreground and background feature categories, p t,s represents the distribution probability corresponding to the fine behavior classification label, S act represents the set of foreground feature categories, S bkg represents the set of background feature categories, i, j represent subscripts, m represents the maximum truncation value of the pre-set surgical behavior features, f n,i represents the predicted probability of the foreground feature, f n,j represents the predicted probability of the background feature.
[0036] In a second aspect of the present invention, a multi-modal data surgical behavior recognition system based on multi-level supervision is proposed. The system includes: a data acquisition module and a category recognition module;
[0037] The data acquisition module is configured to acquire a sequence of video frames to be recognized containing specific surgical behaviors and their corresponding voice sequences within a set time as input data;
[0038] The category recognition module is configured to obtain the category recognition results of each specific surgical behavior in the sequence of video frames to be recognized through a trained multi-level surgical behavior category recognition neural network model based on the input data;
[0039] Among them, the multi-level surgical behavior category recognition neural network model is constructed based on a deep neural network.
[0040] In a third aspect of the present invention, a device is proposed, including at least one processor; and a memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned multi-modal data surgical behavior recognition method based on multi-level supervision.
[0041] In a fourth aspect of the present invention, a computer-readable storage medium is proposed. The computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned multi-modal data surgical behavior recognition method based on multi-level supervision.
[0042] Advantages of the present invention:
[0043] The present invention improves the accuracy and robustness of surgical behavior recognition.
[0044] 1) The present invention comprehensively uses weakly supervised annotation and strongly supervised annotation data of surgical behaviors for surgical behavior recognition, thereby greatly saving the human resources of medical staff and accelerating the formation of a mature surgical behavior recognition technology for actual clinical applications. It can be used to recognize and decompose the surgical process videos of doctors, thereby helping to guide junior medical staff to master and improve surgical techniques more easily and quickly.
[0045] 2) The present invention makes full use of multi-modal data and obtains robust surgical behavior recognition results by utilizing the compatibility between multi-modal data;
[0046] 3) The present invention adopts a category balance method based on multi-level surgical behavior category merger to alleviate the problem of uneven distribution ratios of various behavior categories in surgical behaviors. Description of the Drawings
[0047] Other features, objects, and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments read in conjunction with the accompanying drawings.
[0048] Figure 1 is a schematic flowchart of a method for identifying surgical behaviors of multimodal data based on multi-level supervision according to an embodiment of the present invention;
[0049] Figure 2 is a schematic framework diagram of a system for identifying surgical behaviors of multimodal data based on multi-level supervision according to an embodiment of the present invention;
[0050] Figure 3 is a schematic flowchart of surgical behavior categories and a brief process of surgical behavior identification according to an embodiment of the present invention;
[0051] Figure 4 is a schematic framework diagram of a weak supervision module according to an embodiment of the present invention;
[0052] Figure 5 is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application according to an embodiment of the present invention. Detailed Embodiments
[0053] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] The present application will be further described in detail below in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. Additionally, it should be noted that, for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings.
[0055] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0056] A method for identifying surgical behaviors of multimodal data based on multi-level supervision according to the present invention, as Figure 1 shown, the method includes the following steps:
[0057] S100, obtaining a sequence of video frames to be recognized containing specific surgical behaviors and their corresponding speech sequences within a set time as input data;
[0058] S200. Based on the input data, obtain the classification results of each specific surgical behavior in the video frame sequence to be recognized through the trained multi-level surgical behavior classification neural network model;
[0059] Among them, the multi-level surgical behavior classification neural network model is constructed based on a deep neural network.
[0060] To more clearly illustrate the multi-modal data surgical behavior recognition method based on multi-level supervision of the present invention, the following will elaborate on each step in an embodiment of the method of the present invention in conjunction with the accompanying drawings.
[0061] In the following embodiments, first, the training process of the multi-level surgical behavior classification neural network model will be described, and then the process of obtaining the classification results of specific surgical behaviors by the multi-modal data surgical behavior recognition method based on multi-level supervision will be elaborated in detail.
[0062] 1. Training process of the multi-level surgical behavior classification neural network model
[0063] A100. Obtain training sample data and construct a training set; the training sample data includes a video frame sequence containing surgical behaviors, a voice sequence, and a ground truth label of the classification results of each surgical behavior in the video frame sequence;
[0064] In this embodiment, first obtain the training sample data of the model. The training sample data includes a video frame sequence containing surgical behaviors, a voice sequence (i.e., multi-modal data), and a ground truth label of the classification results of each surgical behavior in the video frame sequence, and construct a training set based on the training sample data.
[0065] A200. According to the training sample data, statistically calculate the probability histogram corresponding to each type of surgical behavior, and calculate the distribution entropy based on the probability histogram;
[0066] A300. Determine whether the distribution entropy of the current type of surgical behavior is greater than the set lower threshold of the distribution entropy. If it is greater, jump to A500; otherwise, jump to A400;
[0067] A400. Select the two surgical behavior categories with the smallest corresponding probabilities in the probability histogram for merging, and after merging, jump to step A300;
[0068] In this embodiment, due to the unbalanced distribution of various types of surgical behaviors, an iterative multi-level supervision method is used to handle the problem of class imbalance, transforming an unbalanced behavior recognition problem into multiple relatively balanced behavior classification recognition problems, improving the overall recognition effect. That is, set the lower threshold of the distribution entropy of the category (abbreviated as the lower limit of the distribution entropy): S low, based on the assumption that the statistical probability of the behavior frequencies of various categories in the training data and the test data is similar, the probability histogram of the occurrence of various categories in the training data is statistically calculated, and the distribution entropy of each category is calculated. According to the classification entropy, it is judged whether the distribution of the current classification is balanced enough. If it is not balanced, the method of merging (i.e., merging the categories of long-tail behaviors) the behaviors with the minimum occurrence probability is adopted to enhance the classification balance, as Figure 3 shown. The original surgical behavior recognition problem is decomposed into multi-level category recognition that meets the lower limit of the distribution entropy, and at each level, the behavior recognition is carried out according to the steps introduced in A500 respectively.
[0069] A500, after the merging is completed, the surgical behaviors of each category are input into the pre-constructed multi-level surgical behavior category recognition neural network model to obtain the category recognition result of the surgical behavior as the prediction result;
[0070] Based on the prediction result and the true value label of the category recognition result of each surgical behavior, the loss value is calculated through the pre-constructed loss function, and the model parameters of the multi-level surgical behavior category recognition neural network model are updated;
[0071] In this embodiment, if the distribution entropy of the categories after several rounds of category merging is greater than the distribution entropy lower limit threshold, according to the current category merging result, the behavior recognition is carried out on the labeled classes of the current merged training set, and the surgical behavior recognition of the first layer is completed to obtain the category recognition result of the surgical behavior, that is, the prediction result.
[0072] After the recognition is completed, based on the prediction result and the true value label of the category recognition result of each surgical behavior, the loss value is calculated through the pre-constructed loss function, and the model parameters of the multi-level surgical behavior category recognition neural network model are updated. The multi-level surgical behavior category recognition neural network model is constructed based on the deep neural network.
[0073] Among them, a key point of the present invention lies in designing a loss function for a behavior recognition model that reasonably reflects the constraints between modalities, multi-level annotation, and constraints between tasks. The loss function includes: strong supervision constraints: a fine behavior discrimination constraint that reflects the fine annotation constraints of basic short surgical actions such as pulling, sharp dissection, blunt dissection, coagulation, etc., and a medical device discrimination constraint that reflects the accuracy constraint of the instrument standard; weak supervision constraints: a category annotation discrimination error that reflects the accuracy of coarse-grained video category annotations (i.e., coarse-grained video classification labels) with clinical significance such as exploration, freeing the posterior mediastinal pleura, dividing the interlobar fissure, etc., a foreground behavior class spacing maximization constraint that increases the distribution distance of different types of behavior features, and a prior assumption that reflects the uniformity of the background behavior feature distribution; enhanced weak supervision constraints: a classification ratio annotation error that reflects the strengthening of weak supervision annotation; a speech / video embedding feature similarity constraint and a multi-modal classification consistency constraint that reflect the consistency between the speech modality signal and the video modality signal. The construction process of the pre-constructed loss function is specifically as follows:
[0074] A510, design and construct a constraint on the accuracy of video fine annotation discrimination and tool recognition. We use the logistic regression loss function constraint to obtain the predicted probability p corresponding to each type of surgical behavior successively through the LSTM module, fully connected layer, and Softmax of the model t,c1 and the relationship with a small amount of manually annotated fine behavior annotations (i.e., fine behavior classification labels: the true value labels of the surgical behavior categories containing short actions) y t,c1 ; the constructed constraint function (i.e., the loss function) is used as the first loss function, and the first loss function is:
[0075]
[0076] Among them, represents the first loss function, t represents the subscript of the traversed surgical time, and c1 represents the set of surgical behavior categories.
[0077] A520, design and construct a discriminant loss function using the coarse-grained video classification annotation in the weak supervision constraint. Through weak supervision (i.e., the backpropagation of the neural network), the infinity norm of the recognition results of each frame's behavior category is summed in the time domain to obtain a binary recognition vector for the long video segment: Z t , combined with the coarse-grained video classification label in the weak supervision constraint, the constructed constraint function (i.e., the loss function) is used as the second loss function, and the second loss function is:
[0078]
[0079] Among them, represents the second loss function, Represents the prediction result corresponding to the coarse-grained video behavior, where the coarse-grained video behavior is a set of behaviors constructed from short actions contained in each frame of a video frame sequence of a set duration, Y t Represents the classification label of the coarse-grained video behavior, that is, the true value label of the coarse-grained video behavior category. T represents the duration of the entire surgical behavior.
[0080] A530, introducing a loss function that reflects the difference between weakly supervised foreground behavior and background behavior. By introducing two prior criteria: First, the norm of the foreground behavior feature is larger, while the norm of the background behavior is smaller; Second, the distribution of the background behavior feature tends to be uniformly distributed. By introducing a loss function that characterizes these two constraints, the overall classification optimization result is constrained. The constraint function for the foreground norm (i.e., the third loss function) is:
[0081]
[0082] Where, Represents the third loss function, f n,i Represents the predicted probability of the foreground feature, f n,j Represents the predicted probability of the background feature, S act Represents the set of foreground feature categories, S bkg Represents the set of background feature categories, i, j represent subscripts, and m represents the maximum truncation value of the pre-set surgical behavior feature.
[0083] The constraint function for the background norm (i.e., the fourth loss function) is:
[0084]
[0085] Where, Represents the fourth loss function, s represents the subscript for traversing the categories of background feature occurrences, S represents the set of foreground and background feature categories, p t,s Represents the distribution probability corresponding to the fine-grained behavior classification label.
[0086] A540, for enhancing the weak supervision behavior recognition annotation in the weak supervision enhanced label constraint, the distribution probability q of the enhanced label (i.e., the enhanced label, and the enhanced label is the pseudo-label of each short action in the coarse-grained video behavior obtained by optimizing through weak supervision constraint) t,c1 (where the distribution probability is obtained by a mixture Gaussian function in the present invention) and the prediction probability p of each category of surgical behavior t,c1 Calculate the KL divergence. The constraint function (i.e., the fifth loss function) is shown as follows:
[0087]
[0088] Where, Represents the fifth loss function.
[0089] A550, using the constraint of discriminant result fusion of surgical multi-modal speech and video modal data, and using the recognition result vectors V of different modal data t and A t The cosine distance of is used as the loss function (i.e., the sixth loss function):
[0090]
[0091] where represents the sixth loss function, represents the transpose of the prediction result corresponding to the video modal data at time t, V t represents the prediction result corresponding to the video modal data at time t. The video modal data is the video data of each frame in the video frame sequence, and A t represents the prediction result corresponding to the speech modal data at time t. The speech modal data is the speech data of each frame in the speech sequence.
[0092] The framework diagram of the weakly supervised part module is as Figure 4 shown. The weakly supervised part module includes a prior standard constraint optimization module that utilizes foreground and background behavior, and a consistency constraint optimization module for the enhanced labels generated by weak supervision and the strong supervision labels. Specifically: First, extract the L3D features of the input data (multi-modal data including video frame sequences and speech sequences) through a deep neural network, then perform feature embedding processing on the L3D features. After processing, according to the time series, extract the foreground and background metrics through the L2 norm. After extraction, perform behavior segmentation according to the time series (including pseudo-foreground behavior segmentation and pseudo-background behavior segmentation). After segmentation, based on the foreground and background features, construct the third loss function (i.e., Figure 4 the uncertainty loss function in), and based on the distribution probability corresponding to the fine-grained behavior classification labels, construct the fourth loss function (i.e., Figure 4 the background distribution information entropy loss function in) to distinguish the foreground behavior features and background behavior features, and output the background distribution and the distribution of the target background.
[0093] In addition, based on the weakly supervised constraint of coarse-grained behavior classification, construct the second loss function (i.e., Figure 4 the coarse-grained classification loss function in), and based on the prior of the consistency of the distribution probability between the fine-grained behavior pseudo-labels generated from the coarse-grained behavior labels (i.e., video coarse-grained classification labels) and the fine-grained behaviors (i.e., short actions included in the video coarse-grained behaviors) annotation labels, construct the fifth loss function (i.e., Figure 4 the enhanced annotation loss function in) to generate fine-grained behavior pseudo-labels.
[0094] A560. Due to the strong correlation between surgical behaviors and instruments, in terms of multi-task learning, a linear mapping is used to map the results in the behavior recognition space (i.e., surgical behavior recognition results) to the tool recognition space. The loss function used (i.e., the seventh loss function) is as follows:
[0095]
[0096] Among them, represents the seventh loss function, p t represents the predicted result corresponding to the surgical behavior category, represents the predicted probability corresponding to each category of surgical behavior at time t, represents the surgical tool recognition result.
[0097] A570. By comprehensively using the constraints introduced in A510, A520, A530, A540, A550, and A560, a multi-constraint weakly supervised multi-modal surgical behavior multi-task recognition constraint function (i.e., the pre-constructed loss function) is constructed:
[0098]
[0099] Among them, represents the total loss.
[0100] A600. After the merger of A500 is completed, the merged surgical behavior categories are decomposed. The decomposed surgical behaviors of each category are used as new training sample data, and A200 - A500 are repeatedly executed;
[0101] In this embodiment, each small category of surgical behavior in the merged surgical behavior category is regarded as one category (i.e., restoring the merged category to the state before merger). For each decomposed merged category, A200 - A500 are repeated for downward-level category merger and decomposition.
[0102] That is, for the behavior recognition problem that has passed category merger, it is further decomposed into refined recognition sub-problems. On the basis of ensuring the balance of the recognition category distribution, fine-grained behavior recognition with the same granularity as the original unbalanced recognition problem can be achieved.
[0103] A700. A600 is iteratively executed until there are no new merged categories, and a trained multi-level surgical behavior category recognition neural network model is obtained.
[0104] In this embodiment, A600 is iteratively executed to classify and recognize surgical behaviors at each level until there are no new merged categories, and the training of the multi-level surgical behavior category recognition neural network model is completed.
[0105] 2. Multi-modal data surgical behavior recognition method based on multi-level supervision
[0106] S100. Obtain a sequence of video frames to be recognized containing specific surgical behaviors and its corresponding speech sequence within a set time as input data;
[0107] In this embodiment, a sequence of video frames to be recognized containing specific surgical behaviors and its corresponding speech sequence is obtained.
[0108] S200. Based on the input data, obtain the category recognition results of each specific surgical behavior in the sequence of video frames to be recognized through a trained multi-level surgical behavior category recognition neural network model.
[0109] In this embodiment, the obtained sequence of video frames to be recognized containing specific surgical behaviors and its corresponding speech sequence are input into a trained multi-level surgical behavior category recognition neural network model to obtain the category recognition results of each specific surgical behavior in the sequence of video frames to be recognized.
[0110] A multi-modal data surgical behavior recognition system based on multi-level supervision according to the second embodiment of the present invention, as Figure 2 shown, specifically includes the following modules: a data acquisition module 100 and a category recognition module 200;
[0111] The data acquisition module 100 is configured to obtain a sequence of video frames to be recognized containing specific surgical behaviors and its corresponding speech sequence within a set time as input data;
[0112] The category recognition module 200 is configured to, based on the input data, obtain the category recognition results of each specific surgical behavior in the sequence of video frames to be recognized through a trained multi-level surgical behavior category recognition neural network model;
[0113] Among them, the multi-level surgical behavior category recognition neural network model is constructed based on a deep neural network.
[0114] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process and related explanations of the above-described system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.
[0115] It should be noted that the multi-modal data surgical behavior recognition system based on multi-level supervision provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing each module or step, and are not regarded as improper limitations of the present invention.
[0116] A multi-modal data surgical behavior recognition device according to the third embodiment of the present invention includes a collection device and a central processing device;
[0117] The collection device includes a camera, a camera, and a voice collection device (such as a microphone), which are used to obtain a sequence of video frames to be recognized containing specific surgical behaviors and their corresponding voice sequences within a set time as input data;
[0118] The central processing device includes a GPU, which is configured to obtain the category recognition results of specific surgical behaviors in the sequence of video frames to be recognized based on the input data through a trained multi-level surgical behavior category recognition neural network model; wherein, the multi-level surgical behavior category recognition neural network model is constructed based on a deep neural network.
[0119] A device according to the fourth embodiment of the present invention includes at least one processor; and a memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned multi-modal data surgical behavior recognition method based on multi-level supervision.
[0120] A computer-readable storage medium according to the fifth embodiment of the present invention stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned multi-modal data surgical behavior recognition method based on multi-level supervision.
[0121] Those skilled in the art can clearly understand that for the sake of convenience and conciseness of description, the specific working processes and related descriptions of the above-described multi-modal data surgical behavior recognition device, device, and computer-readable storage medium can refer to the corresponding processes in the foregoing method examples, and will not be repeated here.
[0122] Next, refer to Figure 5 , which shows a schematic structural diagram of a computer system of a server suitable for implementing the method, system, and device embodiments of the present application. Figure 5The server shown is only an example and should not impose any limitations on the functions and scope of use of the embodiments of the present application.
[0123] As Figure 5 shown, the computer system includes a central processing unit (CPU, Central Processing Unit) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 502 or the program loaded from the storage section 508 into the random access memory (RAM, Random Access Memory) 503. In the RAM 503, various programs and data required for system operation are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The input / output (I / O, Input / Output) interface 505 is also connected to the bus 504.
[0124] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT, Cathode Ray Tube), a liquid crystal display (LCD, Liquid Crystal Display), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as required so that a computer program read from it can be installed into the storage section 508 as required.
[0125] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 509 and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU 501), the above-mentioned functions defined in the methods of the present application are performed. It should be noted that the computer-readable medium in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0126] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0128] The terms "first", "second", etc. are used to distinguish similar objects and not to describe or represent a specific order or sequence.
[0129] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to those process, method, article, or apparatus / device.
[0130] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A method for identifying surgical behaviors of multimodal data based on multi-level supervision, characterized in that, the method comprises the following steps: S100, obtaining a sequence of video frames to be recognized containing specific surgical behaviors and their corresponding speech sequences within a set time as input data; S200, based on the input data, obtaining the category recognition results of each specific surgical behavior in the sequence of video frames to be recognized through a trained multi-level surgical behavior category recognition neural network model; wherein, the multi-level surgical behavior category recognition neural network model is constructed based on a deep neural network, and its training method is: A100, obtaining training sample data and constructing a training set; A200, according to the training sample data, statistically calculating the probability histograms corresponding to various categories of surgical behaviors, and calculating the distribution entropy based on the probability histograms; A300, determining whether the distribution entropy of the surgical behavior of the current category is greater than a set lower threshold of the distribution entropy. If it is greater, jump to A500; otherwise, jump to A400; A400, selecting the two surgical behavior categories with the smallest corresponding probabilities in the probability histogram for merging, and after merging, jumping to step A300; A500, after the merging is completed, inputting the surgical behaviors of each category into a pre-constructed multi-level surgical behavior category recognition neural network model to obtain the category recognition results of the surgical behaviors as prediction results; A600, decomposing each merged surgical behavior category obtained after A500 is completed, and using the decomposed surgical behaviors of each category as new training sample data, and repeating A200 - A500; A700, iteratively executing A600 until there are no new merged categories, and obtaining a trained multi-level surgical behavior category recognition neural network model.
2. The method for identifying surgical behaviors of multimodal data based on multi-level supervision according to claim 1, characterized in that, the training sample data includes a sequence of video frames containing surgical behaviors, a speech sequence, and the ground truth labels of the category recognition results of each surgical behavior in the sequence of video frames; in step A500, based on the prediction results and the ground truth labels of the category recognition results of each surgical behavior, calculating a loss value through a pre-constructed loss function, and updating the model parameters of the multi-level surgical behavior category recognition neural network model.
3. The method for identifying surgical behaviors of multimodal data based on multi-level supervision according to claim 2, characterized in that, the method for decomposing each merged surgical behavior category obtained after A500 is completed is: regarding each small surgical behavior category in the merged surgical behavior category as one category.
4. The method for identifying surgical behaviors of multimodal data based on multi-level supervision according to claim 2, characterized in that, the method for constructing the pre-constructed loss function is: Based on the prediction probabilities of various categories of surgical behaviors output by the Softmax layer of the multi-level surgical behavior category recognition neural network model and the pre-annotated fine behavior labels, constructing constraints on the discriminative accuracy of video fine annotation and the recognition accuracy of tools as the first loss function; the fine behavior labels are the ground truth labels of the surgical behavior categories containing short actions. Sum the infinity norms of the recognition results of the behavior categories of each frame in the video frame sequence in the time domain through weak supervision to obtain the corresponding binary recognition vector, and combine the video coarse-grained behavior classification label to construct the second loss function; the video coarse-grained behavior is a set of behaviors constructed from short actions contained in each frame of a video frame sequence of a set duration. Based on the foreground behavior feature and background behavior feature corresponding to each frame in the video frame sequence, respectively construct a constraint function for the foreground norm length and a constraint function for the background norm length as the third loss function and the fourth loss function. Calculate the KL divergence between the predicted probabilities of various surgical behavior categories output by the Softmax layer of the multi-level surgical behavior category recognition neural network model and the distribution probabilities of the pre-obtained enhanced labels, and then construct the fifth loss function; the enhanced label is the pseudo label of each short action in the video coarse-grained behavior obtained by optimizing through weak supervision constraints. Obtain the prediction results corresponding to the video modality data and the speech modality data, and calculate the cosine distance between the prediction results corresponding to the two modality data at the same time point, and then construct the sixth loss function; the video modality data is the video data of each frame in the video frame sequence; the speech modality data is the speech data of each frame in the speech sequence. Use linear mapping to map the surgical behavior recognition result to the surgical tool recognition result, and then construct the seventh loss function.
5. The multi-modal data surgical behavior recognition method based on multi-level supervision according to claim 4, characterized in that the pre-constructed loss function is: Among them, represents the total loss, represents the first loss function, represents the second loss function, represents the sixth loss function, represents the seventh loss function, represents the fifth loss function, represents the fourth loss function, represents the third loss function, p t,c1 represents the predicted probability corresponding to each type of surgical behavior output by the Softmax layer of the multi-level surgical behavior category recognition neural network model. t represents the subscript of the traversed surgical time, c1 represents the set of surgical behavior categories, y t,c1 represents the fine behavior classification label, T represents the entire duration of the surgical behavior, Z t represents the binarized recognition vector obtained by performing an infinite norm summation in the time domain on the results of recognizing the behavior categories of each frame in the video frame sequence through weak supervision, represents the predicted result corresponding to the video coarse-grained behavior. The video coarse-grained behavior is a set of behaviors constructed from short actions contained in each frame of a video frame sequence with a set duration, Y t represents the video coarse-grained behavior classification label, that is, the true value label of the video coarse-grained behavior category, V t represents the predicted result corresponding to the video modality data at time t, represents the transpose of the predicted result corresponding to the video modality data at time t, A t represents the predicted result corresponding to the speech modality data at time t, p t represents the predicted result corresponding to the surgical behavior category, represents the predicted probability corresponding to each type of surgical behavior at time y, represents the surgical tool recognition result, q t,c1 represents the distribution probability of the enhanced label. The enhanced label is the pseudo-label of each short action in the video coarse-grained behavior obtained by optimizing through weak supervision constraints. s represents the subscript of the traversed background feature occurrence category, S represents the set of foreground and background feature categories, p t,s represents the distribution probability corresponding to the fine behavior classification label, S act represents the set of foreground feature categories, S bkg represents the set of background feature categories. i, j represent subscripts, m represents the maximum truncation value of the pre-set surgical behavior features, f n,i represents the predicted probability of the foreground feature, f n,j represents the predicted probability of the background feature.
6. A multi-modal data surgical behavior recognition system based on multi-level supervision, characterized in that the system includes: a data acquisition module and a category recognition module; the data acquisition module is configured to acquire a video frame sequence to be recognized containing a specific surgical behavior and its corresponding speech sequence within a set time as input data; the category recognition module is configured to obtain the category recognition result of each specific surgical behavior in the video frame sequence to be recognized through a trained multi-level surgical behavior category recognition neural network model based on the input data; wherein, the multi-level surgical behavior category recognition neural network model is constructed based on a deep neural network, and its training method is: A100. Obtain training sample data and construct a training set; A200. According to the training sample data, count the probability histograms corresponding to various surgical behavior categories, and calculate the distribution entropy based on the probability histograms; A300. Judge whether the distribution entropy of the surgical behavior of the current category is greater than the set distribution entropy lower threshold. If it is greater, jump to A500; otherwise, jump to A400; A400. Select two surgical behavior categories with the smallest corresponding probabilities in the probability histogram for merging, and after merging, jump to step A300; A500. After the merging is completed, input the surgical behaviors of each category into the pre-constructed multi-level surgical behavior category recognition neural network model to obtain the category recognition result of the surgical behavior as the prediction result. A600, decompose each merged surgical behavior category obtained after the merger of A500 is completed, and use the decomposed surgical behaviors of each category as new training sample data, and repeat the execution of A200 - A500; A700, iteratively execute A600 until there are no new merged categories, and obtain a trained multi-level surgical behavior category recognition neural network model.
7. A multi-modal data surgical behavior recognition device based on multi-level supervision, characterized in that, it includes: at least one processor; and a memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the multi-modal data surgical behavior recognition method according to any one of claims 1 - 5.
8. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the multi-modal data surgical behavior recognition method according to any one of claims 1 - 5.
Citation Information
Patent Citations
Abnormal behavior recognition method, device and apparatus based on voice and image features
CN111460889A