Method for automatically editing surgical video, computing device and surgical system
By using computing devices to perform machine learning analysis and multiple labeling of surgical videos, low-value segments are automatically cut out, solving the problems of time-consuming and inaccurate surgical video editing, achieving efficient and accurate surgical video editing, and improving video usage efficiency.
Patent Information
- Application Number
- CN202410501491.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-24
- Publication Date
- 2025-10-24
AI Technical Summary
Existing technologies for surgical video editing are time-consuming and inaccurate, resulting in a waste of time and resources for medical workers and limiting the widespread application of surgical videos in teaching, training, and research.
The system uses computing devices to perform machine learning analysis, identifies scene types in surgical video frames, tags them multiple times, automatically divides them into time segments, cuts out low-value segments, retains high-value segments, and generates edited surgical videos.
It enables efficient and accurate automatic editing of surgical videos, reducing manual time consumption, lowering storage pressure, improving video usage efficiency, and facilitating subsequent viewing and learning.
Smart Images

Figure CN120835180A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a method for automatically identifying technical assistance for video editing based on surgical videos. More particularly, the present application relates to a method performed by a computing device for automatically editing surgical videos, a computing device capable of performing the method, a computer-readable medium, and a surgical system comprising the computing device. BACKGROUND
[0002] Surgical operation processes are usually recorded by a camera device in a camera system, for example, a video is taken. In a minimally invasive surgery process, a surgical video is recorded by a laparoscope camera system, and the laparoscopic surgery video is an important record form of the surgery, which is necessary material for surgery review, training, quality control, scientific research, etc.
[0003] However, a complete original video file of a surgery can be as large as 10G or more, which puts pressure on storage. On the other hand, there are a series of specific time periods in a surgical operation process, which may not have much value for review, training, quality control, and scientific research, such as: time period of wiping in the body by a laparoscope, time period without surgical instruments appearing in the field of view, time period without surgical operation, and / or time period of fine-tuning the operation position of surgical instruments, etc. If these time periods are retained, it will make it inefficient to review a video, as it is impossible to determine which are these "specific time periods" in the original surgical video, and the observer has to wait patiently and watch all the videos.
[0004] Therefore, in order to efficiently use the minimally invasive surgery video, it is necessary to edit the original surgical video to streamline the video. At present, only manual methods are generally used to edit or cut / clip an original video to remove these low-value or valueless time periods. However, manual clipping is very time-consuming, and a typical surgical video of a surgery can be as long as several hours or even several dozen hours, and reviewing these original surgical videos or manually editing or clipping the video greatly occupies the valuable time of medical workers and limits the widespread use of surgical operation videos in teaching, training, communication, scientific research, etc. In addition, manual editing / clipping of the video may not be very accurate, and there is a risk of unstable level and many errors.
[0005] Therefore, it is desirable to provide an intelligent surgical video editing or clipping method, a computer-readable medium capable of performing the method, a computing device, and a surgical system, to automatically analyze, edit, clip, export, store, and / or display the original surgical video recorded by the camera device in an artificial intelligence manner, so as to realize convenient, efficient, and accurate automatic editing of surgical videos, reduce the time-consuming of manual work, and reduce the hardware requirements. SUMMARY
[0006] The present application aims to provide a method executed by a computing device for automatically editing a surgical video, a related computing device, a computer readable medium, and a surgical system to at least partially solve the above problems of the prior art, thereby allowing a user to obtain a surgical video with high efficiency and high accuracy of automatic editing / cutting with greatly reduced human input, to reduce storage pressure, save time, and facilitate subsequent viewing, editing, and learning of the user.
[0007] To at least achieve the above-mentioned purposes, according to one aspect of the present application, a method executed by a computing device for automatically editing a surgical video is provided, the method comprising:
[0008] receiving an original surgical video configured to be captured by a camera device associated with the computing device located in a surgical scene;
[0009] pre-processing the received original surgical video;
[0010] performing machine learning analysis on each video frame of a plurality of video frames of the pre-processed surgical video to identify whether the surgical scene of the video frame is a first type of surgical scene or a second type of surgical scene;
[0011] performing at least one labeling on the surgical video according to the identification result of each video frame of the plurality of video frames;
[0012] determining to retain or delete each time period of a plurality of time periods in the surgical video according to the result of the at least one labeling;
[0013] generating and exporting an edited surgical video based on the retained time periods;
[0014] storing the edited surgical video; and
[0015] displaying the edited surgical video on a display device associated with the computing device.
[0016] In some embodiments, the step of performing at least one labeling on the surgical video comprises:
[0017] grouping a plurality of video frames of a surgical video into a plurality of video frame groups in chronological order, wherein each video frame group comprises a predetermined number of video frames;
[0018] performing a first labeling on each video frame group respectively according to the identification result of each video frame of the plurality of video frames;
[0019] automatically dividing the surgical video into a plurality of time periods according to the result of the first labeling, wherein each of the plurality of time periods comprises one or more video frame groups;
[0020] respectively labeling each of the plurality of time periods; and
[0021] determining whether to keep or delete each of the plurality of time periods according to the result of the second labeling.
[0022] In some embodiments, the first labeling of each of the plurality of video frame groups is performed by a voting algorithm, and after the first labeling, the time periods are automatically divided and the second labeling is performed by a tolerance value algorithm according to the result of the first labeling obtained by the voting algorithm.
[0023] In some embodiments, the step of first labeling each of the plurality of video frame groups by a voting algorithm comprises:
[0024] labeling each of the plurality of video frame groups by a voting algorithm according to a ratio between a number of video frames in each of the plurality of video frame groups that are identified as the first type of surgical scene and a total number of video frames in the video frame group.
[0025] In some embodiments, the step of labeling each of the plurality of video frame groups by a voting algorithm comprises:
[0026] labeling the video frame group as a key frame group when the ratio between the number of video frames in the video frame group that are identified as the first type of surgical scene and the total number of video frames in the video frame group is not lower than a predetermined ratio threshold value;
[0027] labeling the video frame group as a non-key frame group when the ratio between the number of video frames in the video frame group that are identified as the first type of surgical scene and the total number of video frames in the video frame group is lower than the predetermined ratio threshold value.
[0028] In some embodiments, the step of automatically dividing the time periods and performing the second labeling by a tolerance value algorithm according to the result of the first labeling comprises:
[0029] determining a tolerance value of each of the plurality of video frame groups labeled by the voting algorithm according to whether the video frame group is a key frame group or a non-key frame group, in combination with considering a predetermined upper threshold value V max and a predetermined lower threshold value V min of the tolerance value; and
[0030] According to the determined tolerance value of each video frame group in the plurality of video frame groups, the surgical video is automatically divided into a plurality of time periods and each time period is marked for a second time.
[0031] In some embodiments, the step of marking each time period using a tolerance value algorithm includes:
[0032] From the initial time T of the first video frame group in each time period n Start, and cycle through the initial time T of each video frame group. n+i Determine whether the video frame group is a key frame group or a non-key frame group;
[0033] According to the judgment result, the tolerance value V of the video frame group is determined n+i :
[0034] When at the initial time T of the video frame group n+i When the video frame group is judged as a key frame group, the tolerance value V of the previous video frame group is n+i-1 and the sum V of the predetermined first value a n+i-1 +a and the upper limit threshold V of the predetermined tolerance value max Compare and determine the tolerance value of the video frame group based on the comparison result:
[0035] If V n+i-1 +a≤V max , let the tolerance value V of the video frame group be n+i =V n+i-1 +a;
[0036] If V n+i-1 +a>V max , let the tolerance value V of the video frame group be n+i-1 =V max ;
[0037] When at the initial time T of the video frame group n+i When the video frame group is judged as a non-key frame group, the tolerance value V of the previous video frame group is n+i-1 The value obtained by subtracting the predetermined second value b from the predetermined tolerance value lower limit threshold V min Compare and determine the tolerance value of the video frame group based on the comparison result:
[0038] If V n+i-1 -b>V min , let the tolerance value V of the video frame group be n+i =V n+i-1 -b;
[0039] If V n+i-1 -b≤V min , let the tolerance value V of the video frame group ben+i = V min ;
[0040] only if the tolerance value V n+i is determined to be V min , the loop is ended and the time period T n ~ T n+i is divided.
[0041] In some embodiments, the step of marking each time period by the tolerance value algorithm further comprises:
[0042] after the loop is ended and the time period T n ~ T n+i is divided, it is determined whether the time period contains one or more video frame groups with the tolerance value V min , so as to mark the time period for the second time:
[0043] if the time period contains one or more video frame groups with the tolerance value V min , it is marked as a low-value time period;
[0044] if the time period does not contain one or more video frame groups with the tolerance value V min , it is marked as a high-value time period.
[0045] In some embodiments, the method comprises a primary editing, and the primary editing comprises:
[0046] retaining the high-value time period and deleting the low-value time period; and
[0047] generating and exporting a primary-edited surgical video based on the retained high-value time period.
[0048] In some embodiments, the step of identifying whether the surgical scene of the video frame is a first type of surgical scene or a second type of surgical scene comprises target detection to identify whether the surgical scene is a scene in or out of a patient's body, whether a surgical instrument appears in the surgical scene, and the position and / or range of the surgical instrument, wherein,
[0049] the surgical scene is identified as the first type of surgical scene when the identified surgical scene meets the following requirements: the surgical scene itself is identified as being in the patient's body, while a surgical instrument appears in the surgical scene and the surgical instrument is in the patient's body; and
[0050] The surgical scene is identified as the second type of surgical scene when the identified surgical scene comprises one or more of the following: the surgical scene itself is identified as being outside of a patient, no surgical instrument is present in the surgical scene, and / or a surgical instrument present in the surgical scene is outside of a patient.
[0051] In some embodiments, when the surgical scene of a video frame is identified as the first type of surgical scene, the method further comprises:
[0052] identifying a type of surgical instrument present in the surgical scene; and
[0053] generating an instrument type label that labels corresponding video frames in the surgical video.
[0054] In some embodiments, the method further comprises generating one or more branch videos from the edited surgical video by type of surgical instrument, each branch video comprising an aggregation of video frames labeled by an instrument type label of the same type of surgical instrument.
[0055] In some embodiments, the step of at least one labeling of the surgical video further comprises a third labeling, wherein the third labeling comprises:
[0056] determining an operational state of a surgical instrument in the surgical scene by comparing each video frame in the plurality of video frames of the high value time period retained in the primary edit to its preceding video frame;
[0057] when the operational state of the surgical instrument is determined to be a firing state, generating a firing label that labels corresponding video frames in the high value time period retained in the primary edit; and
[0058] labeling the time period in the high value time period retained in the primary edit with the firing label as a high value active time period.
[0059] In some embodiments, the method further comprises a secondary edit, the secondary edit comprising:
[0060] retaining the high value active time period and deleting all other time periods in the high value time period retained in the primary edit other than the high value active time period; and
[0061] generating and exporting a secondary edited surgical video based on the retained high value active time period.
[0062] In some embodiments, the firing state of the surgical instrument is predetermined based on the type of the surgical instrument.
[0063] In some embodiments, when the surgical instrument is an ultrasonic scalpel, a stapler, and / or a hemostatic clip, the firing state is identified by the switching action between the open state and the closed state of the end effector of the surgical instrument; and when the surgical instrument is an electric scalpel, the firing state is identified by the smoke generation situation.
[0064] In some embodiments, the computing device is configured to allow a user to click on the firing label and / or instrument type label in the edited surgical video to jump to the location of the video frame in the surgical video corresponding to the firing label and / or instrument type label.
[0065] In some embodiments, the computing device is configured to allow a user to further edit the edited surgical video.
[0066] In some embodiments, the step of pre-processing the received raw surgical video includes:
[0067] Unifying the resolution of the original surgical video into a predetermined resolution; and
[0068] Perform frame extraction on the surgical video with unified resolution; and
[0069] The size of each of the multiple video frames of the surgical video after the frame decimation process is unified into a predetermined size.
[0070] According to another aspect of the present invention, a computing device is provided, wherein the computing device is configured to execute the aforementioned method.
[0071] According to another aspect of the present invention, a computer-readable medium stored in a computing device is provided, wherein the computer-readable medium is configured to execute the aforementioned method.
[0072] According to yet another aspect of the present invention, there is provided a surgical system, comprising:
[0073] According to the aforementioned computing device;
[0074] one or more surgical instruments associated with the computing device;
[0075] a camera device located in the surgical scene and associated with the computing device, the camera device being configured to capture raw surgical video; and
[0076] A display device associated with the computing device, the display device configured to display the edited surgical video.
[0077] By means of the method, the computing device, the computer readable medium and the surgical system according to the present application, the important segments in the surgical video can be automatically identified and marked by the artificial intelligence technology, the marked important segments can be reserved, the non-key segments can be removed, and the surgical video which can be initially edited or secondarily edited according to the selection of a user can be obtained, so that the surgical videos which are edited to different degrees are provided, the artificial time consumption and the storage pressure in the editing process of the surgical video are greatly reduced, and the use efficiency of the surgical video is effectively improved. BRIEF DESCRIPTION OF DRAWINGS
[0078] For better understanding of the above and other objects, features, advantages and functions of the present application, reference can be made to the preferred embodiments illustrated in the drawings. The same reference numerals in the drawings refer to the same components. It should be understood by those skilled in the art that the drawings are intended to illustrate the preferred embodiments of the present application schematically, and have no limiting effect on the scope of the present application, and the components in the drawings are not drawn to scale.
[0079] Figure 1 A method performed by a computing device for automatically editing a surgical video according to an aspect of the present application is shown;
[0080] Figure 2 Specific steps of pre-processing a received original surgical video in the method according to the present application are shown;
[0081] Figure 3 Specific steps of marking the surgical video at least once in the method according to the present application are shown;
[0082] Figure 4 A flowchart of the second marking in the method according to the present application is shown;
[0083] Figure 5 Specific steps of the third marking in the method according to the present application are shown. DETAILED DESCRIPTION
[0084] Now, referring to the drawings, the specific embodiments of the present application are described in detail. The preferred embodiments described herein are only based on the preferred embodiments of the present application, and other ways which can achieve the present application can be conceived by those skilled in the art based on the preferred embodiments, and the other ways also fall within the scope of the present application.
[0085] In surgical work, it is usually necessary to manually edit or clip surgical operations, such as the original surgical videos recorded by a shooting device set in the surgical scene in the operating room during the surgical operation, such as the laparoscope in the minimally invasive surgical system, to meet the purposes of subsequent viewing, review, teaching, improvement, etc. Since surgical operations are usually time-consuming, manual editing is also very time-consuming, resulting in waste of manpower, time, cost, etc. In addition, manual editing may not be very accurate, and there is a risk of unstable level and many erroneous operations. Therefore, the present invention provides an intelligent surgical video editing or clipping method, computing device and system to at least partially solve the above-mentioned problems in the prior art.
[0086] refer to Figure 1 One aspect of the present invention provides a method 10 executed by a computing device for automatically editing a surgical video, the method 10 comprising the following steps:
[0087] receiving 110 a raw surgical video configured to be captured by a camera device associated with a computing device located in a surgical scene;
[0088] Pre-processing 120 the received raw surgical video;
[0089] performing a machine learning algorithm analysis on each of the plurality of video frames of the pre-processed surgical video to identify 130 whether the surgical scene in the video frame is a first type of surgical scene or a second type of surgical scene;
[0090] marking the surgical video at least once according to the recognition result of each video frame in the plurality of video frames 140;
[0091] determining 150 whether to retain or delete each of a plurality of time periods in the surgical video based on a result of at least one marking;
[0092] generating and exporting 160 an edited surgical video based on the retained time period;
[0093] storing 170 the edited surgical video; and
[0094] The edited surgical video is displayed 180 on a display device associated with the computing device.
[0095] It can be appreciated that the above camera device can be any camera device that can record a video of a surgical procedure that can be conceived in a surgical application, operation, such as but not limited to a laparoscope camera device used in a minimally invasive surgical procedure in an operating room. The above computing device can be any computer, computing system associated with a corresponding surgical operation. For example, the computing device can be an auxiliary computing device for laparoscopic surgery, which can form an artificial intelligence (AI) model suitable for the present application, capable of performing the method of the present application, by retraining an existing open-source pre-trained model (e.g. YOLO) using training data obtained, for example, in previous surgical operations, through machine learning, thereby automatically editing the raw surgical video from the camera device.
[0096] Further preferably, the pre-processing 120 of the received raw surgical video can comprise Figure 2 the steps shown:
[0097] uniformizing 121 the resolution of the raw surgical video into a predetermined resolution; and
[0098] frame decimating 122 the surgical video after uniformizing the resolution; and
[0099] uniformizing 123 the size of each video frame of the surgical video after frame decimation to a predetermined size.
[0100] In the step of uniformizing the resolution, preferably, the resolution of the raw surgical video, regardless of its resolution size (e.g. 4k or 2k), is uniformized into a predetermined resolution, for example, all into 1080p. In this step, the size of the data source and the frame rate of the original video are preserved. In the frame decimation step, preferably, frame pictures are extracted for the surgical video after uniformizing the resolution, i.e. only part of the video frames of the surgical video are uniformly or non-uniformly kept. For example, preferably, assuming that the frame rate of the original surgical video is 60Hz, in the frame decimation step, only 9 video frame pictures are uniformly kept per second, so that the number of video frames required to be processed is greatly reduced (e.g. to 15% of the original video), reducing the storage and processing burden of the computing device, without loss of subsequent recognition accuracy. In the step of uniformizing the video frame size, preferably, the image of each video frame of the multiple video frames after frame decimation is embedded into a canvas of a specified size using the Letterbox method, so that the image is expanded to the network input size while maintaining the aspect ratio. This step of uniformizing the size facilitates the standardization of each frame image in the surgical video, so as to improve the accuracy of target detection in the subsequent analysis and recognition steps.
[0101] Further, in step 130 of the method 10 according to the present application, each of the plurality of video frames of the pre-processed surgical video is analyzed, e.g. target detection, by a machine-learned correlation algorithm model, e.g. a model obtained by retraining an existing open-source pre-trained model (e.g. YOLO) with training data obtained in previous surgical operations as described above, to identify whether the surgical scene shown in each video frame is a surgical scene of a first type or a surgical scene of a second type, e.g. to identify whether the surgical scene shown in the video frame being processed is a scene inside the patient or outside the patient, whether a surgical instrument appears in the surgical scene, and the position and / or extent of the surgical instrument. Illustratively, a surgical scene is identified as the surgical scene of the first type when the surgical scene itself is identified as being inside the patient, while a surgical instrument appears in the surgical scene and is identified as being inside the patient; and a surgical scene is identified as the surgical scene of the second type when the surgical scene itself is identified as being outside the patient, no surgical instrument appears in the surgical scene; and / or the surgical instrument appearing in the surgical scene is identified as being outside the patient.
[0102] Preferably, when the surgical scene of a video frame is identified as the surgical scene of the first type as described above, the type of the surgical instrument appearing in the surgical scene can be further identified, and an instrument type label is generated to mark the corresponding video frame in the surgical video, e.g. the final edited surgical video. Such a label allows subsequent processing to represent the surgical instruments of the corresponding type in the edited surgical video distinctively from other instruments, e.g. to generate different video branches to represent the surgical instruments of the corresponding type collectively, each branch representing all segments in which a surgical instrument of the same type appears in the entire surgical video (i.e. each branch video comprises an aggregation of the video frames marked by the instrument type label of the same type of surgical instrument). The instrument type label can further allow a user to click in the edited surgical video to jump to the corresponding video frame or time period.
[0103] Further, in step 140 of the method 10 according to the present application, the surgical video is at least once labeled according to the identification result of the target detection on each of the plurality of video frames of the pre-processed surgical video as described above. Preferably, with reference to Figure 3 the step of at least once labeling comprises the following specific steps:
[0104] According to the time sequence, the plurality of video frames of the surgical video processed by the above steps 110-130 are grouped 141 into a plurality of video frame groups, wherein each video frame group comprises a predetermined number of video frames;
[0105] According to the above identification result of each video frame in the plurality of video frames, each video frame group is respectively first labeled 142;
[0106] According to the result of the first labeling, the surgical video is automatically divided 143 into a plurality of time periods, wherein each time period in the plurality of time periods comprises one or more video frame groups;
[0107] Each time period in the plurality of time periods is respectively second labeled 144;
[0108] According to the result of the second labeling, it is determined 145 whether to keep or delete each time period in the plurality of time periods.
[0109] More preferably, in the above at least one labeling step, each video frame group in the plurality of video frame groups is respectively first labeled by a voting algorithm, and after the first labeling, according to the result of the first labeling obtained by the voting algorithm, a time period is automatically divided and second labeled by a tolerance value algorithm.
[0110] Preferably, in step 141, the plurality of video frames are grouped into a plurality of video frame groups, wherein the number of video frames in each group is predetermined. The number of video frames in each group can be the same as or different from other groups. The number of video frames can be predetermined according to one or more factors such as surgical operation experience data, actual situation, user specific needs, accuracy considerations, machine training parameters, etc. For example, in the present application, grouping can be performed in seconds, i.e. according to the time sequence, the plurality of video frames of the surgical video processed by the above steps 110-130 are grouped by seconds, and each video frame group comprises 9 video frames sorted in time sequence. It can be understood that the grouping method is not limited to the above-mentioned method of grouping 9 video frames per second, but other numbers of video frames can be grouped according to the needs in actual use.
[0111] Further, preferably, in step 142, after the grouping step 141, according to the above-mentioned identification result of each video frame of the plurality of video frames of the surgical video in step 130, each video frame group is sequentially marked for the first time 142. Preferably, for each video frame group in the plurality of video frame groups, the ratio between the number of video frames identified as the above-mentioned first type of surgical scene and the total number of video frames in the video frame group is calculated, and each video frame group in the plurality of video frame groups is sequentially marked by a voting algorithm. The voting algorithm is specifically that when the ratio between the number of video frames identified as the first type of surgical scene in each video frame group and the total number of video frames in the video frame group is not less than a predetermined proportion threshold, the video frame group is marked as a key frame group; when the above-mentioned ratio is lower than the predetermined proportion threshold, the video frame group is marked as a non-key frame group.
[0112] The above-mentioned predetermined proportion threshold, for example, depends on the data distribution in the algorithm training data and the degree of feature extraction, for example, can be determined in advance according to one or more factors such as surgical medical experience data, training parameters of machine learning model, actual application needs, etc. and stored in the computing device of the present application. In addition, the proportion threshold can also be fine-tuned by manual assistance. The proportion threshold can be predetermined according to the type of instrument, for example, different surgical instruments can correspond to different proportion thresholds, so as to correspondingly ensure the identification accuracy of each type of surgical instrument, reduce false judgments and omissions, while improving the identification efficiency. In one preferred embodiment, the proportion threshold can be 2 / 3, that is, as long as there are more than 6 video frames identified as the first type of surgical scene in each group of 9 video frames, the video frame group is marked as a key frame group, and if less than 6, it is marked as a non-key frame group (a non-key frame group means that there is valid information in the video frame group, for example, the proportion of surgical instruments located in the patient's body is low, so it is not key, which can be marked in this step, so as to be excluded in the subsequent processing process).
[0113] After the first marking step 142 by the voting algorithm, according to the marking result, the surgical video processed by the foregoing steps is automatically divided 143 to form a plurality of time periods, each time period formed by the division includes one or more video frame groups, and each time period is marked for the second time 144.
[0114] Preferably, the above-mentioned time period division and the related steps of the second marking introduce the concept of tolerance value, which uses the tolerance value algorithm to perform tolerance value judgment on multiple video frame groups in the surgical video in chronological order (the judgment result of the tolerance value algorithm indicates the distribution of key frame groups and non-key frame groups in each time period, which can provide a basis for the subsequent processing of retaining or deleting time periods), and divides the time period accordingly and marks each time period. Specifically, based on the result of the first marking, the steps of automatically dividing the time period and performing the second marking by the tolerance value algorithm include: according to whether each video frame group in the multiple video frame groups marked by the voting algorithm is a key frame group or a non-key frame group, combined with consideration of a predetermined upper limit threshold value V of the tolerance value. max and the predetermined lower limit threshold V of the tolerance value min , determining a tolerance value for the video frame group; and automatically dividing the surgical video into a plurality of time periods and marking each time period a second time according to the tolerance value for each of the determined plurality of video frame groups.
[0115] More specifically, refer to Figure 4 The flowchart shown below details the steps related to the above tolerance values:
[0116] From the initial time T of the first video frame group in each of the multiple time periods n Start, and cycle through the initial time T of each video frame group. n+i Determine whether the video frame group is a key frame group or a non-key frame group;
[0117] According to the judgment result, the tolerance value V of the video frame group is determined n+i :
[0118] When at the initial time T of the video frame group n+i When the video frame group is judged as a key frame group, the tolerance value V of the previous video frame group is n+i-1 and the sum of the predetermined first value a V n+i-1 +a and the upper limit threshold V of the predetermined tolerance value max Compare and determine the tolerance value of the video frame group based on the comparison result:
[0119] If V n+i-1 +a≤V max , let the tolerance value V of the video frame group be n+i =V n+i-1 +a;
[0120] If V n+i-1 +a>V max , let the tolerance value V of the video frame group be n+i-1 =V max ;
[0121] When the initial time T n+i , the tolerance value V n+i-1 of the previous video frame group is subtracted by a predetermined second value b, and the obtained value is compared with a predetermined lower threshold value V min of the tolerance value, based on the comparison result, the tolerance value of the video frame group is determined:
[0122] If V n+i-1 -b > V min , the tolerance value V n+i of the video frame group is set as V n+i-1 -b;
[0123] If V n+i-1 -b≤V min , the tolerance value V n+i of the video frame group is set as V min ;
[0124] Only when the tolerance value V n+i of the video frame group is determined as V min , the cycle of the time period T n ~T n+i is ended and the time period T n ~T n+i is divided, otherwise, the cycle of the time period is continued until the tolerance value V min is encountered.
[0125] It should be noted that in the above tolerance value determination process, T represents time, V represents tolerance value, and the subscript represents the i-th video frame group in the n-th time period in time sequence, for example, T n represents the start time of the first video frame group in the n-th time period, V n represents the tolerance value of the first video frame group in the n-th time period; T n+1 represents the start time of the second video frame group in the n-th time period, V n+1 represents the tolerance value of the second video frame group in the n-th time period…T n+i represents the start time of the i-1-th video frame group in the n-th time period, V n+i represents the tolerance value of the i-1-th video frame group in the n-th time period. Both n and i are integers. In addition, the above predetermined first value a is positive, and the predetermined second value b is negative.
[0126] After ending the cycle and dividing the time period T n ~T n+i , it is determined whether the tolerance value V minOne or more video frame groups to mark the time period for the second time: if the time period contains a tolerance value of V min One or more video frame groups are marked as low-value time periods; if the time period does not contain a tolerance value of V min One or more video frame groups are marked as high-value time periods.
[0127] It is understood that the steps of dividing the time periods and marking the second time can be performed in parallel or alternately in real time, or the second marking can be performed after the entire video is divided into time periods. The algorithm can be adjusted according to actual needs.
[0128] After the second marking, the surgical video is initially edited, that is, the high-value time period is retained and the low-value time period is deleted, and based on the retained high-value time period, a surgical video that has undergone initial editing is generated and exported.
[0129] Next, combine Figure 5 , the steps of the second marking as described above are illustrated. In this example, only the time period within 8 seconds after the start of the video is exemplified to help understand the above tolerance value algorithm. Assuming that starting from time 0, in seconds (that is, according to a number of video frames per second as a group, for example, the above 9 video frames per second as a video frame group), the tolerance value is judged in chronological order. Among them, each video frame group has been marked as a key frame group or a non-key frame group after the first marking. Assume that the first marking results of the 8 video frame groups during 0-8 seconds are shown in the following table. Assume that the initial tolerance value is 0. The above predetermined first value a is equal to 1, and the predetermined second value b is equal to -1, that is, when encountering a key frame group, the tolerance value is refreshed based on the current tolerance value, that is, the value is increased by 1, and when encountering a non-key frame group, the tolerance value is consumed, that is, the value is reduced by 1. In addition, in this example, it is assumed that the upper limit threshold value V of the above tolerance value max The lower limit of the tolerance value is 0.
[0130] Table 1 The first labeling results of each video frame group
[0131]
[0132] As shown in Table 1, during 0-1 seconds, the first time marking result of the first video frame group is judged as "key frame group", thus, the initial tolerance value 0 is increased by 1, i.e. the tolerance value is refreshed as 1, then, the first time marking result of 1-2 seconds is also judged as "key frame group", thus, the tolerance value is refreshed again, and the tolerance value becomes 2. During 2-3 seconds, the first time marking result of the third video frame group is still "key frame group", but at this time, the sum of the current tolerance value 2 and the first predetermined value 1 is greater than the upper limit threshold value 2 of the predetermined tolerance value, thus, the tolerance value is not refreshed, but remains as 2. Then, during 3-4 seconds, the first time marking result of the fourth video frame group is "non-key frame group", thus, the tolerance value is consumed, i.e. decreased by 1, and the tolerance value is reduced to 1. During 4-5 seconds, the first time marking result of the fifth video frame group is still "non-key frame group", thus, the tolerance value is continuously consumed, and the tolerance value is reduced to 0. At this time, since the current tolerance value is not greater than the lower limit threshold value V min Thus, the loop is ended and a time period is divided, i.e. a time period containing 0-4 seconds is formed.
[0133] After the loop is ended and the time period is divided, the tolerance value is continuously judged according to the time advancement, starting from the sixth video frame group of 5-6 seconds. Since the first time marking result of the sixth video frame group is still "non-key frame group", the tolerance value is still not greater than 0, the loop is ended and a time period is divided, i.e. a time period containing 4-5 seconds is formed.
[0134] The tolerance value is continuously judged according to the time advancement, starting from the seventh video frame group of 6-7 seconds. Since the first time marking result of the seventh video frame group is "key frame group", the tolerance value is refreshed as 1, and the loop is continuously executed. Since the first time marking result of the eighth video frame group of 7-8 seconds is still "key frame group", the tolerance value becomes 2, and the loop of the current time period is continuously executed, until the tolerance value is consumed as 0, the loop is ended and the time period is divided.
[0135] After the above steps, the division forms a plurality of time periods as described above, for example, a time period containing 0-4 seconds and a time period containing 4-5 seconds. Then, it is judged whether each time period contains a video frame group with a tolerance value of 0 to perform the second labeling. For example, as described above, the time period of 0-4 seconds is judged to not contain a video frame group with a tolerance value of 0, and the tolerance value of each video frame group corresponding thereto is 1, 2, 2, and 1 in time sequence, so that the second labeling result of the time period is "high-value time period"; the time period of 4-5 seconds contains only one video frame group, and the corresponding tolerance value is 0, so that the second labeling result of the time period is "low-value time period". The steps of dividing the time period and the second labeling can be performed in parallel and alternately in real time, or after the entire video is divided into time periods, the second labeling is performed. The algorithm can be adjusted according to actual needs. For example, the second labeling can be performed immediately after each time period is divided, or the second labeling can be uniformly performed after the entire video is divided into time periods.
[0136] After the second labeling, the surgical video is initially edited, that is, the high-value time period is retained, and the low-value time period is deleted, and a preliminarily edited surgical video is generated and exported based on the retained high-value time period.
[0137] After the above steps, a preliminarily edited surgical video containing only the high-value time period is obtained, which eliminates the low-value and unnecessary time period in the original video by means of artificial intelligence, for example, a variety of user-unwanted segments such as no surgical instrument in the field of view, the endoscope is located outside the patient's body, the surgical instrument is located outside the body, etc. Therefore, the user does not need to perform manual editing to obtain an automatically edited surgical video with greatly shortened length, and the edited surgical video only retains useful high-value segments, which is convenient for the user to review, consult, learn, and greatly reduces the memory pressure.
[0138] Further, the method according to the present application can also perform more refined processing on the surgical video, for example, on the basis of the surgical video processed in the above steps, a third labeling 146 is performed. The third labeling includes Figure 5 the following steps:
[0139] By comparing each video frame in the plurality of video frames of the high-value time period retained in the preliminary editing with the previous video frame, the operation state of the surgical instrument in the surgical scene is determined 1461.
[0140] When the operation state of the surgical instrument is determined to be the firing state, a firing label is generated 1462, which labels the corresponding video frame in the high-value time period retained in the preliminary editing; and
[0141] The time period with the firing label in the high-value time period reserved in the primary editing is marked 1462 as a high-value valid time period.
[0142] Wherein, the firing state of the surgical instrument can be determined in advance according to different types of surgical instruments. As a non-limiting example, when the surgical instrument is an ultrasonic knife, a stapler, and / or a hemostatic clamp, the firing state is identified by a switching action between an open state and a closed state of an end effector of the surgical instrument (for example, whether the video frame is in the firing state can be identified by the opening and closing of the end effector of the surgical instrument in the adjacent front and rear two video frame pictures); and when the surgical instrument is an electric knife, the firing state is identified by the smoke generation condition (for example, whether the video frame is in the firing state can be identified by the smoke change condition in the adjacent front and rear two video frame pictures).
[0143] Similar to the instrument type label described above, when the user uses and watches the surgical video of the primary editing, the firing label can be clicked to directly jump to the position of the video frame or time period corresponding to the label.
[0144] It should be noted that the above three steps can be performed in sequence, alternately, or in parallel, for example, the third marking step can be performed in parallel or alternately with the second marking step, or can be performed after the completion of the second marking step of the entire video and the primary editing step, and can be selected according to actual needs.
[0145] After the third marking step, secondary editing is performed on the basis of the primary editing. The secondary editing includes: reserving the high-value valid time period, and deleting other time periods in the high-value time period reserved in the primary editing except for the high-value valid time period; and generating and exporting the secondary edited surgical video based on the reserved high-value valid time period. The secondary edited surgical video is a more refined edited surgical video, which only retains the corresponding time period of the firing state of the surgical instrument, thereby further shortening the video time and facilitating the user to view the most interesting and most important firing operation. The method of the present application allows the user to choose to export the primary edited surgical video containing the surgical instrument related time period, or the secondary edited surgical video which is more refined, shorter in time, and only contains the firing time period of the surgical instrument, thereby being more flexible.
[0146] Further preferably, the computing device capable of performing the method of the present application is further configured to allow the user to further manually edit the surgical video of each stage, in particular the edited surgical video (primary or secondary editing). In this way, the user is provided with higher freedom to manually mark, fine-tune, edit, etc. according to his needs, to generate a more personalized surgical video.
[0147] According to another aspect of the present application, there is provided a computing device configured to enable the aforementioned method provided by the present application. The computing device can be any possible device within a surgical scene associated with a surgical operation having video receiving, processing, deriving, saving, forwarding, editing, etc. functionalities, e.g. a computer associated with a surgical device, a scope, etc.
[0148] According to yet another aspect of the present application, there is provided a computer readable medium stored in a computing device configured to enable the aforementioned method provided by the present application.
[0149] According to still another aspect of the present application, there is provided a surgical system comprising: the aforementioned computing device provided by the present application; one or more surgical instruments associated with the computing device; a camera device located in a surgical scene and associated with the computing device, the camera device being configured to capture raw surgical video; and a display device associated with the computing device, the display device being configured to display the edited surgical video.
[0150] The above description of various embodiments of the present application is provided for the purpose of describing the relevant teachings of the present application to one of ordinary skill in the art. The present application is not intended to be exclusive or limited to a single disclosed embodiment. As above, one of ordinary skill in the art will appreciate various alternatives and modifications to the present application. Therefore, although some alternative embodiments have been described in particular detail, other embodiments and modifications will be readily apparent to those of ordinary skill in the art. The present application is intended to include all such alternatives, modifications and variations as falling within the scope of the present application as described above.
Claims
1. A method performed by a computing device for automatically editing surgical video, the method comprising: The method comprises: receiving a raw surgical video configured to be captured by a camera associated with the computing device and located in a surgical scene; pre-processing the received raw surgical video; performing machine learning analysis on each video frame of a plurality of video frames of the pre-processed surgical video to identify whether the surgical scene of the video frame is a first type of surgical scene or a second type of surgical scene; performing at least one labeling on the surgical video according to the identification result of each video frame of the plurality of video frames; determining whether to retain or delete each time period of a plurality of time periods in the surgical video according to a result of the at least one labeling; generating and exporting an edited surgical video based on the retained time period; storing the edited surgical video; and displaying the edited surgical video on a display device associated with the computing device.
2. The method of claim 1, wherein, The step of performing at least one labeling on the surgical video comprises: grouping a plurality of video frames of a surgical video into a plurality of video frame groups in a time sequence, wherein each video frame group comprises a predetermined number of video frames; performing a first labeling on each video frame group respectively according to the identification result of each video frame of the plurality of video frames; automatically dividing the surgical video into a plurality of time periods according to a result of the first labeling, wherein each time period of the plurality of time periods comprises one or more video frame groups; performing a second labeling on each time period of the plurality of time periods respectively; and determining whether to retain or delete each time period of the plurality of time periods according to a result of the second labeling.
3. The method of claim 2, wherein, The first labeling on each video frame group of the plurality of video frame groups is performed by a voting algorithm, and after the first labeling, the time periods are automatically divided and the second labeling is performed by a tolerance value algorithm according to a result of the first labeling obtained by the voting algorithm.
4. The method of claim 3, wherein, The step of performing the first labeling on each video frame group of the plurality of video frame groups by a voting algorithm comprises: labeling each video frame group by a voting algorithm according to a proportion between a number of video frames identified as the first type of surgical scene in each video frame group of the plurality of video frame groups and a total number of video frames in the video frame group.
5. The method of claim 4, wherein, The step of labeling each video frame group by a voting algorithm comprises: labeling the video frame group as a key frame group when the proportion between the number of video frames identified as the first type of surgical scene in each video frame group and the total number of video frames in the video frame group is not lower than a predetermined proportion threshold; labeling the video frame group as a non-key frame group when the proportion between the number of video frames identified as the first type of surgical scene in each video frame group and the total number of video frames in the video frame group is lower than the predetermined proportion threshold.
6. The method of claim 5, wherein, The step of automatically dividing the time periods and performing the second labeling by a tolerance value algorithm according to a result of the first labeling comprises: determining a tolerance value of the video frame group according to whether each of the plurality of video frame groups marked by the voting algorithm is a key frame group or a non-key frame group, in combination with considering an upper threshold value V of the predetermined tolerance value max and a lower threshold value V of the predetermined tolerance value min ; and automatically dividing the surgical video into a plurality of time segments and secondly labeling each time segment based on the determined tolerance value of each of the plurality of video frame groups.
7. The method of claim 6, wherein, The step of labeling each time segment by the tolerance value algorithm comprises: from the initial time T of the first video frame group of each time period n Initially, the initial time T of each video frame group is sequentially and circularly judged as a key frame group or a non-key frame group n+i Initially, the initial time T of each video frame group is sequentially and circularly judged as a key frame group or a non-key frame group According to the judgment result, the tolerance value V of the video frame group is determined n+i : When at the initial time T of the video frame group n+i When the video frame group is judged as a key frame group, the tolerance value V of the previous video frame group is n+i-1 and the sum V of the predetermined first value a n+i-1 +a and the upper limit threshold V of the predetermined tolerance value max Compare and determine the tolerance value of the video frame group based on the comparison result: If V n+i-1 +a≤V max , let the tolerance value V n+i of the video frame group be V n+i-1 +a; If V n+i-1 +a>V max , let the tolerance value V of the video frame group be n+i-1 =V max ; When at the initial moment T of the video frame group n+i When the video frame group is judged as a non-key frame group, the tolerance value V of the previous video frame group is n+i-1 The value obtained by subtracting the predetermined second value b from the predetermined tolerance value lower limit threshold V min Compare and determine the tolerance value of the video frame group based on the comparison result: If V n+i-1 -b min , let the tolerance value V n+i of the video frame group be V n+i-1 -b; If V n+i-1 -b≤V min , let the tolerance value V n+i of the video frame group be V min ; only when the tolerance value V of the video frame group is determined as V n+i min the cycle is ended, and the time period T n ~T n+i is divided. 8. The method of claim 7, wherein, The step of labeling each time segment by the tolerance value algorithm further comprises: At the end of the cycle and division of the time period T n ~T n+i After that, it is judged whether one or more video frame groups with a tolerance value V min are contained in the time period to mark the time period for the second time: If the time period contains one or more video frame groups with a tolerance value of V min , it is marked as a low value time period; If the time period does not contain one or more video frame groups with a tolerance value of V min , it is marked as a high value time period.
9. The method of claim 8, wherein, The method comprises a primary edit, which comprises: retaining the high-value time segments and deleting the low-value time segments; and generating and exporting a primary-edited surgical video based on the retained high-value time segments.
10. The method of claim 9, wherein, The step of identifying whether the surgical scene of a video frame is a first type of surgical scene or a second type of surgical scene comprises object detection to identify whether the surgical scene is a scene inside a patient, whether surgical instruments appear in the surgical scene, and the location and / or extent of surgical instruments appearing in the surgical scene, wherein, the surgical scene is identified as the first type of surgical scene when the identified surgical scene meets the requirements that the surgical scene itself is identified as being inside a patient, and surgical instruments appear in the surgical scene and are located inside the patient; and the surgical scene is identified as the second type of surgical scene when the identified surgical scene comprises one or more of the following: the surgical scene itself is identified as being outside a patient, surgical instruments do not appear in the surgical scene, and / or surgical instruments appearing in the surgical scene are located outside the patient.
11. The method of claim 10, wherein, When the surgical scene of a video frame is identified as the first type of surgical scene, the method further comprises: identifying the type of surgical instruments appearing in the surgical scene; and generating an instrument type label that labels the corresponding video frame in the surgical video.
12. The method of claim 11, wherein, The method further comprises generating one or more branch videos from the edited surgical video according to the type of surgical instruments, each branch video comprising an aggregation of video frames labeled by the instrument type label of the same type of surgical instruments.
13. The method of claim 10, wherein, The step of labeling the surgical video at least once further comprises a third label, wherein the third label comprises: determining the operational state of the surgical instruments in the surgical scene by comparing each video frame in the plurality of video frames of the retained high-value time segment in the primary edit with its previous video frame; generating a firing label that labels the corresponding video frame in the retained high-value time segment in the primary edit when the operational state of the surgical instruments is determined to be a firing state; and labeling the time segment with the firing label in the retained high-value time segment in the primary edit as a high-value valid time segment.
14. The method of claim 13, wherein, The method further comprises a secondary edit, which comprises: retaining the high-value valid time segments and deleting other time segments in the retained high-value time segment in the primary edit other than the high-value valid time segments; and generating and exporting a secondary-edited surgical video based on the retained high-value valid time segments.
15. The method of claim 13, wherein, The firing state of the surgical instruments is predetermined according to the type of the surgical instruments.
16. The method of claim 15, wherein, When the surgical instrument is an ultrasonic knife, a stapler, and / or a hemostatic clip, the firing state is identified by a switching action between an open state and a closed state of an end effector of the surgical instrument; and when the surgical instrument is an electric knife, the firing state is identified by a smoke generation situation.
17. The method of claim 13, wherein, The computing device is configured to allow a user to click on the firing label and / or the instrument type label in the edited surgical video to jump to a location of a video frame in the surgical video corresponding to the firing label and / or the instrument type label.
18. The method of claim 1, wherein, The computing device is configured to allow a user to make further edits to the edited surgical video.
19. The method of claim 1, wherein, The step of pre-processing the received original surgical video comprises: uniformizing resolutions of the original surgical video into a predetermined resolution; and performing frame extraction processing on the surgical video after the resolutions are uniformized; and uniformizing sizes of each of a plurality of video frames of the surgical video after the frame extraction processing.
20. A computing device comprising: The computing device is configured to be capable of performing the method of any one of claims 1-19.
21. A computer readable medium stored in a computing device, comprising: The computer readable medium is configured to be capable of performing the method of any one of claims 1-19.
22. A surgical system, characterized by, The surgical system comprises: a computing device according to claim 20; one or more surgical instruments associated with the computing device; a camera device located in a surgical scene and associated with the computing device, the camera device being configured to be capable of capturing an original surgical video; and a display device associated with the computing device, the display device being configured to be capable of displaying an edited surgical video. The computing device is configured to allow a user to click on the firing label and / or the instrument type label in the edited surgical video to jump to a location of a video frame in the surgical video corresponding to the firing label and / or the instrument type label. The computing device is configured to allow a user to make further edits to the edited surgical video. The step of pre-processing the received original surgical video comprises: uniformizing resolutions of the original surgical video into a predetermined resolution; and performing frame extraction processing on the surgical video after the resolutions are uniformized; and uniformizing sizes of each of a plurality of video frames of the surgical video after the frame extraction processing. The computing device is configured to be capable of performing the method of any one of claims 1-19. The computer readable medium is configured to be capable of performing the method of any one of claims 1-19. The surgical system comprises: a computing device according to claim 20; one or more surgical instruments associated with the computing device; a camera device located in a surgical scene and associated with the computing device, the camera device being configured to be capable of capturing an original surgical video; and a display device associated with the computing device, the display device being configured to be capable of displaying an edited surgical video.