Method, computing device and surgical system for automatically editing surgical video
The method employs AI to analyze surgical video frames, identifying critical scenes and instrument manipulation for efficient and accurate editing, addressing the inefficiencies of manual editing and reducing storage needs.
Patent Information
- Application Number
- PCT/EP2025/060965
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-24
- Filing Date
- 2025-04-23
- Publication Date
- 2025-10-30
AI Technical Summary
Manual editing of surgical videos is time-consuming and inefficient, often resulting in inaccurate clipping and increased storage requirements due to the inclusion of non-value-added segments, which hinders their effective use in teaching, training, and research.
A method utilizing machine learning and artificial intelligence to analyze surgical video frames, identify critical and non-critical scenes, and automatically edit the video by retaining or deleting time periods based on scene type and instrument manipulation, allowing for efficient and accurate video editing.
Reduces manpower and storage demands by automatically identifying and labeling important segments, enabling efficient editing and clipping of surgical videos, improving their usability and reducing time consumption.
Smart Images

Figure EP2025060965_30102025_PF_FP_ABST
Abstract
Description
[0001] METHOD, COMPUTING DEVICE AND SURGICAL SYSTEM FOR AUTOMATICALLY EDITING SURGICAL VIDEO
[0002] TECHNICAL FIELD
[0003] The present invention relates to a method for assisted video editing on the basis of an automatic surgical video recognition technique. More particularly, the present invention relates to a method implemented by a computer device for automatically editing a surgical video, a computing device capable of implementing the method, a computer-readable medium, and a surgical system comprising the computing device.
[0004] BACKGROUND OF THE INVENTION
[0005] A surgical operation process is usually recorded by means of an image capture device in an image capture system, e.g., shooting a video. During a minimally invasive surgical process, a surgical video is recorded by means of an endoscopic imaging system. The endoscopic surgical video is an important form of recording of the surgery, and is a necessary material for various tasks such as surgical review, training, quality control, and scientific research.
[0006] However, a raw video document of the entire surgery is huge, possibly up to 10 G or more, which places pressure on storage. In addition, there are a range of specific time periods during the surgical operation process which may be of little value for review, training, quality control, and scientific research, for example: a time period in which the endoscope is being scrubbed outside the body, a time period in which a surgical instrument does not appear in the visual field, a time period in which a surgical manipulation is not performed, and / or a time period in which the position at which the surgical instrument being manipulated is being fine-tuned, etc. Were these time periods retained, reviewing of a video would become very inefficient. Since it is not possible to determine which time periods in the raw surgical video are such “specific time periods”, an observer must wait patiently and watch the entire video.
[0007] Therefore, in order to use the video of a minimally invasive surgery in an efficient way, it is essential to edit so as to reduce the raw surgical video. At present, a raw video can generally only be edited or cut / clipped to eliminate these low-value or valueless time periods manually. However, manual editing is very time-consuming. The surgical video of a typical surgery is as long as a few hours or even ten hours or more. It would take a medical worker a great amount of precious time to re-watch these raw surgical videos or manually edit or clip the videos, which limits extensive use of the videos of surgical operations in teaching, training, exchange, scientific research, etc. In addition, it may not be possible to perform manual editing or clipping of a video very accurately, and there are risks such as instability of the level of performance and high frequency of erroneous operations.
[0008] Accordingly, there is a desire to provide a smart method for editing or clipping a surgical video, as well as a computer-readable medium, a computing device, and a surgical system capable of implementing the method, so as to allow automatic performance, by means of artificial intelligence, of analysis, editing, clipping, exporting, storage, and / or display of a raw surgical video recorded by a shooting device, and thereby achieve automatically editing the surgical video in a convenient, efficient and accurate manner, reduce labor time consumption, and lower the requirements for hardware.
[0009] SUMMARY OF THE INVENTION
[0010] Objects of the present invention lie in providing a method implemented by a computing device for automatically editing a surgical video, and a related computing device, computer-readable medium and surgical system, to at least partially solve the described problems in the prior art, so as to allow a user to obtain an automatically edi ted / clipped surgical video with high efficiency and high accuracy, and with greatly reduced manpower input, thereby reducing pressure on storage, reducing time consumption, and facilitating subsequent viewing, editing and learning by a user.
[0011] To at least achieve the described objects, according to an aspect of the present invention, there is provided a method implemented by a computing device for automatically editing a surgical video, said method comprising:
[0012] Receiving a raw surgical video, said raw surgical video being configured to be captured by an image capture device which is associated with said computing device and which is located in a surgical scene;
[0013] Pre-processing the received raw surgical video;
[0014] Performing machine learning analysis on each video frame of a plurality of video frames of the pre-processed surgical video, so as to identify whether the surgical scene of each video frame is a first type of surgical scene or a second type of surgical scene; Performing labeling on said surgical video at least once, according to the result of the identification for each video frame of said plurality of video frames;
[0015] Determining retention or deletion of each time period of a plurality of time periods in said surgical video according to the result of said labeling;
[0016] Generating and exporting an edited surgical video on the basis of retained time periods;
[0017] Storing said edited surgical video; and
[0018] Displaying said edited surgical video on a display device associated with said computing device.
[0019] In some embodiments, the step of performing labeling on said surgical video at least once comprises:
[0020] Grouping, in time sequence, the plurality of video frames of the surgical video into a plurality of video frame groups, each video frame group comprising a predetermined number of video frames;
[0021] Performing first labeling on each video frame group individually, according to the result of the identification for each video frame of said plurality of video frames;
[0022] Automatically dividing said surgical video into a plurality of time periods according to the result of the first labeling, each time period of said plurality of time periods comprising one or more video frame groups;
[0023] Performing second labeling on each time period of the plurality of time periods individually; and
[0024] Determining retention or deletion of each time period of the plurality of time periods according to the result of the second labeling.
[0025] In some embodiments, said first labeling is performed on each video frame group of said plurality of video frame groups individually by means of a voting algorithm, and after said first labeling, the surgical video is automatically divided into time periods by means of a tolerance value algorithm according to the result of the first labeling obtained by means of the voting algorithm, and said second labeling is performed on the time periods. In some embodiments, the step of performing the first labeling on each video frame group of said plurality of video frame groups individually by means of a voting algorithm comprises:
[0026] According to the ratio of the number of video frames identified as said first type of surgical scene in each video frame group of said plurality of video frame groups to the total number of video frames in the video frame group, performing labeling on each video frame group by means of the voting algorithm.
[0027] In some embodiments, the step of performing labeling on each video frame group by means of the voting algorithm comprises:
[0028] When said ratio of the number of video frames identified as said first type of surgical scene in each video frame group to the total number of video frames in the video frame group is not lower than a predetermined ratio threshold value, labeling the video frame group as a critical frame group;
[0029] When said ratio of the number of video frames identified as said first type of surgical scene in each video frame group to the total number of video frames in the video frame group is lower than the predetermined ratio threshold value, labeling the video frame group as a non-critical frame group.
[0030] In some embodiments, the steps of automatically dividing into time periods by means of a tolerance value algorithm according to the result of the first labeling and performing said second labeling comprise:
[0031] According to whether each video frame group of the plurality of video frame groups is a critical frame group or a non-critical frame group as labeled by means of the voting algorithm, and in combination with a consideration of a predetermined tolerance value upper-limit threshold value Vmax and a predetermined tolerance value lower-limit threshold value Vmin, determining the tolerance value of the video frame group; and
[0032] According to the determined tolerance value of each video frame group of said plurality of video frame groups, automatically dividing the surgical video into a plurality of time periods, and performing the second labeling on each time period.
[0033] In some embodiments, the step of performing labeling on each time period by means of a tolerance value algorithm comprises: Starting from an initial time Tnof the first video frame group in each time period, sequentially and recurrently judging, at an initial time Tn+i of each video frame group, whether the video frame group is a critical frame group or a non-critical frame group;
[0034] Determining the tolerance value Vn+i of the video frame group according to the result of the judgment:
[0035] When, at the initial time Tn+i of the video frame group, the video frame group is judged to be a critical frame group, comparing the sum Vn+i-i+a of the tolerance value Vn+i-i of the previous video frame group and a predetermined first numerical value a with the predetermined tolerance value upper-limit threshold value Vmax, and on the basis of the result of the comparison, determining the tolerance value of the video frame group:
[0036] If Vn+i-i+a < Vmax, let the tolerance value of the video frame group Vn+i = Vn+i-i+a; and
[0037] If Vn+i-i+a > Vmax, let the tolerance value of the video frame group Vn+i-l Vmax,
[0038] When, at the initial time Tn+i of the video frame group, the video frame group is judged to be a non-critical frame group, comparing the value obtained by subtracting a predetermined second numerical value b from the tolerance value Vn+i- i of the previous video frame group with the predetermined tolerance value lower- limit threshold value and on the basis of the result of the comparison, determining the tolerance value of the video frame group:
[0039] If Vn+i-i-b > Vmin, let the tolerance value of the video frame group Vn+i = Vn+i-i-b; and
[0040] If Vn+i-i-b < Vmin, let the tolerance value of the video frame group Vn+i Vmin,
[0041] Only when the tolerance value Vn+i of the video frame group is determined to be Vmin, terminating the recurrence and performing division to form the time period Tnto Tn+i.
[0042] In some embodiments, the step of performing labeling on each time period by means of a tolerance value algorithm further comprises: After terminating the recurrence and performing division to form the time period Tnto Tn+i, determining whether the time period comprises one or more video frame groups having a tolerance value of Vmin, so as to perform the second labeling on the time period:
[0043] If the time period comprises one or more video frame groups having a tolerance value of labeling the time period as a low-value time period; and
[0044] If the time period does not comprise one or more video frame groups having a tolerance value of Vmin, labeling the time period as a high-value time period.
[0045] In some embodiments, said method comprises first clipping, said first clipping comprising:
[0046] Retaining said high-value time period and deleting said low-value time period; and
[0047] On the basis of the retained high-value time period, generating and exporting a first clipped surgical video.
[0048] In some embodiments, the step of identifying whether the surgical scene of the video frame is a first type surgical scene or a second type surgical scene comprises target detection to identify whether said surgical scene is a scene inside or outside the body of a patient, whether a surgical instrument appears in said surgical scene, and the position and / or area at which the surgical instrument appears, wherein,
[0049] When the surgical scene being identified meets the following requirements, the surgical scene is identified as the first type of surgical scene: the surgical scene itself is identified as being located inside the body of a patient, and in addition, a surgical instrument appears in the surgical scene, and the surgical instrument is located inside the body of the patient; and
[0050] When the surgical scene being identified involves one or more of the following situations, the surgical scene is identified as the second type of surgical scene: the surgical scene itself is identified as being located outside the body of the patient, no surgical instrument appears in the surgical scene, and / or the surgical instrument that appears in the surgical scene is located outside the body of the patient.
[0051] In some embodiments, when the surgical scene of the video frame is identified as said first type of surgical scene, said method further comprises:
[0052] Identifying the type of the surgical instrument that appears in said surgical scene; and Generating an instrument type tag, the instrument type tag labeling the corresponding video frame in said surgical video.
[0053] In some embodiments, said method further comprises generating one or more video subsets from the clipped surgical video according to the type of the surgical instrument, each video subset comprising an aggregation of the video frames labeled with the instrument type tag for the same type of surgical instrument.
[0054] In some embodiments, the step of performing labeling on said surgical video at least once further comprises third labeling, wherein the third labeling comprises:
[0055] Determining a manipulation state of the surgical instrument in the surgical scene by comparing each video frame of the plurality of video frames in the high-value time period retained in said first clipping with the previous video frame;
[0056] Generating a firing tag when the manipulation state of said surgical instrument is determined to be a firing state, the firing tag labeling a corresponding video frame in the high-value time period retained in said first clipping; and
[0057] Labeling the firing tag-bearing time period in the high-value time period retained in said first clipping as a high-value effective time period.
[0058] In some embodiments, said method further comprises second clipping, said second clipping comprising:
[0059] Retaining said high-value effective time period, and deleting the time periods in the high-value time period retained in said first clipping other than said high-value effective time period; and
[0060] On the basis of the retained high-value effective time period, generating and exporting a second clipped surgical video.
[0061] In some embodiments, the firing state of said surgical instrument is predetermined according to the type of said surgical instrument.
[0062] In some embodiments, when said surgical instrument is an ultrasonic scalpel, an anastomat / stapler, and / or a hemostatic clamp, said firing state is identified in view of a switching action between an open state and a closed state of an end effector of said surgical instrument; and when said surgical instrument is an electrical knife, said firing state is identified in view of smoke generation. In some embodiments, said computing device is configured to allow a user to click on said firing tag and / or instrument type tag in the edited surgical video, so as to jump to the location of the video frame in said surgical video that corresponds to said firing tag and / or instrument type tag.
[0063] In some embodiments, said computing device is configured to allow a user to perform further editing on said edited surgical video.
[0064] In some embodiments, the step of pre-processing the received raw surgical video comprises:
[0065] Normalizing the resolution of said raw surgical video to a predetermined resolution;
[0066] Performing frame extraction processing on the surgical video after the resolution normalization; and
[0067] Normalizing the size of each video frame of the plurality of video frames of the surgical video after the frame extraction processing to a predetermined size.
[0068] According to another aspect of the present invention, there is provided a computing device, said computing device being configured to be capable of implementing the described method.
[0069] According to another aspect of the present invention, there is provided a computer- readable medium stored in the computing device, said computer-readable medium being configured to be capable of implementing the described method.
[0070] According to a yet another aspect of the present invention, there is provided a surgical system, said surgical system comprising:
[0071] The described computing device;
[0072] One or more surgical instruments associated with the computing device;
[0073] An image capture device located in a surgical scene and associated with said computing device, said image capture device being configured to be capable of capturing a raw surgical video; and a display device associated with said computing device, said display device being configured to be capable of displaying an edited surgical video. Using the method, computing device, computer-readable medium and surgical system according to the present invention, it is possible to, by means of artificial intelligence technology, automatically identify and label important segments in a surgical video, retain the labeled important segments and eliminate non-critical segments, so as to obtain a surgical video on which first clipping or second clipping can be performed according a user’s choice, thereby providing surgical videos having undergone different degrees of clipping. This greatly reduces the time consumption of manpower and the pressure on storage in the course of editing a surgical video, and effectively improves the use efficiency of the surgical video.
[0074] BRIEF DESCRIPTION OF THE DRAWINGS
[0075] To better understand the above and other objectives, features, advantages and functions of the present invention, reference may be made to preferred embodiments shown in the accompanying drawings. The same reference numerals refer to the same components in the accompanying drawings. It should be understood by those skilled in the art that the drawings are intended to schematically illustrate the preferred embodiments of the present invention and make no limitation on the scope of the present invention, and various members are not drawn to scale.
[0076] FIG. 1 illustrates a computing device-implemented method for automatically editing a surgical video in accordance with an aspect of the present invention;
[0077] FIG. 2 illustrates specific steps of pre-processing the received raw surgical video in the method according to the present invention;
[0078] FIG. 3 illustrates specific steps of performing labeling on the surgical video at least once in the method according to the present invention;
[0079] FIG. 4 illustrates a flow chart of second labeling in the method according to the present invention; and
[0080] FIG. 5 illustrates specific steps of third labeling in the method according to the present invention.
[0081] DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
[0082] Preferred embodiments of the present invention will be described in detail with reference to the drawings. The embodiments described herein are merely preferred embodiments according to the present invention, and on the basis of the preferred embodiments, those skilled in the art could conceive of other modes capable of implementing the present invention, which also fall within the scope of the present invention.
[0083] In a surgical task, it is often necessary, for purposes of subsequent viewing, review, teaching, improvement, etc., to manually edit or clip a raw surgical video recorded during a surgical manipulation, such as a surgical operation, by a shooting device, such as an endoscope in a minimally invasive surgical system, disposed in a surgical operation scene within a surgical operating room. Since a surgical operation often takes a long time, manual editing is correspondingly very time-consuming, resulting in a waste of manpower, time, costs, etc. In addition, it may not be possible to perform manual editing or clipping very accurately, and there are risks such as instability of the level of performance and high frequency of erroneous operations. Accordingly, the present invention provides a smart method for editing or clipping a surgical video, a computing device, and a system, so as to at least partially solve the described problems existing in the prior art.
[0084] With reference to FIG. 1, an aspect of the present invention provides a computing device-implemented method 10 for automatically editing a surgical video, the method 10 comprising the following steps:
[0085] Receiving 110 a raw surgical video, the raw surgical video being configured to be captured by an image capture device which is associated with the computing device and which is located in a surgical scene;
[0086] Pre-processing 120 the received raw surgical video;
[0087] Performing machine learning algorithm analysis on each video frame of a plurality of video frames of the pre-processed surgical video, so as to identify 130 whether the surgical scene of each video frame is a first type of surgical scene or a second type of surgical scene;
[0088] Performing labeling 140 on the surgical video at least once according to the result of identification for each video frame of the plurality of video frames;
[0089] Determining 150 retention or deletion of each time period of a plurality of time periods in the surgical video according to the result of the labeling;
[0090] Generating and exporting 160 an edited surgical video on the basis of retained time periods;
[0091] Storing 170 the edited surgical video; and Displaying 180 the edited surgical video on a display device associated with the computing device.
[0092] It will be appreciated that the described image capture device may be any image capture device capable of recording a video of a surgical process, which could be conceived of for a surgical application or operation, including, but not limited to, an image capture device for an endoscope used in a minimally invasive surgical process in a surgical operating room. The described computing device may be any computer or computing system associated with the corresponding surgical manipulation. For example, the computing device may be an auxiliary computing device for use in endoscopic surgery, which can perform retraining on an existing open source pre-trained model (such as YOLO) by means of machine learning using training data obtained, e.g., in a previous surgical manipulation, so as to form an artificial intelligence (Al) model that is suitable for the present invention and can implement the method of the present invention, thereby automatically editing a raw surgical video from the image capture device.
[0093] Further preferably, the pre-processing 120 the received raw surgical video may comprise the steps shown in FIG. 2:
[0094] Normalizing 121 the resolution of the raw surgical video to a predetermined resolution;
[0095] Performing frame extraction 122 processing on the surgical video after the resolution normalization; and
[0096] Normalizing 123 the size of each video frame of the plurality of video frames of the surgical video after the frame extraction processing to a predetermined size.
[0097] In the step of resolution normalization, preferably, no matter how high the resolution of the raw surgical video is (e.g., the resolution is 4k or 2k), the resolution is normalized to the predetermined resolution, e.g., to 1080p. In this step, the size of the data source and the frame rate of the raw video are retained. In the step of frame extraction, preferably, extraction of frame pictures is performed on the surgical video after resolution normalization; that is, only some video frames of the plurality of video frames of the surgical video are retained at a uniform interval or at non-uniform intervals. For example, preferably, assuming that the frame rate of the raw surgical video is 60 Hz, only nine video frame pictures are uniformly retained per second in the step of frame extraction, so that the number of video frames that need to be processed is greatly decreased (e.g., decreased to 15% of the raw video) without losing accuracy in subsequent recognition, thereby reducing the burden of storage and processing of the computing device. In the step of video frame size normalization, preferably, the image of each video frame of the plurality of video frames after the frame extraction processing is embedded into a canvas of a specified size using a letterboxing method so that the image is expanded to a network input size while maintaining an aspect ratio. This step of size normalization facilitates standardization of each frame image in the surgical video to improve the accuracy of target detection in the subsequent step of identifying by performing analysis.
[0098] Further, in step 130 of the method 10 according to the present invention, by means of a machine learning-associated algorithm model, e.g., an algorithm model as described above which is obtained by performing re-training on an existing open source pre-trained model (such as YOLO) by using training data obtained in a previous surgical manipulation, and which is capable of accurately identifying e.g., a surgical instrument and the position at which it is located (e.g., located inside or outside the body of a patient, etc.), analysis, e.g., target detection, is performed on each video frame of the plurality of video frames in the pre- processed surgical video, so as to identify whether the surgical scene shown in each video frame belongs to the first type of surgical scene or the second type of surgical scene, e.g., to identify whether the surgical scene shown in the video frame being processed is a scene inside or outside the body of the patient, whether the surgical instrument appears in the surgical scene, and the position and / or area at which the surgical instrument appears. Exemplarily, when the surgical scene being identified meets the following requirements, the surgical scene is identified as the first type of surgical scene: the surgical scene itself is identified as being located inside the body of a patient, and in addition, a surgical instrument appears in the surgical scene and the surgical instrument is identified as being located inside the body of the patient; and when the surgical scene being identified involves the following situations, the surgical scene is identified as the second type of surgical scene: the surgical scene itself is identified as being located outside the body of the patient, and no surgical instrument appears in the surgical scene; and / or the surgical instrument that appears in the surgical scene is identified as being located outside the body of the patient.
[0099] Preferably, when the surgical scene of the video frame is identified as the described first type of surgical scene, the type of the surgical instrument that appears in the surgical scene may be further identified, and an instrument type tag generated, the instrument type tag labeling the corresponding video frame in the surgical video (e.g., a final clipped surgical video). Such a tag allows subsequent processing to represent the corresponding type of surgical instrument by distinguishing same from other instruments in the clipped surgical video. For example, different video subsets are generated to collectively represent corresponding types of surgical instrument, each subset representing all segments of the entire surgical video in which one type of surgical instrument appears (that is, each subset of the video comprising an aggregation of the video frames labeled with the instrument type tag for the same type of surgical instrument). The instrument type tag may further allow a user to click on the clipped surgical video, so as to jump to the corresponding video frame or time period.
[0100] Further, in step 140 of the method 10 according to the present invention, labeling is performed on the surgical video at least once according to the result of the described identification for target detection performed on each video frame of the plurality of video frames of the pre-processed surgical video. Preferably, with reference to FIG. 3, the step of labeling comprises the following specific steps:
[0101] Grouping 141, in time sequence, the plurality of video frames of the surgical video after the described steps 110 through 130 into a plurality of video frame groups, each video frame group comprising a predetermined number of video frames;
[0102] Performing first labeling 142 on each video frame group individually according to the result of the described identification of each video frame of the plurality of video frames;
[0103] Automatically dividing 143 the surgical video into a plurality of time periods according to the result of the first labeling, each time period of the plurality of time periods comprising one or more video frame groups;
[0104] Performing second labeling 144 on each time period of the plurality of time periods individually; and
[0105] Determining 145 retention or deletion of each time period of the plurality of time periods according to the result of the second labeling;
[0106] More preferably, in the described step of labeling, the first labeling is performed on each video frame group of the plurality of video frame groups individually by means of a voting algorithm, and after the first labeling, the surgical video is automatically divided into time periods by means of a tolerance value algorithm according to the result of the first labeling obtained by means of the voting algorithm, and the second labeling is performed on the time periods.
[0107] Preferably, in step 141, the plurality of video frames are grouped into a plurality of video frame groups, the number of video frames in each group being predetermined. The number of video frames in each group may be the same as or different from the other groups. The number of video frames may be predetermined according to one or more factors, such as empirical data of surgical manipulation, actual conditions, user-specific requirements, accuracy consideration, machine training parameters, etc. For example, in the present invention, grouping may be performed in units of a second. That is, the plurality of video frames of the surgical video after being processed in the above-mentioned steps 110 through 130 are grouped in units of a second in time sequentice, each video frame group comprising nine video frames arranged in time sequence. It is appreciated that the means of grouping is not limited to the described grouping of nine video frames in one second, but rather, other numbers of video frames may be grouped into one group according to need in practical use.
[0108] Further, preferably, in step 142 after the described grouping step 141, the first labeling 142 is performed on each video frame group sequentially according to the result of the described identification of each video frame of the plurality of video frames of the surgical video in step 130. Preferably, for each video frame group of the plurality of video frame groups, the ratio of the number of video frames identified as the described first type of surgical scene to the total number of video frames in the video frame group is calculated, and labeling is performed on each video frame group of the plurality of video frame groups sequentially by means of the voting algorithm. The voting algorithm is specifically such that when the calculated ratio of the number of video frames identified as the first type of surgical scene in each video frame group to the total number of video frames in the video frame group is not lower than a predetermined ratio threshold value, the video frame group is labeled as a critical frame group; when the described ratio is lower than the predetermined ratio threshold value, the video frame group is labeled as a non-critical frame group.
[0109] The described predetermined ratio threshold value depends, for example, on the data distribution in the algorithm training data and the degree of feature extraction, and may, for example, be predetermined according to one or more factors such as data of surgical medical experience, training parameters of a machine learning model, and requirements of actual use, and stored in the computing device of the present invention. Furthermore, the ratio threshold value may also be fine-tuned with artificial aid. The ratio threshold value may be predetermined according to the type of instrument. For example, different surgical instruments may correspond to different ratio threshold values, so that, correspondingly, efficiency of identification is improved while ensuring the identification accuracy rate of each surgical instrument and reducing mistakes or omission in judgment. In one preferred embodiment, the ratio threshold value may be 2 / 3. That is, as long as six or more out of the nine video frames in each group are identified as video frames of the first type of surgical scene, the video frame group is labeled as a critical frame group, and if less than six, then labeled as a non-critical frame group. (By non-critical frame group, it is meant that the proportion of the video frames in which effective information occurs, e.g., the surgical instrument located in the body of a patient occurs, in the video frame group is relatively low, and therefore the video frame group is not critical, and can be labeled in this step so as to be eliminated in the subsequent processing process).
[0110] After performing the first labeling step 142 by means of the voting algorithm, automatic division 143 is performed, according to the result of the labeling, on the surgical video having been processed in the aforementioned steps so as to form a plurality of time periods, each time period formed from the division comprising one or more video frame groups, and the second labeling 144 is performed on each time period respectively.
[0111] Preferably, the concept of a tolerance value is introduced in the described related steps of dividing into time periods and performing the second labeling. By means of a tolerance value algorithm, judgment with respect to the tolerance value is performed on the plurality of video frame groups in the surgical video sequentially in time sequence (the result of the judgment by means of the tolerance value algorithm indicates the distribution of critical frame groups and non-critical frame groups in each time period, which can provide a basis for the subsequent processing process with respect to retention or deletion of the time periods), and on that basis, the surgical video is divided into time periods and labeling is performed on the time periods. In particular, the steps of automatically dividing into time periods by means of a tolerance value algorithm according to the result of the first labeling and performing the second labeling comprise: according to whether each video frame group of the plurality of video frame groups is a critical frame group or a non-critical frame group, as labeled by means of the voting algorithm, and in combination with a consideration of a predetermined tolerance value upper-limit threshold value Vmax and a predetermined tolerance value lower-limit threshold value Vmin, determining the tolerance value of the video frame group; and, according to the determined tolerance value of each video frame group of the plurality of video frame groups, automatically dividing the surgical video into a plurality of time periods, and performing the second labeling on each time period.
[0112] More specifically, with reference to the flow chart shown in FIG. 4, the steps associated with the described tolerance value are explained in detail:
[0113] Starting from an initial time Tnof the first video frame group in each time period of the plurality of time periods, sequentially and recurrently judging, at an initial time Tn+i of each video frame group, whether the video frame group is a critical frame group or a non- critical frame group;
[0114] Determining the tolerance value Vn+i of the video frame group according to the result of the judgment:
[0115] When, at the initial time Tn+i of the video frame group, the video frame group is judged to be a critical frame group, comparing the sum Vn+i-i+a of the tolerance value Vn+i-i of the previous video frame group and a predetermined first numerical value a with the predetermined tolerance value upper-limit threshold value Vmax, and on the basis of the result of the comparison, determining the tolerance value of the video frame group:
[0116] If Vn+i-i+a < Vmax, let the tolerance value of the video frame group Vn+i = Vn+i-i+a; and
[0117] If Vn+i-i+a > Vmax, let the tolerance value of the video frame group Vn+i-l Vmax,
[0118] When, at the initial time Tn+i of the video frame group, the video frame group is judged to be a non-critical frame group, comparing the value obtained by subtracting a predetermined second numerical value b from the tolerance value Vn+i- i of the previous video frame group with the predetermined tolerance value lower- limit threshold value Vmin, and on the basis of the result of the comparison, determining the tolerance value of the video frame group:
[0119] If Vn+i-i-b > Vmin, let the tolerance value of the video frame group Vn+i
[0120] = Vn+i-i-b; and If Vn+i-i-b < Vmin, let the tolerance value of the video frame group Vn+i
[0121] Vmin,
[0122] Only when the tolerance value Vn+i of the video frame group is determined to be Vmin. terminating the recurrence for the time period Tnto Tn+i and performing division to form the time period Tnto Tn+i, and otherwise, continuing the recurrence for the time period until the tolerance value becomes Vmin.
[0123] It should be noted that, in the described flow path forjudging the tolerance value, T represents time, V represents tolerance value, and the subscripts thereof represent a certain video frame group in a certain time period progressing in time sequence. For example, Tnrepresents the start time of the first video frame group in the nth time period, and Vnrepresents the tolerance value of the first video frame group in the nth time period; Tn+i represents the start time of the second video frame group of the nth time period, Vn+i represents the tolerance value of the second video frame group of the n-th time period..., Tn+i represents the start time of the (i-l)th video frame group of the nth time period, and Vn+i represents the tolerance value of the (i-l)th video frame group of the nth time period. The above-mentioned n and i are both integers. In addition, the described predetermined first numerical value a is a positive number, and the described predetermined second numerical value b is a negative number.
[0124] After terminating the recurrence and performing dividing to form the time period Tnto Tn+i, determining whether the time period comprises one or more video frame groups having a tolerance value of Vmin, so as to perform the second labeling on the time period: if the time period comprises one or more video frame groups having a tolerance value of Vmin, labeling same as a low-value time period; and if the time period does not comprise one or more video frame groups having a tolerance value of Vmin, labeling same as a high-value time period.
[0125] It may be appreciated that the described steps of performing division to form a time period and performing the second labeling may be performed in parallel or, alternately, in real time, or the second labeling may be performed after the entire video has been divided into time periods. The algorithms may be adjusted according to practical requirements.
[0126] After the second labeling, first clipping is performed on the surgical video; that is, the high-value time period is retained and the low-value time period is deleted, and on the basis of the high-value time period retained, a first clipped surgical video is generated and exported.
[0127] Next, with reference to FIG. 5, the described step of the second labeling is explained by way of example. In this example, only a time period of eight seconds after the start of the video is explained, to aid in understanding the described tolerance value algorithm. Assume that, starting from time 0, tolerance values are sequentially judged in units of a second (that is, the plurality of video frames in each second is grouped as one group, e.g., the described nine video frames in each second are grouped as one video frame group) in time sequence. Each video frame group has been subjected to the first labeling to be labeled as a critical frame group or a non-critical frame group. Assume that the results of the first labeling of the eight video frame groups during the period of 0 to 8 seconds are as shown in the table below. Assume that the initial tolerance value is 0. The described predetermined first numerical value a is equal to 1, and the described predetermined second numerical value b is equal to -1. That is, when a critical frame group is encountered, the tolerance value is refreshed on the basis of the current tolerance value, i.e., numerical value 1 is added to the tolerance value, and when a non-critical frame group is encountered, the tolerance value is consumed, i.e., the numerical value 1 is subtracted from the tolerance value. In addition, it is assumed in this example that the described upper-limit tolerance value threshold value Vmax is 2, and the described tolerance value lower-limit threshold value is 0.
[0128] Table 1. Results of the first labeling of respective video frame groups
[0129] As shown in Table 1, at 0 to 1 seconds, it is judged that the result of the first labeling of the first video frame group is “critical frame group”. Therefore, 1 is added on the basis of the initial tolerance value 0, i.e., the tolerance value is refreshed to become 1. Thereafter, it is judged that the result of the first labeling at 1 to 2 seconds is also “critical frame group”. Therefore, the tolerance value is refreshed again, and the tolerance value becomes 2. At 2 to 3 seconds, the result of the first labeling of the third video frame group is still “critical frame group”, but at this time, the sum of the current tolerance value 2 and the predetermined first numerical value 1 is greater than the predetermined tolerance value upper-limit threshold value 2. Therefore, the tolerance value is no longer refreshed, but is kept unchanged and remains at 2. Thereafter, at 3-4 seconds, the result of the first labeling of the fourth video frame group is “non-critical frame group”. Therefore, the tolerance value is consumed, i.e., 1 is subtracted from the tolerance value, and the tolerance decreases to 1. At 4-5 seconds, the result of the first labeling of the fifth video frame group is still “non-critical frame group”. Therefore, the tolerance value is further consumed, i.e., the tolerance decreases to 0. At this time, since the current tolerance value is not greater than the predetermined tolerance value lower-limit threshold value the recurrence is terminated and division is performed to form a time period; that is, a time period comprising 0 to 4 seconds is formed.
[0130] After terminating the recurrence and performing division to form the time period, judgment of the tolerance value is continued by continuing in time sequence and starting from the sixth video frame group at 5 to 6 seconds. Since the result of the first labeling of the sixth video frame group is still “non-critical frame group”, the tolerance value is still not greater than 0. The recurrence is terminated and division is performed to form a time period; that is, a time period comprising 4 to 5 seconds is formed.
[0131] Judgment of the tolerance value is continued by continuing in time sequence and starting from the seventh video frame group at 6 to 7 seconds. Since the result of the first labeling of the seventh video frame group is “critical frame group”, the tolerance value is refreshed to become 1, and the recurrence is continued. Since the result of the first labeling of the eighth video frame group at 7 to 8 seconds is still “critical frame group”, the tolerance value becomes 2. The recurrence for this time period is continued until the tolerance value is consumed to become 0. The recurrence is terminated and division is performed to form a time period.
[0132] After the described steps, the plurality of time periods as described above have been formed by performing division, e.g., the time period comprising 0 to 4 seconds and the time period comprising 4 to 5 seconds. Thereafter, each of the time periods is judged to determine whether it comprises a video frame group having a tolerance value of 0, so as to perform the second labeling. For example, as described above, the time period of 0 to 4 seconds is judged not to comprise a video frame group having a tolerance value of 0, and the tolerance values corresponding to the respective video frame groups of the time period are 1, 2, 2, and 1 sequentially in time sequence. Therefore, the result of the second labeling of the time period is “high-value time period”. The time period of 4 to 5 seconds comprises only one video frame group, and the corresponding tolerance value is 0. Therefore, the result of the second labeling for this time period is “low-value time period”. The steps of performing division to form a time period and performing the second labeling may be performed in parallel or alternately in real time, or the second labeling may be performed after the entire video has been divided into time periods. The algorithms may be adjusted according to practical requirements. For example, the second labeling can be performed immediately after each division to obtain a time period, or the second labeling can be performed all at once after the entire video has been divided into time periods.
[0133] After the second labeling, first clipping is performed on the surgical video; that is, the high-value time period is retained and the low-value time period is deleted, and on the basis of the high-value time period retained, a first clipped surgical video is generated and exported.
[0134] After undergoing the described steps, a clipped surgical video is obtained which only comprises the high-value time period, and from which low-value unnecessary time periods, e.g., a plurality of segments that a user does not need, such as those in which there is no surgical instrument in the visual field, the endoscope is located outside the body of a patient, and the surgical instrument is located outside the body, have been eliminated from the raw video by means of artificial intelligence. Thus, without needing to perform manual clipping, the user can obtain an automatically clipped surgical video which has a greatly reduced length. Such a clipped surgical video only retains useful high-value segments, making it convenient for the user to review, search and learn, and greatly reducing the pressure on a memory.
[0135] Furthermore, the method according to the present invention also enables more elaborate processing of the surgical video. For example, on the basis of the surgical video having been processed in the described steps, third labeling 146 is performed. The third labeling comprises the following steps as shown in FIG. 5: Determining 1461 a manipulation state of the surgical instrument in the surgical scene by means of comparing each video frame of the plurality of video frames in the high- value time period retained in the first clipping with the previous video frame preceding same;
[0136] Generating 1462 a firing tag when the manipulation state of the surgical instrument is determined to be a firing state, the firing tag labeling a corresponding video frame in the high-value time period retained in the first clipping; and
[0137] Labeling 1462 the firing tag-bearing time period in the high-value time period retained in the first clipping as a high-value effective time period.
[0138] The firing state of the surgical instrument may be predetermined according to different types of surgical instrument. As a non-limiting example, when the surgical instrument is an ultrasonic scalpel, an anastomat, and / or a hemostatic clamp, the firing state is identified in view of a switching action between an open state and a closed state of an end effector of the surgical instrument (e.g., whether the video frame is in a firing state can be identified in view of the opening / closing of the end effector of the surgical instrument in two adjacent video frame pictures); and when the surgical instrument is an electrical knife, the firing state is identified in view of smoke generation (e.g., whether the video frame is in a firing state can be identified in view of a change in smoke in two adjacent video frame pictures).
[0139] Similar to the described instrument type tag, when the user uses and views the first clipped surgical video, the user can click on the firing tag so as to directly jump positions to the position of the video frame or time period to which the tag corresponds.
[0140] It should be noted that the described steps of labeling three times can be performed sequentially, alternately, or in parallel. For example, the third labeling step can be performed in parallel or alternately with respect to the second labeling step, or can be performed after completing the second labeling step for the entire video and the first clipping step, which may be selected according to actual requirements.
[0141] After the third labeling step, second clipping on the basis of the first clipping is performed. The second clipping comprises: retaining the high-value effective time period, and deleting the time periods in the high-value time period retained in the first clipping other than the high-value effective time period; and on the basis of the retained high-value effective time period, generating and exporting a second clipped surgical video. The second clipped surgical video is a more elaborately edited surgical video which retains only the time period corresponding to the firing state of the surgical instrument, thereby further reducing the time of the video and making it convenient for the user to view the firing manipulation of utmost interest and highest importance. The method of the present invention allows the user to choose whether to export the first edited surgical video, which comprises relevant time periods of the surgical instrument, or to export the second edited surgical video, which is more elaborate, has a shorter duration, and only comprises the time period of the surgical instrument at a firing stage, thereby allowing more flexibility.
[0142] Further preferably, the described computing device capable of implementing the method of the present invention is further configured to allow the user to perform further manual editing on surgical videos of various stages, in particular an edited (first or second edited) surgical video. This provides the user with a higher degree of freedom to perform manipulations such as manual labeling, fine tuning, and clipping according to need, so as to generate a more personalized surgical video.
[0143] According to another aspect of the present invention, there is provided a computing device, the computing device being configured to be capable of the described method provided by the present invention. The computing device may be any possible device which is associated with a surgical manipulation in a surgical scene and which has functions such as receiving, processing, exporting, storing, forwarding and editing a video, e.g., a computer, etc., associated with a surgical operation device, an endoscope, etc.
[0144] According to yet another aspect of the present invention, there is provided a computer-readable medium stored in the computing device, the computer-readable medium being configured to be capable of implementing the described method provided by the present invention.
[0145] According to a further aspect of the present invention, there is provided a surgical system, the surgical system comprising: the described computing device of the present invention; one or more surgical instruments associated with the computing device; an image capture device located in a surgical scene and associated with the computing device, the image capture device being configured to be capable of capturing a raw surgical video; and a display device associated with the computing device, the display device being configured to be capable of displaying an edited surgical video. The above description of various embodiments of the present invention is provided to one of ordinary skill in the related art for the purpose of description. The present invention is not intended to be exclusive or limited to a single disclosed embodiment. In view of the foregoing, various alternatives and modifications of the present invention will be apparent to those of ordinary skill in the art. Thus, although some alternative embodiments have been specifically described, those of ordinary skill in the art will understand or relatively easily develop other embodiments. The present invention is intended to include all alternatives, modifications, and variations of the present invention described herein, as well as other embodiments that fall within the spirit and scope of the present invention described above.
Claims
CLAIMS1. A method implemented by a computing device for automatically editing a surgical video, characterized in that said method comprises: receiving a raw surgical video, said raw surgical video being configured to be captured by an image capture device which is associated with said computing device and which is located in a surgical scene; pre-processing the received raw surgical video; performing machine learning analysis on each video frame of a plurality of video frames of the pre-processed surgical video, so as to identify whether the surgical scene of each video frame is a first type of surgical scene or a second type of surgical scene; performing labeling on said surgical video at least once according to the result of the identification for each video frame of said plurality of video frames; determining retention or deletion of each time period of a plurality of time periods in said surgical video according to the result of said labeling; generating and exporting an edited surgical video on the basis of retained time periods; storing said edited surgical video; and displaying said edited surgical video on a display device associated with said computing device.
2. The method according to claim 1, wherein the step of performing labeling on said surgical video at least once comprises: grouping, in time sequence, the plurality of video frames of the surgical video into a plurality of video frame groups, each video frame group comprising a predetermined number of video frames; performing first labeling on each video frame group individually, according to the result of the identification for each video frame of said plurality of video frames;automatically dividing said surgical video into a plurality of time periods according to the result of the first labeling, each time period of said plurality of time periods comprising one or more video frame groups; performing second labeling on each time period of the plurality of time periods individually; and determining retention or deletion of each time period of the plurality of time periods according to the result of the second labeling.
3. The method according to claim 2, wherein said first labeling is performed on each video frame group of said plurality of video frame groups individually by means of a voting algorithm, and after said first labeling, the surgical video is automatically divided into time periods by means of a tolerance value algorithm according to the result of the first labeling obtained by means of the voting algorithm, and said second labeling is performed on the time periods.
4. The method according to claim 3, wherein the step of performing the first labeling on each video frame group of said plurality of video frame groups individually by means of a voting algorithm comprises: according to the ratio of the number of video frames identified as said first type of surgical scene in each video frame group of said plurality of video frame groups to the total number of video frames in the video frame group, performing labeling on each video frame group by means of the voting algorithm.
5. The method according to claim 4, wherein the step of performing labeling on each video frame group by means of the voting algorithm comprises: when said ratio of the number of video frames identified as said first type of surgical scene in each video frame group to the total number of video frames in the video frame group is not lower than a predetermined ratio threshold value, labeling the video frame group as a critical frame group;when said ratio of the number of video frames identified as said first type of surgical scene in each video frame group to the total number of video frames in the video frame group is lower than the predetermined ratio threshold value, labeling the video frame group as a non-critical frame group.
6. The method according to claim 5, wherein the steps of automatically dividing into time periods by means of a tolerance value algorithm according to the result of the first labeling and performing said second labeling comprise: according to whether each video frame group of the plurality of video frame groups is a critical frame group or a non-critical frame group, as labeled by means of the voting algorithm, and in combination with a consideration of a predetermined tolerance value upper-limit threshold value Vmax and a predetermined tolerance value lower-limit threshold value Vmin, determining the tolerance value of the video frame group; and according to the determined tolerance value of each video frame group of said plurality of video frame groups, automatically dividing the surgical video into a plurality of time periods, and performing the second labeling on each time period.
7. The method according to claim 6, wherein the step of performing labeling on each time period by means of a tolerance value algorithm comprises: starting from an initial time Tnof the first video frame group in each time period, sequentially and recurrently judging, at an initial time Tn+i of each video frame group, whether the video frame group is a critical frame group or a non-critical frame group; determining the tolerance value Vn+i of the video frame group according to the result of the judgment: when, at the initial time Tn+i of the video frame group, the video frame group is judged to be a critical frame group, comparing the sum Vn+i-i+a of the tolerance value Vn+i-i of the previous video frame group and a predetermined first numerical value a with the predetermined tolerance value upper-limit threshold value Vmax, and on the basis of the result of the comparison, determining the tolerance value of the video frame group:if Vn+i-i+a < Vmax, let the tolerance value of the video frame group Vn+i = Vn+i-i+a; and if Vn+i-i+a > Vmax, let the tolerance value of the video frame group Vn+i-l Vmax, when, at the initial time Tn+i of the video frame group, the video frame group is judged to be a non-critical frame group, comparing the value obtained by subtracting a predetermined second numerical value b from the tolerance value Vn+i- i of the previous video frame group with the predetermined tolerance value lower- limit threshold value and on the basis of the result of the comparison,determining the tolerance value of the video frame group: if Vn+i-i-b > Vmin, let the tolerance value of the video frame group Vn+i = Vn+i-i-b; and if Vn+i-i-b < Vmin, let the tolerance value of the video frame group Vn+i Vmin, only when the tolerance value Vn+i of the video frame group is determined to be Vmin, terminating the recurrence and performing division, to form the time period Tnto Tn+i.
8. The method according to claim 7, wherein the step of performing labeling on each time period by means of a tolerance value algorithm further comprises: after terminating the recurrence and performing division to form the time period Tnto Tn+i, determining whether the time period comprises one or more video frame groups having a tolerance value of Vmin, so as to perform the second labeling on the time period: if the time period comprises one or more video frame groups having a tolerance value of Vmin, labeling the time period as a low-value time period; and if the time period does not comprise one or more video frame groups having a tolerance value of Vmin, labeling the time period as a high-value time period.
9. The method according to claim 8, wherein said method comprises first clipping, said first clipping comprising:retaining said high-value time period and deleting said low-value time period; and based on the retained high-value time period, generating and exporting a first clipped surgical video.
10. The method according to claim 9, wherein the step of identifying whether the surgical scene of the video frame is a first type surgical scene or a second type surgical scene comprises target detection, to identify whether said surgical scene is a scene inside or outside the body of a patient, whether a surgical instrument appears in said surgical scene, and the position and / or area at which the surgical instrument appears, wherein, when the surgical scene being identified meets the following requirements, the surgical scene is identified as the first type of surgical scene: the surgical scene itself is identified as being located inside the body of a patient, and in addition, a surgical instrument appears in the surgical scene, and the surgical instrument is located inside the body of the patient; and when the surgical scene being identified involves one or more of the following situations, the surgical scene is identified as the second type of surgical scene: the surgical scene itself is identified as being located outside the body of the patient, no surgical instrument appears in the surgical scene, and / or the surgical instrument that appears in the surgical scene is located outside the body of the patient.
11. The method according to claim 10, wherein, when the surgical scene of the video frame is identified as said first type of surgical scene, said method further comprises: identifying the type of the surgical instrument that appears in said surgical scene; and generating an instrument type tag, the instrument type tag labeling the corresponding video frame in said surgical video.
12. The method according to claim 11, wherein said method further comprises generating one or more video subsets from the clipped surgical video according to the type of the surgical instrument, each video subset comprising an aggregation of the video frames labeled with the instrument type tag for the same type of surgical instrument.
13. The method according to claim 10, wherein the step of performing labeling on said surgical video at least once further comprises third labeling, wherein the third labeling comprises: determining a manipulation state of the surgical instrument in the surgical scene by comparing each video frame of the plurality of video frames in the high-value time period retained in said first clipping with the previous video frame; generating a firing tag when the manipulation state of said surgical instrument is determined to be a firing state, the firing tag labeling a corresponding video frame in the high-value time period retained in said first clipping; and labeling the firing tag-bearing time period in the high-value time period retained in said first clipping as a high-value effective time period.
14. The method according to claim 13, wherein said method further comprises second clipping, said second clipping comprising: retaining said high-value effective time period, and deleting the time periods in the high-value time period retained in said first clipping other than said high-value effective time period; and on the basis of the retained high-value effective time period, generating and exporting a second clipped surgical video.
15. The method according to claim 13, wherein the firing state of said surgical instrument is predetermined according to the type of said surgical instrument.
16. The method according to claim 15, wherein when said surgical instrument is an ultrasonic scalpel, an anastomat, and / or a hemostatic clamp, said firing state is identified in view of a switching action between an open state and a closed state of an end effector of said surgical instrument; and when said surgical instrument is an electrical knife, said firing state is identified in view of smoke generation.
17. The method according to claim 13, wherein said computing device is configured to allow a user to click on said firing tag and / or instrument type tag in the edited surgical video so as to jump to the location of the video frame in said surgical video that corresponds to said firing tag and / or instrument type tag.
18. The method according to claim 1, wherein said computing device is configured to allow a user to perform further editing on said edited surgical video.
19. The method according to claim 1, wherein the step of pre-processing the received raw surgical video comprises: normalizing the resolution of said raw surgical video to a predetermined resolution; and performing frame extraction processing on the surgical video after the resolution normalization; and normalizing the size of each video frame of the plurality of video frames of the surgical video after the frame extraction processing to a predetermined size.
20. A computing device, characterized in that said computing device is configured to be capable of implementing the method of any one of claims 1-19.
21. A computer-readable medium stored in a computing device, characterized in that the computer-readable medium is configured to be capable of implementing the method of any one of claims 1-19.
22. A surgical system, characterized in that said surgical system comprises: the computing device according to claim 20; one or more surgical instruments associated with the computing device;an image capture device located in a surgical scene and associated with said computing device, said image capture device being configured to be capable of capturing a raw surgical video; and a display device associated with said computing device, said display device being configured to be capable of displaying an edited surgical video.
Citation Information
Patent Citations
Timeline overlay on surgical video
US20200237452A1