Annotating medical videos using object detection and activity estimation
Patent Information
- Application Number
- JP2024503688
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-07-20
- Filing Date
- 2022-07-11
- Publication Date
- 2025-07-18
AI Technical Summary
Existing methods for annotating surgical videos are tedious, repetitive, and prone to errors, particularly for long or routine procedures, leading to information loss and inefficiency.
An apparatus and method utilizing a computer system with machine learning algorithms to automatically annotate surgical videos based on commands, instrument settings, and video parameters, including optical settings, surgical commands, and object detection, with timestamps for accurate and efficient annotation.
Reduces user fatigue and annotation errors by providing accurate, efficient, and informative annotations that support educational and legal documentation of surgical procedures.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Each example relates to a method and apparatus for recording a surgical procedure. [Background technology]
[0002] Medical surgeries are often recorded. For example, surgeries may be filmed as a legal requirement in many countries and as educational material for medical students around the world. Video recordings of surgeries can be generated by devices including computer systems. Summary of the Invention [Means for solving the problem]
[0003] It would be desirable to improve the techniques for determining annotations and / or adding annotations to videos of surgical procedures. Automating annotation can reduce user fatigue from a tedious and repetitive process and can help reduce errors in annotating videos.
[0004] Disclosed herein is an apparatus configured to annotate a video of a surgical procedure based on at least one of commands performed during the surgical procedure, instrument settings, or the video. The apparatus may include a processor and / or memory. The method of annotating may reduce errors in annotation and / or reduce the burden on a user to manually create annotations.
[0005] The command may be for determining / setting at least one of a plurality of settings of a surgical instrument, such as a surgical optics instrument. The surgical optics instrument may be configured for a procedure performed on a human eye. Enabling a faster procedure may reduce trauma to the patient. Enabling annotation based on a command may increase the accuracy of the determined annotation. Basing annotation on a command may allow annotations to be determined with minimal or no user intervention. Having accurate annotations may support the educational value of the annotated video and / or reduce the risk that legal obligations of procedure documentation will not be met.
[0006] The command may be for setting at least one of a plurality of settings of a surgical device, such as a surgical microscope, and / or an optical instrument. By allowing a command for a setting of a surgical device, such as a surgical microscope, and / or an optical instrument to be included in the basis for determining the annotation, the accuracy of the annotation may be increased. The command may be for setting an optical setting.
[0007] The optical settings may include at least one of the following: illumination intensity, focus, magnification, illumination source, color filters (e.g., color filter selection). Alternatively / additionally, the command may be a command to start recording, which may also form at least a partial basis for annotation. By allowing the optical settings to be included in the basis for determining annotation, the accuracy of annotation may be increased. It is desirable to have a "base set" of commands / settings that can assist in accurate annotation of videos, especially when machine learning algorithms are used.
[0008] The command may be at least one of a plurality of surgical commands of a surgical instrument, such as a surgical microscope. By allowing the surgical command to be included in the basis for determining the annotation, the accuracy of the annotation may be increased.
[0009] The surgical commands may be for ultrasound on, ultrasound off, pump on, pump off, infusion and / or suction. By allowing specific surgical commands to be included in the basis for determining the annotation, the accuracy of the annotation may be increased.
[0010] The determination of the annotation may be based on at least one of an image parameter of the video, an object in the video, a relative position of at least two objects in the video, or a motion in the video. By allowing certain parameters, objects, relative positions and / or motions to be included in the basis for determining the annotation, the accuracy of the annotation may be increased.
[0011] The object may be one or more identifiable objects, such as at least one surgical instrument (such as a blade or phacoemulsification machine) or at least one anatomical structure. By allowing specific objects and / or movements to be included in the basis for determining the annotation, the accuracy of the annotation may be increased.
[0012] The apparatus can determine a timestamp associated with the annotation, which can improve a user experience in editing the annotation and / or in searching for content within the video.
[0013] The apparatus can include a memory capable of storing a library from which annotations are selected. The annotation library can standardize annotations, which can be useful for satisfying legal requirements for procedure documentation and / or for providing annotations that are broadly understood.
[0014] A video can be annotated with two or more annotations, each with an associated timestamp. Multiple annotations can help accurately describe a multi-step surgical procedure.
[0015] Machine learning algorithms can be used to determine the annotations. The machine learning algorithms can provide accurate annotations and / or reduce the burden on the user to create / edit annotations. The machine learning algorithms can also adapt and incorporate new information. The machine learning can increase accuracy and / or provide more relevant / more informative annotations.
[0016] The machine learning algorithm can identify at least one of a command, an image parameter, an object, or a motion, or can classify at least one of a command, an image parameter, an object, or a motion. Algorithms that can identify and / or classify can increase the accuracy of annotations and / or provide more relevant / more useful annotations.
[0017] Disclosed herein is a method for annotating a video of a surgical procedure. The method includes determining annotations for a video of a surgical procedure and annotating the video with the annotations. Determining the annotations is based on at least one of determining at least one of commands performed during the procedure recorded by the video, equipment settings or image parameters of the video, objects in the video, or movements in the video. Using the commands, settings, image parameters, object identification, and / or movements can increase the accuracy of the annotations and / or provide more relevant / more informative annotations.
[0018] Disclosed herein is a computer program having a program code for annotating a video of a surgical procedure. The annotation can be based on at least one of commands performed during the surgical procedure, equipment parameters, or video. Using the commands, settings, equipment parameters, and / or video can increase the accuracy of the annotation and / or provide more relevant / more informative annotation.
[0019] Some examples of the present apparatus and / or methods will now be described, by way of example only, with reference to the accompanying drawings in which: [Brief description of the drawings]
[0020] [Figure 1] 1 is a schematic diagram of a system including a computer system. [Diagram 2] FIG. 1 illustrates a method for annotating a video. [Diagram 3] FIG. 1 illustrates a surgical workflow. [Figure 4] FIG. 2 illustrates an annotated video. [Diagram 5] FIG. 1 illustrates a surgical procedure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] Various examples will now be described more fully with reference to the accompanying drawings, in which several examples are shown and in which line thicknesses, layer thicknesses and / or area sizes may be exaggerated for clarity.
[0022] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ". For example, "perfusion / aspiration" can mean perfusion and / or aspiration. As used herein, the suffix "(s)" indicates an optional plural. For example, "frame(s)" can mean one or more frames of a video.
[0023] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, where a block or apparatus corresponds to a step or feature of a step, and similarly, aspects described in the context of a step also represent a description of a corresponding block or item or feature of a corresponding apparatus.
[0024] Annotating medical video clips can be useful to highlight moments / steps during surgery for both documentation and educational purposes.
[0025] Some surgeries follow routine procedures and may have standard steps. For example, in cataract surgery, port incision, second incision, administration of viscoelastic (e.g., Viscoat), serial annular capsulorhexis, flap creation, phacoemulsification, aspiration, irrigation, and intraocular lens insertion may be performed sequentially, and in some cases, at least some of the same steps may take approximately the same amount of time in all surgeries. Thus, annotation for this type of surgery may be similar for each procedure and thus may be repetitive if done manually. Disclosed herein are devices and methods for annotating videos of surgical procedures. The methods and devices described herein may make annotating videos of surgical procedures easier.
[0026] For example, some surgeries, such as neurosurgery, may take a long time to complete. Annotating a 10-hour operation may be tedious using prior art methods. For similar types of surgeries, the effort of having to review the video stream to highlight important steps may be tedious, especially if done manually. For medical video clips, long or short, a user may review the entire video, pause at steps of interest, and manually add annotations. In general, manually annotating a surgical video is time-consuming, repetitive, and tedious. It is prone to errors and may result in unwanted information loss in some cases.
[0027] In the case of short and / or standard medical procedures, such as cataract surgery, each step may be routine, and reviewing very similar video clips to highlight similar or even identical steps may become a repetitive task.
[0028] For relatively long medical procedures, such as neurosurgery, the videos tend to be in the range of 10 hours or more. Reviewing medical video clips having such lengths can be tedious. A user may overlook important steps while annotating, resulting in undesirable information loss. For example, the highlights in a procedure may only constitute about 1% of the duration of the procedure. Searching for specific sequences of video frames within a long video recording to manually annotate the highlights of the procedure can be tedious. Currently, annotating surgical videos can be highly inefficient.
[0029] Disclosed herein are methods, devices, systems, and microscopes that can reduce errors and / or tedious tasks in annotating surgical videos.
[0030] Some embodiments relate to a microscope including a system as described in connection with one or more of Figures 1 to 5. Alternatively, the microscope may be part of or connected to a system as described in connection with one or more of Figures 1 to 5.
[0031] FIG. 1 shows a schematic diagram of a system 100 configured to perform the methods described herein. The system 100 includes a microscope 110 and a computer system 120. The microscope 110 is configured to capture images and is connected to the computer system 120. The computer system 120 is configured to perform at least some of the methods described herein. The computer system 120 may be configured to execute machine learning algorithms. The computer system 120 and the microscope 110 may be separate entities, but may be integrated in one common housing. The computer system 120 may be part of a central processing system of the microscope 110 and / or the computer system 120 may be part of a subordinate part of the microscope 110, such as a sensor, actor, camera or lighting unit of the microscope 110.
[0032] The computer system may include a processor 101. The computer system may include a memory 107. The processor 101 and / or the memory 107 may be configured to operate the microscope and / or perform the methods described herein, particularly annotating videos of surgical procedures, such as in real-time during the recording of the surgery.
[0033] The computer system 120 may be a local computing device (e.g., a personal computer, laptop, tablet computer, or mobile phone) with one or more processors and one or more storage devices, or may be a distributed computing system (e.g., a cloud computing system with one or more processors and one or more storage devices distributed at various locations, such as local clients and / or one or more remote server farms and / or data centers). The computer system 120 may include any circuit or combination of circuits. In one embodiment, the computer system 120 may include one or more processors, which may be of any type. As used herein, a processor may contemplate any type of computing circuit, such as, but not limited to, a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a graphics processor, a digital signal processor (DSP), a multi-core processor, a field programmable gate array (FPGA), or any other type of processor or processing circuit, for example, of a microscope or a microscope component (e.g., a camera). Other types of circuits that may be included in computer system 120 may be custom circuits, application specific integrated circuits (ASICs), such as one or more circuits (such as communications circuits) used in wireless devices such as cell phones, tablet computers, laptop computers, two-way radios, and similar electronic systems. Computer system 120 may also include one or more storage devices, which may include one or more memory elements suitable for a particular application, such as main memory in the form of random access memory (RAM), one or more hard drives and / or one or more drives handling removable media, such as compact discs (CDs), flash memory cards, digital video discs (DVDs), and the like.Computer system 120 may also include a display device, one or more speakers, and a keyboard and / or controller which may include a mouse, a trackball, a touch screen, a voice recognition device, or any other device that enables a user of the system to input information to and receive information from computer system 120.
[0034] Some or all of the steps may be performed by (or using) a hardware apparatus, such as, for example, a processor, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, any one or more of the crucial steps may be performed by such an apparatus.
[0035] Depending on certain implementation requirements, the embodiments of the present invention can be implemented in hardware or software. The implementation can be performed by a non-transitory recording medium, such as a digital recording medium, for example a floppy disk, a DVD, a Blu-ray, a CD, a ROM, a PROM and EPROM, an EEPROM or a FLASH memory, on which electronically readable control signals are stored, which cooperate (or can cooperate) with a programmable computer system to implement the respective methods. Thus, the digital recording medium can be computer readable.
[0036] Some embodiments of the present invention include a data carrier having electronically readable control signals capable of cooperating with a programmable computer system to perform any of the methods described herein.
[0037] Generally, embodiments of the present invention can be implemented as a computer program product comprising program code which is operable to perform any of the methods when the computer program product is run on a computer, the program code may for example be stored on a machine readable carrier.
[0038] Another embodiment comprises the computer program for performing any of the methods described herein, stored on a machine readable carrier.
[0039] In other words, an embodiment of the present invention is, therefore, a computer program having a program code for performing any of the methods described herein, when the computer program runs on a computer.
[0040] Therefore, another embodiment of the present invention is a recording medium (or data carrier or computer readable medium) containing a computer program stored thereon for performing any of the methods described herein when executed by a processor. The data carrier, digital recording medium or recording medium is typically tangible and / or non-transitory. Another embodiment of the present invention is an apparatus as described herein, including a processor and a recording medium.
[0041] A further embodiment of the invention is therefore also a data stream or a sequence of signals representing the computer program for performing any of the methods described herein, the data stream or the sequence of signals being for example adapted to be transmitted via a data communication connection, for example the Internet.
[0042] Another embodiment comprises a processing means, for example a computer, or a programmable logic device configured to or adapted to perform any of the methods described herein.
[0043] Another embodiment comprises a computer having the computer program installed thereon for performing any of the methods described herein.
[0044] Another embodiment of the invention includes an apparatus or system configured to transfer (e.g., electronically or optically) a computer program for implementing any of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a storage device, etc. The apparatus or system may, for example, include a file server to transfer the computer program to the receiver.
[0045] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform any of the methods described herein. In general, the methods are advantageously performed by any hardware apparatus.
[0046] The system 100, microscope 110 and / or computer system 120 may be configured to annotate the video of the surgical procedure. The annotation may be based on commands performed during the surgical procedure, on equipment settings and / or on the video of the surgical procedure. As used herein, commands may be equipment commands, such as commands processed / executed by the computer system 120, microscope 110 and / or other surgical equipment, particularly during the surgical procedure. Alternatively / additionally, the annotation may be based at least in part on equipment settings.
[0047] The annotations may be determined at least in part from the video. For example, a frame and / or a sequence of frames of the video may be at least partially recognized and annotations may be associated with and / or added to the frame. For example, an object in the frame may be identified and the identification used as at least a partial basis for the annotation. Alternatively / additionally, dictation may be used to generate and / or edit the annotations. A language model trained using a medical corpus and manually annotated medical records may also be used to generate the annotations.
[0048] A command that may result in annotation may be, for example, a command to determine optical settings of a surgical optical instrument, such as the microscope 110. There may be more than one command to determine optical settings, such as illumination intensity, focus, and / or magnification. Alternatively / additionally, the settings / commands may include a command to start recording, such as start capturing video. Commands used during surgery may alternatively / additionally be used to annotate the video. For example, the commands (e.g., surgical commands) may include at least one of: turn on ultrasound, turn off ultrasound, turn on a pump, infuse, or aspirate.
[0049] The annotations may alternatively / additionally be determined based on one or more of image parameters (such as light levels and / or average color), objects in the video (such as detection of a surgical instrument appearing in the video), and / or motion in the video (such as motion of an object such as an anatomical feature or a surgical instrument). The processor may be capable of identifying objects including at least one surgical instrument (e.g., a blade or phacoemulsification machine) or at least one anatomical structure. Particular types of blades, such as a paracentesis blade, a keratome, and / or a capsulotome, may be identifiable.
[0050] The apparatus may be configured to determine a timestamp associated with the annotation. The timestamp may be included in and / or associated with the annotation. An annotated video may include one or more annotations, each having an associated timestamp. When playing the annotated video, the timestamp may enable a user to reach the video content at the timestamp by selecting the associated annotation and / or timestamp.
[0051] The apparatus may include a memory that stores a library of annotations. The annotation or at least one of the annotations may be selected from the library.
[0052] In one embodiment, which may be combined with any other embodiment described herein, a machine learning algorithm may be used to determine the annotation. For example, the machine learning algorithm may identify at least one of a command, a setting, an image parameter, an object, or a motion. Alternatively / additionally, the machine learning algorithm may classify at least one of a command, an image parameter, an object, or a motion.
[0053] Computer system 120 may be viewed as one embodiment of an apparatus for annotating videos, as described herein. Computer system 120 may include means for machine learning, such as a processor configured to determine annotations, as described herein. Alternatively / additionally, the computer system may be configured to perform a method for annotating videos, as described herein.
[0054] 2 shows a method for annotating a video. To determine 250 annotations for a video, commands 210, device settings / device parameters 215 and / or a video 220 can be input. Once annotations are determined, the video is annotated 280 with the annotations. The method can be repeated, for example, at different times and / or different frames of the video.
[0055] The method 200 may include determining 250 annotations for the video. The determining 250 annotations may be based on, for example, commands 210 executed during the surgical procedure recorded by the video, equipment parameters / equipment settings, and / or the video 220 itself.
[0056] The commands 210 and / or the video 220 may be in the form of input to a machine learning algorithm that performs the determining annotations 250. Annotating the video 280 may take a variety of forms, such as embedding text in the video, adding text to accompany the video and / or adding text to superimpose / overlay at least one frame of the video. Alternatively / additionally, the annotations may be in the form of subtitles and / or captions. The video 220 may include images of identifiable / recognizable anatomical features and / or surgical equipment. Such identification / recognition may be performed by a machine learning algorithm.
[0057] As an example of determining annotations 250, the machine learning algorithm may determine a number of input characteristics. For example, the algorithm may detect at least one of lighting conditions and / or lighting commands, such as an LED for illumination being in use, a lens movement command, a lens magnification command, for example to detect magnification, a start video recording command, or the presence of a blade, such as a puncture blade, in the image. In one example, based on the determination of lighting conditions, lens position / movement, magnification, and the presence of a puncture blade, the algorithm may determine that the video is from a "port incision" step and annotate the video frame accordingly. Determining the annotation may be based on commands performed at or before the relevant frame of the video, optical settings, and / or identification of objects in the video. Determining the annotation may include identifying a step of the surgical procedure, such as a port incision step. Annotating 280 may be performed by selecting a corresponding annotation from a library, such as by selecting an annotation that identifies the video frame as corresponding to a step of the surgical procedure, such as a port incision step.
[0058] FIG. 3 illustrates a surgical workflow. The surgical workflow 300 may be routine and / or predictable. The surgery can be decomposed into several different steps, which may include several sub-steps. The surgery can be decomposed broadly to include several major steps, as in the example of FIG. 3. The major steps can be decomposed to provide more detailed steps (e.g., sub-steps) of the surgical procedure. FIG. 3 illustrates three steps as an example. In step 1, the top frame of FIG. 3, access to the cataract is made. This access can be performed by creating several small incisions. In step 2, the middle frame of FIG. 3, the cloudy lens can be crushed and / or removed. In step 3, a replacement can be performed. In step 3, a new intraocular lens can be implanted.
[0059] Step 1 is an example of how anatomical structures may appear in a frame of a video. Anatomical structures can be identified, for example, by machine learning algorithms, to help determine annotations. For example, a capsular bag opening can be recognized in a video frame. The shape of the capsular bag opening may be irregular.
[0060] Step 2 is an example of how an object, such as one or more surgical instruments, may appear in a frame of the video. The object is identifiable, for example, by a machine learning algorithm, to assist in determining the annotation. Alternatively / additionally, the surgical procedure may include the execution of a command, such as the execution of an ultrasound, which may be used to determine the annotation. Alternatively / additionally, the shattering of the cloudy lens may be determined in the surgical procedure by a combination of determining the activation of ultrasound (e.g., via a command) and image analysis that specifically recognizes the cloudy lens being exposed to ultrasound. Step 2 may illustrate that the combination of the command and the video / video analysis may be used as a basis for determining the annotation, for example, utilizing a machine learning algorithm. Step 3 may illustrate that an implant, such as a lens, in the video frame is recognized / identified as being at least part of the basis for determining the annotation.
[0061] In one example, the annotation can be determined by a selection from library 340. The library may include multiple selections 341, 342, 343 for use as annotations. For example, the annotation for step 1 is selection 341 in library 340. The annotation for step 2 can be selection 342 in the library. The annotation for step 3 can be selection 343 in the library. The annotations of selections 341, 342, and / or 343 may be editable by the user.
[0062] FIG. 4 illustrates an annotated video. The annotated video 400 of FIG. 4 illustrates a frame 410 of a video 400 having an associated annotation 420. The video 400 may include multiple annotations 431, 432, 433, 434, 435. These annotations 431, 432, 433, 434, 435 may be associated with timestamps. Each annotation 420, 431, 432, 433, 434, 435 may be optionally displayed with a video frame of the video 400 associated with the annotation and / or timestamp. For example, at time t, the video 400 has a frame 410 having an associated annotation 420. At time t+1, the video has a frame associated with annotation 431, which may be displayed at time t+1. Each annotation 432, 433, 434, 435 may have a respective associated timestamp t+2, t+3, t+4, and t+5, eg, the time at which each annotation may be displayed.
[0063] Annotation can be performed, for example, in real time, during video acquisition and / or during a surgical procedure. Time stamps and annotations can be determined and, in some cases, displayed in a menu format, as shown, for example, in table 430 of FIG. 4, which associates annotations 431-435 with times t+1-t+5, respectively. The annotated video can include data, such as metadata, to facilitate editing of the annotations, including, for example, time stamps and annotations. While viewing the video, a user can select an annotation and / or time stamp to skip to a frame of the video associated with the time stamp and / or annotation.
[0064] FIG. 5 illustrates a surgical procedure. The procedure is represented in FIG. 5 as a series of frames of video. A video recording of the surgical procedure may include multiple annotations 510-590, such as at different times in the recording. Each annotation 510-590 may be determined as described herein and included in the video 400. In the example of FIG. 5, a cataract surgery is depicted. Determining the annotations 510-590 may be based at least in part on information provided by a user, such as an identification of a type of surgery, e.g., cataract surgery.
[0065] The first annotation 510 can refer to a port incision, which can be determined based on at least one of a command for the light to be on, a command for focusing, a command for magnification, a command for starting recording, or the identification of a paracentesis 511 in the video (e.g., a frame or a series of frames).
[0066] As with any annotation that follows a preceding annotation (e.g., first annotation 510), the second annotation 520 can be determined at least in part based on the preceding annotation / step being followed by a step of the surgical procedure.
[0067] As in the cataract surgery example, the second annotation 520 can refer to a second incision, which can be determined at least in part based on the identification of an object in the video, such as a keratome 521.
[0068] As in the cataract surgery example of Figure 5, a third annotation 530 can be determined that a viscoelastic substance (e.g., Viscoat) is being administered. Determining the annotation can be based at least in part on pump commands and / or injection commands corresponding to, for example, the injection of a viscoelastic substance, as in annotation 530. Alternatively / additionally, the determination can be based at least in part on the injection of an anesthetic agent, which can be in the form of a pump command.
[0069] The fourth annotation 540 can be determined to be a continuous circular capsulotomy. The fourth annotation 540 can be based at least in part on detection of an object, such as a capsulotome 541, and / or a pump command, such as injection of an anesthetic agent.
[0070] The fifth annotation 550 can be determined to be flap creation. The determination of flap creation can be based at least in part on detection of an object, such as a flap 552 and / or a blade (e.g., a capsulotome 551) and / or detection of a movement, such as a movement to peel the flap 552.
[0071] A sixth annotation 560 can be determined to be phacoemulsification. The determination of flap creation can be based at least in part on determining commands for detection, sonication and / or motion detection of an object, such as a phacoemulsification device 561.
[0072] The seventh annotation 570 can be determined to be aspiration. The determination to be aspiration can be based on detection of an object, such as a phacoemulsification device 571, aspiration commands and / or aspiration movements.
[0073] The eighth annotation 580 is identifiable as irrigation / aspiration. The determination of irrigation / aspiration may be based at least in part on detection of an object, such as an aspiration handpiece 581, detection of a movement, such as a cleaning movement, and / or a lighting command, particularly a command such as red reflex lighting.
[0074] A ninth annotation 590 can be determined to be an intraocular lens insertion. The determination of an intraocular lens insertion can be based at least in part on detection of an object, such as a lens injector 591, and / or detection of a movement, such as adjusting the position of a lens.
[0075] In one embodiment, which may be combined with any other embodiment described herein, the relative positions and / or relative motion of two or more objects identified in a video may be a partial or larger basis for determining annotations.
[0076] FIG. 5 also illustrates annotation insertion 585. For example, a user may insert an annotation into the video. Alternatively / additionally, a machine learning algorithm may determine or suggest annotation insertion 585. For example, the procedure may include offline activity such as lens preparation. Offline activity may be identifiable, for example, by periods of inactivity in the video and / or pauses in recording. Alternatively / additionally, offline activity may be identifiable based at least in part on the absence of illumination, the absence of illumination commands, lack of detected motion in the video and / or movement of surgical equipment, such as movement of a surgical microscope upward from the surgical site. Annotation insertion 585 may be determined to be lens preparation, for example, based at least in part on nearby annotations 580, 590, which may be known to have a correlation with lens preparation occurring during.
[0077] During surgery, objects such as puncture blades and viscoelastic materials may appear. Identification and / or classification of objects may be combined with determined motions associated with particular surgical steps performed by the user to suggest / determine annotations. For example, when annotations are determined by machine learning algorithms, these annotations may be subsequently edited by the surgeon / user, for example to verify accuracy. Activities occurring during surgery, such as flap peeling during the flap creation step in standard cataract surgery, are often surgical highlights, and it is particularly desirable for such highlights to have associated annotations. Convolutional neural networks specifically trained for medical object detection and / or activity classification (e.g., motion detection) may be combined with language models, which may be optionally trained based on medical corpora, to determine / suggest annotations for surgical videos.
[0078] System commands issued during surgery, such as when a user adjusts the system configuration, may be at least a partial basis for determining annotations and / or suggest steps in the surgical process. In combination with a trained neural network, more detailed annotations may be determined that may potentially document medical-surgical activities and / or configuration changes of the system.
[0079] As the user performs the surgery, system commands can be issued during the process, such as controlling the light, zoom factor, etc. Such system commands can be captured to generate information regarding detailed system configuration during a particular medical step. The system commands and / or system states can be included in annotations, such as multiple alternative annotations, that are accessible while viewing the video if the viewer is interested in the system settings during the procedure. Medical objects, blades, viscoelastic materials, etc. that may appear in the video during the procedure can also be included in the annotations, such as identified and highlighted as an overlay in the frame, such as by a trained convolutional neural network. Activities performed by the user, such as incision or peeling of a flap, can likewise be identified and classified, such as by using machine learning algorithms, to correlate movements captured in the video with appropriate annotations, such as annotations describing steps of the surgical procedure.
[0080] Further, it is possible to use language models trained using medical corpora and / or manually annotated medical records to determine annotations such as descriptions of objects appearing in frames of a video and / or descriptions of activities captured in a surgical video.
[0081] For example, the system configuration changes, the highlighted objects appearing in the video, and the identified activities (such as the movements / steps of the surgical procedure) can be combined to generate the annotations. The annotations may appear directly in the video, such as highlighting the actions / objects seen in the frames as subtitles, such as in an overlay / superimposed format. The annotations may include supplemental material associated with the video, such as equipment parameters, settings, and / or equipment commands. Such annotations can greatly assist in describing the surgical procedure, for example, for educational purposes. For example, the annotations can be embedded in the video stream as stop points and / or timestamps. A user can jump to a time-stamped frame of the video, for example, to quickly locate and review the desired content. Alternatively / additionally, the annotations may be in the form of subtitles, such as srt (subRip), ssa (Substation Alpha), and / or ass (Advanced Substation Alpha). Alternatively / additionally, the annotations may be embedded in the video / video container (e.g., MP4, MKV, etc.).
[0082] The embodiments may be based on the use of machine learning models or algorithms. Instead of relying on models and inferences, machine learning may refer to algorithms and statistical models that a computer system may use to perform a particular task without using explicit instructions. For example, machine learning may use data transformations inferred from analysis of past data and / or training data instead of rule-based data transformations. For example, image content may be analyzed using a machine learning model or using a machine learning algorithm. For the machine learning model to analyze image content, the machine learning model may be trained with training images as input and training content information as output. By training the machine learning model with a large number of training images and / or training sequences (e.g., words or sentences) and associated training content information (e.g., labels or annotations), the machine learning model "learns" to recognize image content, such that image content not included in the training data becomes recognizable using the machine learning model. The same principle may be used for other types of sensor data in a similar manner: by training the machine learning model with training sensor data and a desired output, the machine learning model "learns" a transformation between sensor data and output, which can be used to provide an output based on the non-training sensor data provided to the machine learning model. The provided data (e.g., sensor data, metadata and / or image data) may be pre-processed to obtain feature vectors that are used as input to a machine learning model.
[0083] The machine learning model may be trained using training input data. The above example uses a training method called "supervised learning". In supervised learning, the machine learning model is trained using multiple training samples, where each sample may include multiple input data values and multiple desired output values, i.e., each training sample is associated with a desired output value. By specifying both the training samples and the desired output value, the machine learning model "learns" during training which output value to provide based on input samples that are similar to the provided sample. Besides supervised learning, semi-supervised learning may be used. In semi-supervised learning, some of the training samples lack a corresponding desired output value. Supervised learning may be based on supervised learning algorithms (e.g. classification algorithms, regression algorithms or similarity learning algorithms). Classification algorithms may be used when the output is restricted to a limited set of values (categorical variables), i.e., the input is classified into one of a limited set of values. Regression algorithms may be used when the output may have any numerical value (within a range). Similarity learning algorithms may be similar to both classification and regression algorithms, but are based on learning from examples with a similarity function that measures how similar or related two objects are. In addition to supervised or semi-supervised learning, unsupervised learning may be used to train machine learning models. In unsupervised learning, input data may be (only) provided, and unsupervised learning algorithms may be used to find structure in the input data (e.g., by grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data containing multiple input values into multiple subsets (clusters), such that input values in the same cluster are similar according to one or more (predefined) similarity criteria, but are not similar to input values contained in another cluster.
[0084] Reinforcement learning is a third group of machine learning algorithms. In other words, reinforcement learning may be used to train machine learning models. In reinforcement learning, one or more software actors (referred to as "software agents") are trained to take actions in their surroundings. Based on the actions taken, rewards are calculated. Reinforcement learning is based on training one or more software agents to select actions that result in an increasing cumulative reward (as manifested by an increasing reward), resulting in the software agent becoming better at a given task.
[0085] Furthermore, some techniques may be applied to parts of the machine learning algorithm. For example, feature representation learning may be used. In other words, the machine learning model may be trained at least in part with feature representation learning, and / or the machine learning algorithm may include a feature representation learning component. A feature representation learning algorithm, which may be referred to as a representation learning algorithm, may not only preserve information in its input, but may also transform the information to make it useful, often as a pre-processing step before performing classification or prediction. Feature representation learning may be based on, for example, principal component analysis or cluster analysis.
[0086] In some examples, anomaly detection (i.e., outlier detection) may be used, which aims to provide identification of input values that raise suspicion by differing significantly from the majority of the input or training data. In other words, a machine learning model may be trained at least in part with anomaly detection and / or a machine learning algorithm may include an anomaly detection component.
[0087] In some examples, the machine learning algorithm may use a decision tree as a predictive model. In other words, the machine learning model may be based on a decision tree. In a decision tree, an observation about an item (e.g., a set of input values) may be represented by a branch of the decision tree, and an output value corresponding to this item may be represented by a leaf of the decision tree. The decision tree may support both discrete and continuous values as output values. If discrete values are used, the decision tree may be represented as a classification tree, and if continuous values are used, the decision tree may be represented as a regression tree.
[0088] Association rules are another technique that may be used in machine learning algorithms. In other words, a machine learning model may be based on one or more association rules. Association rules are created by identifying relationships between variables in large amounts of data. A machine learning algorithm may identify and / or utilize one or more association rules that represent knowledge derived from the data. These rules may be used, for example, to store, manipulate, or apply the knowledge.
[0089] Machine learning algorithms are typically based on machine learning models. In other words, the term "machine learning algorithm" may refer to a set of instructions that can be used to create, train, or use a machine learning model. The term "machine learning model" may refer to a set of data structures and / or rules that represent learned knowledge (e.g., based on training performed by a machine learning algorithm). In embodiments, the use of machine learning algorithm may refer to the use of an underlying machine learning model (or underlying machine learning models). The use of machine learning model may refer to the machine learning model and / or the set of data structures / rules that are the machine learning model being trained by a machine learning algorithm.
[0090] For example, the machine learning model may be an artificial neural network (ANN). An ANN is a system influenced by biological neural networks, such as those found in the retina or the brain. An ANN contains multiple interconnected nodes and multiple junctions, so-called edges, between the nodes. Typically, there are three types of nodes: input nodes that receive input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node may represent an artificial neuron. Each edge may convey information from one node to another. The output of a node may be defined as a (non-linear) function of its inputs (e.g. the sum of its inputs). The inputs of a node may be used in a function based on the "weights" of the edges or nodes that provide the inputs. The weights of the nodes and / or edges may be adjusted during the learning process. In other words, training an artificial neural network may involve adjusting the weights of the nodes and / or edges of the artificial neural network to obtain a desired output for a given input.
[0091] Alternatively, the machine learning model may be a support vector machine, a random forest model, or a gradient boosting model. A support vector machine (i.e., a support vector network) is a supervised learning model with an associated learning algorithm that may be used to analyze data (e.g., in classification or regression analysis). A support vector machine may be trained by providing input with multiple training input values that belong to one of two categories. A support vector machine may be trained to assign new input values to one of two categories. Alternatively, the machine learning model may be a Bayesian network, which is a probabilistic directed acyclic graphical model. A Bayesian network may represent a set of random variables and their conditional dependencies using a directed acyclic graph. Alternatively, the machine learning model may be based on a genetic algorithm, which is a heuristic method that mimics search algorithms and the process of natural selection.
[0092] It is specifically envisioned that the annotation method and apparatus disclosed herein can dynamically improve accuracy in some cases and dynamically reduce annotation error, such as when data for the machine learning algorithm is expanded. Training data may alternatively / additionally be used initially and / or added to a memory that can be accessed by the machine learning algorithm at a later time. Data, particularly additional data that may become available to the machine learning algorithm after the device is operational, can further expand the possible annotations (e.g., increase the library of possible annotations 340) and / or improve the accuracy of annotation determination / selection.
[0093] The examples described herein are for illustrative purposes only. The present invention is defined by the following claims and their equivalents. [Explanation of symbols]
[0094] System 100 Processor 101 Memory 107 Microscope 110 Computer Systems 120 How to Annotate 200 Command 210 Device parameters / device settings 215 Video 220 Annotation Decision 250 Annotating Videos 280 Surgical Workflow 300 Library 340 First Choice 341 Second Choice 342 Third Choice 343 Annotated Videos 400 Frame 410 Comments 420 Table of Notes 430 First Note 431 Second Note 432 Third Note 433 Fourth Note 434 Fifth Note 435 First comment 510 Puncture knife 511 Second Note 520 Corneal Scalpel 521 Third Note 530 Fourth Note 540 Capsulotomy knife 541 Fifth comment 550 Capsulotomy knife 551 Flap 552 Note 6: 560 Phacoemulsification device 561 Seventh Note 570 Phacoemulsification device 571 Commentary No. 8 580 Suction Handpiece 581 Annotation 585 Commentary No. 9 590 Lens injector 591
Claims
1. configured to annotate (280) a surgical procedure video (220) based on at least one of commands (210), device settings, or the video (220) executed during the surgical procedure device (120).
2. The command (210) is for setting at least one of a plurality of optical settings of an optical device The device (120) according to claim 1.
3. The plurality of optical settings includes at least one of illumination intensity, focus, illumination source, color filter, or magnification The device (120) according to claim 2.
4. The command (210) is at least one of a plurality of surgical commands of a surgical device (120) The device (120) according to claim 1.
5. The plurality of surgical commands is for at least two of ultrasonic activation, ultrasonic deactivation, pump activation, pump deactivation, injection, or suction The device (120) according to claim 4.
6. The device (120) is configured to determine the annotation based on at least one of image parameters of the video (220), an object (511) within the video, a relative position of at least two objects (551, 552) within the video, or movement within the video The device (120) according to claim 1.
7. The object is at least one of a plurality of distinguishable objects including at least one surgical device (511) or at least one anatomical structure (552) The device (120) according to claim 6.
8. The at least one surgical device is a scalpel (521) or a phacoemulsification and aspiration device (561) The device (120) according to claim 7.
9. The device (120) is configured to determine a timestamp (t + 1) associated with the annotation (431) The device (120) according to claim 1.
10. The device (120) further includes a memory (107) The annotation (420) is selected from a library (340) stored in the memory (107) The device (120) according to claim 1.
11. The device is configured such that the video (220) is annotated using a plurality of annotations (431, 432, 433) Each annotation has an associated timestamp (t+1, t+2, t+3), the apparatus (120) according to claim 1.
12. The apparatus (120) is configured to determine (250) the annotation (420) by means of a machine learning algorithm, the apparatus (120) according to claim 1.
13. The machine learning algorithm identifies at least one of the command (210), the image parameters of the video (220), an object or a movement, or classifies at least one of the command (210), the image parameters, the object or the movement, and is configured to perform at least one of the above, the apparatus (120) according to claim 12.
14. A method of annotating a video (220) of a surgical procedure, the method comprising: determining (250) an annotation for the video (220) of the surgical procedure, and annotating (280) the video (220) using the annotation wherein determining the annotation is based on at least one of a command (210), a device setting, executed during the procedure recorded by the video (220), or determination of at least one of the image parameters of the video (220), an object within the video (220), or a movement within the video, Method.
15. A computer program having program code for annotating a video (220) of a surgical procedure, wherein the annotation is based on at least one of a command (210), a device parameter or the video (220) executed during the surgical procedure, Computer program.