Automated annotation of endoscopic videos

An endoscopic system automatically annotates frames using clinician audio to enhance efficiency and reduce costs in generating training data for polyp detection, addressing the inefficiencies of manual annotation.

JP2026509761APending Publication Date: 2026-03-25GYRUS ACMI INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-13
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Conventional endoscopic video annotation is manual, time-consuming, and resource-intensive, particularly for tasks like polyp detection, due to noise and uncertainty in video quality, necessitating a more efficient method for generating training data.

Method used

An endoscopic system that automatically annotates frames using audio recordings of clinicians' spoken comments during procedures, integrating a camera, microphones, natural language processor, and controller to synchronize and annotate usable frames.

Benefits of technology

Enhances efficiency and reduces costs by automatically generating high-quality training data for machine learning algorithms, improving polyp detection and classification tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509761000001_ABST
    Figure 2026509761000001_ABST
Patent Text Reader

Abstract

A method for the automatic annotation of individual frames of a procedure video may include receiving a video stream captured by an endoscope camera during an endoscopic procedure via a controller's processing circuit. The video stream may include a first timestamp. The method may also include receiving an audio recording captured during an endoscopic procedure. The audio recording may include a second timestamp. The method may also include receiving transcribed text from the audio recording. The transcribed text may also include a second timestamp. The method may also include annotating the video stream with the transcribed text by associating the transcribed audio with the video stream when the first and second timestamps match.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Priority Claim This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 486,698, filed Feb. 24, 2023, which is hereby incorporated by reference in its entirety.

[0002] The present disclosure generally relates to endoscopes, and more particularly to the automatic annotation of individual frames of endoscopic video.

Background Art

[0003] Conventional endoscopes can be used in a variety of clinical procedures. For example, an endoscope can be used to elucidate, image, detect, and diagnose one or more medical conditions, provide fluid delivery (e.g., saline or other preparations via a fluid channel) to an anatomical area, provide a passage for one or more treatment devices for sampling or treating an anatomical area (e.g., via a working channel), provide a suction channel for collecting fluid (e.g., saline or other preparations), and the like. Such anatomical areas can include the gastrointestinal tract (e.g., esophagus, stomach, duodenum, pancreaticobiliary ducts, intestine, colon, and the like), renal regions (e.g., kidneys, ureters, bladder, urethra), other internal organs (e.g., reproductive organs, sinuses, submucosal regions, respiratory systems), and the like.

Summary of the Invention

Means for Solving the Problems

[0004] Various examples are illustrated in the figures of the accompanying drawings. Such examples are illustrative and are not intended to be an exhaustive or exclusive embodiment of the subject matter.

Brief Description of the Drawings

[0005] [Figure 1]This is a schematic diagram of an endoscopic system, as an example of the disclosure. [Figure 2] Figure 1 is a schematic diagram of an imaging and control system connected to an endoscope, illustrating an example of an imaging and control system according to this disclosure. [Figure 3] This is a block diagram of an example of a control unit for an endoscope system for automatic annotation of individual frames of endoscope video, according to an example of the present disclosure. [Figure 4] This is a flowchart illustrating a method according to an example of this disclosure. [Figure 5] This flowchart further illustrates the method shown in Figure 4, as an example of this disclosure. [Figure 6] This flowchart further illustrates the method shown in Figure 4, as an example of this disclosure. [Figure 7] This flowchart further illustrates the method shown in Figure 4, as an example of this disclosure. [Figure 8] This flowchart further illustrates the method shown in Figure 4, as an example of this disclosure. [Figure 9] This flowchart further illustrates the method shown in Figure 4, as an example of this disclosure. [Figure 10] This flowchart further illustrates the method shown in Figure 4, as an example of this disclosure. [Figure 11] This flowchart further illustrates the method shown in Figure 4, as an example of this disclosure. [Figure 12] This flowchart further illustrates the method shown in Figure 4, as an example of this disclosure. [Figure 13] This is a schematic diagram of an example of annotated images from a video stream captured during a medical procedure. [Figure 14] This is a block diagram illustrating an example of a machine in which one or more embodiments may be implemented. [Modes for carrying out the invention]

[0006] Endoscopic videos may contain a lot of noise and unusable frames caused by camera movement or splashes. For example, colonoscopy videos may contain bubbles from sprays or residual stool due to insufficient bowel preparation. During a colonoscopy, if polyps are detected, polypectomy may be performed, which can obscure the video stream due to medical instruments and blood from polypectomy. Because there is uncertainty in the quality of the video stream in colonoscopy videos, colonoscopy video frames are selected and annotated before they can serve as training data when training algorithms assist in tasks such as polyp detection or classification.

[0007] The selection and annotation of these images are manual processes that may be performed after the procedure. For example, generating training data may involve an endoscopist reviewing hours of recorded video and manually selecting a subset of usable frames corresponding to moments when the camera is stable and free from noise, debris, instruments, or similar. After manually selecting a subset of usable frames, the endoscopist can annotate the frames with clinical findings from the video captured during the colonoscopy. Manually annotating a subset of usable frames can be time-consuming and resource-intensive, and therefore very expensive. Well-annotated image data is crucial for the proper training of artificial intelligence systems that use machine learning algorithms to help endoscopists detect and classify abnormalities during the procedure. The larger the training dataset, the more likely the performance of the machine learning algorithms is to improve after training. Therefore, the inventors of this disclosure have found a need to increase the efficiency and reduce the cost associated with generating training data for use in medical image analysis.

[0008] This disclosure relates to an endoscopic system capable of automatically annotating endoscopic video. For example, this disclosure generally relates to a system capable of automatically identifying available frames during an examination and annotating the available frames with information extracted from procedure-performing audio spoken by the clinician while the clinician is viewing the images during the examination. In an examination, such as a colonoscopy, the performing clinician tends to speak aloud about the clinical findings or medical procedure regarding any abnormalities detected during the procedure. During a colonoscopy, the clinician may find a polyp or other abnormality, and the clinician tends to speak aloud about it (for example, to the team). In another example, sometimes during a colonoscopy, the clinician may perform a polypectomy. Here, the clinician typically speaks about the polyp and the removal of such a polyp. Finally, when the colon appears healthy, the clinician typically speaks about the colon appearing healthy while performing the colonoscopy. The utterances of clinicians who have completed a medical procedure typically ensure that the medical team performing the procedure is informed of the procedure's status and whether any intervention measures (e.g., polyp removal) should be taken. The inventors of this invention recognized that the utterances of clinicians performing medical procedures can contain rich clinical information. An example of this disclosure enhances efficiency in generating training data by extracting this rich information and then using the extracted information to automatically annotate images that the system has determined to be usable frames, thereby creating training data that can be used to train an algorithm to perform tasks such as polyp detection and classification.

[0009] In one example, the endoscopic system may include a camera connected to the distal end of an elongated member, microphones mounted around the endoscope in a position to capture sounds surrounding the medical procedure, a natural language processor configured to process the audio recordings captured by the microphones, and a controller configured to receive signals from the camera, microphones, and natural language processor for automatic annotation of the endoscopic video.

[0010] Figure 1 is a schematic diagram of an endoscopic system 10 which may include an imaging and control system 12 and an endoscope 14. System 10 is an exemplary example of an endoscopic system suitable for use with the systems, devices, and methods described herein, such as a colonoscopy system for automatic annotation of endoscopic video.

[0011] The endoscope 14 is insertable into an anatomical region for imaging, or can provide a passage or attachment (e.g., via tethering) to one or more sampling devices for biopsy or therapeutic devices for the treatment of a medical condition associated with the anatomical region. The endoscope 14 can interface with and connect to an imaging and control system 12. The endoscope 14 may also include a colonoscope, but other types of endoscopes may also be used with the features and teachings of this disclosure. The imaging and control system 12 may comprise a control unit 16, an output unit 18, an input unit 20, a light source unit 22, a fluid source 24, and a suction pump 26.

[0012] The imaging and control system 12 may have various ports for coupling with the endoscope system 10. For example, the control unit 16 may have data input / output ports for receiving data from and transmitting data to the endoscope 14. The light source unit 22 may include an output port for transmitting light to the endoscope 14, for example, via a fiber optic link. The fluid source 24 may include a port for supplying fluid to the endoscope 14. The fluid source 24 may include, for example, a pump and a fluid tank, or may be connected to an external tank, container, or storage unit. The suction pump 26 may include a port for drawing a vacuum from the endoscope 14 to generate suction, such as for drawing fluid out of the anatomical region into which the endoscope 14 is inserted. The output unit 18 and input unit 20 may be used by the operator of the endoscope system 10 to control the functions of the endoscope system 10 and to view the output of the endoscope 14. The control unit 16 may also generate signals or other outputs from the anatomical region into which the endoscope 14 is inserted. In some examples, the control unit 16 can generate electrical, acoustic, fluid, and similar outputs for treating anatomical regions using operations such as cauterization, cutting, freezing, and similar methods.

[0013] The endoscope 14 may include an insertion section 28, a functional section 30, and a handle section 32, which may be coupled to a cable section 34 and a coupler section 36. The insertion section 28 may extend distally from the handle section 32, and the cable section 34 may extend proximal to the handle section 32. The insertion section 28 may be elongated and include a bend section and a distal end to which the functional section 30 may be attached. The bend section may be controllable (e.g., by a control knob 38 on the handle section 32) to maneuver the distal end through a winding anatomical passage (e.g., stomach, duodenum, kidney, ureter, etc.). The insertion section 28 may also be elongated and include one or more working channels (e.g., lumens) that can support the insertion of one or more therapeutic instruments of the functional section 30, such as a cholangioscopy. The working channels may extend between the handle section 32 and the functional section 30. Additional functionality, such as fluid passages, guide wires, and pull wires, may also be provided by the insertion section 28 (for example, via suction or perfusion passages or similar).

[0014] The coupler section 36 is connected to the control unit 16, thereby allowing the endoscope 14 to be connected to several features of the control unit 16, such as the input unit 20, the light source unit 22, the fluid source 24, and the suction pump 26.

[0015] The handle section 32 may include a knob 38 and a port 40A. The knob 38 may be connected to a pull wire or other operating mechanism that can pass through the insertion section 28. Other ports, such as port 40A and even port 40B (Figure 2), may be configured to connect various electrical cables, guide wires, auxiliary scopes, tissue collection devices, fluid tubes, and the like to the handle section 32 for coupling with the insertion section 28, etc.

[0016] According to an example, the imaging and control system 12 can be provided on a mobile platform (e.g., cart 41) having shelves for housing a light source unit 22, a suction pump 26, an image processing unit 42 (FIG. 2), etc. Alternatively, some components of the imaging and control system 12 (shown in FIGS. 1 and 2) can be provided directly on the endoscope 14, thereby making the endoscope "self - contained".

[0017] The functional section 30 can comprise components for treating and diagnosing a patient's anatomical structure. The functional section 30 can include an imaging device, a lighting device, and an elevator. The functional section 30 can further include devices for collecting and removing optically enhanced biological materials and tissues, as described herein. For example, the functional section 30 can include one or more electrodes conductively connected to the handle section 32 and functionally connected to the imaging and control system 12, whereby biological materials contacting the electrodes can be analyzed based on comparative biological data stored in the imaging and control system 12.

[0018] FIG. 2 is a schematic diagram of the endoscope system 10 of FIG. 1 including the imaging and control system 12 and the endoscope 14. FIG. 2 schematically illustrates the components of the imaging and control system 12 coupled to the endoscope 14, which in the example illustrated includes a colonoscope. The imaging and control system 12 can include a control unit 16, which can include or be coupled to an image processing unit 42, a treatment generator 44, and a drive unit 46, as well as a light source unit 22, an input unit 20, and an output unit 18. The control unit 16 can include or communicate with an endoscope, a surgical instrument 48, and an endoscope system, which can include devices configured to engage tissue and collect and store a portion of that tissue, through which an imaging device (e.g., a camera) can view a target tissue by including optically enhanced materials and components. The control unit 16 can be configured to operate a camera to view a target tissue distal to the endoscope system. Similarly, the control unit 16 can be configured to operate the light source unit 22 to illuminate the surgical instrument 48, which can include selected components configured to reflect light in a particular manner, such as an enhanced tissue cutter having reflective particles.

[0019] The coupler section 36 is connected to the control unit 16, whereby it can be connected to the endoscope 14 for a plurality of features of the control unit 16, such as the image processing unit 42 and the treatment generator 44. In the example, the port 40A can be used to insert another surgical instrument 48 or device, such as a daughter scope or an auxiliary scope, into the endoscope 14. Such instruments and devices can be independently connected to the control unit 16 via a cable 47. In the example, the port 40B can be used to connect the coupler section 36 to various inputs and outputs, such as video, air, light, and electricity.

[0020] The image processing unit 42 and the light source unit 22 can each interface with the endoscope 14 (for example, in the functional section 30) by wired or wireless electrical connections. The imaging and control system 12 can accordingly illuminate an anatomical region, collect signals representing the anatomical region, process the signals representing the anatomical region, and display images representing the anatomical region on the display unit 18. The imaging and control system 12 may include a light source unit 22 for illuminating the anatomical region with light of a desired spectrum (e.g., broadband white light, narrowband imaging using preferred electromagnetic wavelengths, and so on). The imaging and control system 12 can be connected to the endoscope 14 (for example, via an endoscope connector) for signal transmission (e.g., light output from the light source, video signals from the imaging system at the distal end, diagnostic and sensor signals from diagnostic devices, and so on).

[0021] A fluid source 24 (shown in Figure 1) can communicate with the control unit 16 and may include one or more sources of air, saline, or other fluids, as well as associated fluid pathways (e.g., air channels, irrigation channels, suction channels, or similar) and connectors (barb fittings, fluid seals, valves, or similar). The fluid source 24 may be used as activation energy for the biasing or pressure-applying devices of this disclosure. The imaging and control system 12 may also include a drive unit 46 which may include an electric drive for advancing the distal section of the endoscope 14.

[0022] Figure 3 is a block diagram illustrating an example of a system 300 for the automatic annotation of individual frames of a colonoscopy video according to an example of the present disclosure. The system 300 may comprise an endoscope 302, a microphone 316, a natural language processor 320, a control system 322, and a memory 328. The endoscope 302 may include an elongated member 304, a control mechanism 310, and a camera 312. As best shown in Figure 2, the elongated member 304 (e.g., an insertion section 28 and a functional section 30 (Figures 1 and 2)) may extend from a proximal portion 306 to a distal portion 308. The elongated member 304 may be insertable into the patient's cavity.

[0023] A control mechanism 310 (for example, a knob 38 or a handle section 32 (both in Figures 1 and 2)) may be coupled to the proximal portion 306 of the elongated member 304. The control mechanism 310 may be configured to navigate the elongated member 304 during the procedure. In one example, the control mechanism 310 may be configured to be operated by a physician or other healthcare professional completing a medical procedure. In another example, the control mechanism 310 may be controlled by a robot or any other controller that can be used to help navigate the endoscope 302 within the patient's cavity.

[0024] Camera 312 may be mounted on the distal portion 308 of the elongated member 304. Camera 312 may be configured to capture a video stream 314 during a medical procedure. An image processing unit 42 (Figure 2) can process the video stream 314 and display it on a display unit 18 (Figures 1 and 2), so that a physician or other medical professional can see the front of the distal portion 308 of the elongated member 304 during a medical procedure. Camera 312 can also transmit the video stream 314 to multiple components simultaneously. For example, camera 312 can transmit the video stream 314 to a display unit to provide a live feed of the video stream 314 on a display for a physician, an image processing unit or control system 322 for processing, and on memory 328 for storing a raw version of the video stream 314. Any example of the video stream 314 may include a first timestamp 334 to help synchronize the video stream 314 with other signals of the system 300.

[0025] One or more microphones 316 are connected to the system 300, thereby enabling the capture of audio recordings 318. In one example, a microphone 316 may be attached to the endoscope 302. For example, a microphone 316 may be attached to the control mechanism 310. A microphone 316 may be attached to the handle section 32 (Figure 1), the knob 38 (also Figure 1), or any other location along the endoscope 302 where it can detect words spoken by the operator of the endoscope 302 during a medical procedure.

[0026] In another example, the microphone 316 may be attached to a part of the system 300 that is detached from the endoscope 302. For example, one or more microphones 316 may be attached to the bed or table on which the patient is lying during the procedure. One or more microphones 316 may be attached throughout the room, for example, to the wall or other fixtures.

[0027] In yet another example, the microphone 316 may be mounted on an imaging and control system (e.g., imaging and control system 12 (Figure 1)), a display or output device (e.g., output unit 18 (Figure 1)), an input device (e.g., input unit 20 (Figure 1)), or any other location on a medical cart (e.g., cart 41 (Figure 1)). System 300 may include one or more of the microphones 316 that communicate wirelessly with other components of System 300. For example, System 300 may include a wireless receiver configured to convert sound into electrical signals that can be transmitted to a natural language processor 320 or a control system 322 for processing.

[0028] The audio recording 318 may include spoken words, sounds, or any other noise that occurs around the system 300 during the procedure. The microphone 316 can transmit the audio recording 318 to the natural language processor 320, the control system 322, or any other component of the system 300 for analysis and editing. For example, the audio recording 318 may be transmitted to multiple components at once by the microphone 316. For example, the microphone 316 may simultaneously transmit the audio recording 318 to the natural language processor 320 or the control system 322 for processing and to memory 328 for storage. Any example of the audio recording 318 may include a second timestamp 338 to help synchronize the audio recording 318 with other signals around the system 300.

[0029] The natural language processor 320 may be configured to receive the audio recording 318, analyze the audio recording 318 using natural language processing techniques, and generate a transcribed audio recording 340. In one example, the natural language processor 320 may run live during an endoscopic procedure. When the natural language processor 320 is running during an endoscopic procedure, it may be delayed to some extent after the endoscopic procedure is performed so that the natural language processor 320 has data from the audio recording 318 when it is started. In another example, the natural language processor 320 may run offline. For example, the video stream 314 and the audio recording 318 may be sent to the natural language processor 320 after the endoscopic procedure is completed.

[0030] In one example, the natural language processor 320 can detect a single word from the speech recording 318. In another example, the natural language processor 320 can detect complete sentences, phrases, or paragraphs, which can be grouped together and stored in one or more text files.

[0031] The transcribed audio recording 340 may be a complete transcription of the audio recording 318. For example, the transcribed audio recording 340 may include all the recognized words found in the audio recording 318 by the natural language processor 320. In another example, the natural language processor 320 or the control system 322 may edit, rearrange, or otherwise modify the text from the natural language processor 320 to produce a more focused version of the transcribed audio recording 340. Variations of the portions of the audio recording 318 that may be used by the natural language processor 320 to create the transcribed audio recording 340 are described in further detail herein.

[0032] The control system 322 (for example, control unit 16) may be one or more controllers configured to operate the system 300. The memory 328 may contain instructions 330 that, when executed by the control system 322, can cause the processing circuitry of the control system 322 to perform an operation or complete a procedure. For example, the processing circuitry of the control system 322 may be configured by instruction 330 to receive a video stream 314 from the camera 312, receive an audio recording 318 from the microphone 316, receive a transcribed audio recording 340 from the natural language processor 320, and annotate one or more images of the video stream by completing a procedure such as that instructed by instruction 330 to annotate frames of endoscopic video. The control system 322 is described in more detail herein.

[0033] Instruction 330 can then cause the processing circuit of the control system 322 to complete a procedure or task. For example, instruction 330 can guide the control system 322 to annotate one or more images 324 of the video stream 314 with the transcribed text from the transcribed audio recording 340 by associating the transcribed audio recording 340 with the video stream 314 when a first timestamp 334 and a second timestamp 338 match. The first timestamp 334 and the second timestamp 338 can match if the first timestamp 334 and the second timestamp 338 are the same. In another example, there can be a range, for example, the first timestamp 334 and the second timestamp 338 may be considered to match when they are within each other's thresholds. One or more annotated images may include still images of the video stream 314 with annotated text from the transcribed audio recording 340 or audio recording 318. The instructions 330 and their interactions with the control system 322 will be described in more detail herein with reference to Figures 4 to 12.

[0034] Figure 4 is a flowchart illustrating Method 400 according to an example of the present disclosure. Method 400 can automatically annotate endoscopic video. As described above with reference to Figure 3, System 300 may be used to capture audio recordings, capture video recordings, and generate transcribed text files from the audio or video recordings while a medical procedure is being performed. In the example, the annotated images may be displayed on a display unit visible to the physician who has completed the medical procedure, overlaid on a video stream of the medical procedure, transmitted to a database, or stored in memory. Method 400 is described below with reference to Figures 4 to 12.

[0035] In step 410, method 400 may include receiving a video stream 314 captured by an endoscope camera (e.g., camera 312 in Figure 3) during the performance of an endoscopic procedure via a processing circuit of a controller (e.g., natural language processor 320 or control system 322 in Figure 3). For example, the video stream 314 may be a continuous feed transmitted from a camera on the endoscope. In another example, the video stream 314 may be one or more images that can be stitched together to form the video stream 314. Here, the video stream 314 can be sent to a video processor (e.g., image processing unit 42 (Figure 2)) to analyze the video stream 314 and generate one or more images of the video stream 314 that best capture the medical procedure. For example, one or more images may erase debris, blood, instruments, or other obstructions so that one or more images best represent the medical procedure. Each of the one or more images may have a first timestamp 334 so that the time of each of the one or more images can be determined after the performance of the procedure.

[0036] In step 420, method 400 may include receiving an audio recording 318 captured during the endoscopic procedure. The audio recording 318 may be one or more signals detected from one or more microphones placed around the operating room. For example, the audio recording 318 may be a single recording combining each signal detected from each microphone around the room. In another example, the audio recording 318 may be individual recordings of each recording from one or more microphones around the operating room. Nevertheless, each recording of the audio recording 318 may include a second timestamp 338. The control system 322 may receive the audio recording 318 and transmit the audio recording 318 to one or more components of system 300, for example, a natural language processor 320 or a memory 328 for storage.

[0037] In step 430, method 400 may include receiving a transcribed text of the audio recording 318 or a transcribed audio recording 340. In the example, the transcribed audio recording 340 may include a transcription from either of the audio recordings 318. A natural language processor (e.g., natural language processor 320 (Figure 3)), or any other language processor, may be connected to a system that transcribes speech to produce the transcribed audio recording 340. The transcribed audio recording 340 may also include a second timestamp 338.

[0038] In step 440, method 400 may include annotating the video stream 314 with the transcribed audio recording 340 by associating the transcribed audio recording 340 with the video stream when the first timestamp 334 and the second timestamp 338 match. For example, the control system 322 may overlay the transcribed audio recording 340 with the video stream or one or more images of the video stream such that the first timestamp 334 on the transcribed audio recording 340 matches the second timestamp 338 on the video stream 314.

[0039] Figure 5 is a flowchart illustrating the method 400 from Figure 4, as illustrated by an example of the present disclosure. In one example, step 440 of the method 400 from Figure 4 may optionally include steps 510-540, which can be performed on the processing circuitry of a natural language processor 320 or a control system 322 that processes a video stream (e.g., video stream 314) together with a first edited text file 342.

[0040] In step 510, method 400 may include converting the audio recording to a text file using natural language processing. For example, the control system 322 may send the audio recording to the natural language processor 320 to generate a first edited text file 342. The natural language processor 320 may use natural language processing to convert the audio recording 318 to a text file 344 (e.g., transcribed audio recording 340), and may send the text file 344 back to the control system 322 or save the text file 344 to memory 328 for further processing.

[0041] In step 520, method 400 may include determining a portion of the audio recording by identifying information about the patient by analyzing a text file. The natural language processor 320 or control system 322 may analyze the text file 344 to determine that a portion of the audio recording 318 contains identifying information about the patient. The identifying information can be any description of the patient that makes it easier to identify the patient. For example, the identifying information may include name, age, race or ethnicity, or any other factor that may be used to identify the patient. In one example, the natural language processor 320 or control system 322 may be configured to customize the words edited from the text file 344. For example, slang, jargon, or other non-technical terms that may affect the integrity of the training data of the annotated images may be edited from the text file 344.

[0042] In step 530, method 400 may include removing portions of the audio recording containing patient identification information from text file 344 to generate a first edited text file 342. For example, the natural language processor 320 or control system 322 may modify text file 344 by removing portions of the audio recording 318 containing patient identification information to generate the first edited text file 342. As described above, the natural language processor 320 or control system 322 may remove or edit any other words that the natural language processor 320 or control system 322 is configured to detect and edit. Thus, the first edited text file 342 can be a clean text file ready to be annotated at any time to generate training data. The control system 322 may save the first edited text file 342 separately from text file 344 so that both text file 344 and the first edited text file 342 can be processed later. Each of the first edited text files 342 and 344 may include a timestamp to help synchronize the text within the first edited text files 342 and 344 with other samples taken during a medical procedure.

[0043] In step 540, method 400 may include annotating the video stream 314 in the first edited text file 342 by associating the video stream 314 with the first edited text file 342 when the first timestamp and the second timestamp match. For example, the control system 322 may then annotate the video stream 314, or one or more images of the video stream 314, in the first edited text file 342 by associating the video stream 314 with the first edited text file 346 when the first timestamp 334 and the second timestamp 338 match.

[0044] Figure 6 is a flowchart further illustrating an additional optional operation performed as part of Method 400 from Figure 4, according to an example of the present disclosure. In one example, step 440 of Method 400 from Figure 4 may optionally include processing circuitry of a controller (e.g., natural language processor 320 or control system 322) annotating a video stream (e.g., video stream 314) with the associated speech-text file by completing steps 610-660.

[0045] In step 610, method 400 may include converting the speech recording into a text file using natural language processing. For example, method 400 may complete step 510 as described with reference to Figure 5.

[0046] In step 620, method 400 may include splicing a primary text file (for example, a transcribed audio recording 340 (Figure 3)) into two or more secondary text files 372. The two or more secondary text files 372 may be spliced ​​from the transcribed audio recording 340 based on timing, context, or any other indicators that can help decipher the speech of a medical professional performing a medical procedure.

[0047] In step 630, method 400 may include generating a relevance score 374 for each of the two or more secondary text files 372 by detecting keywords in each of the two or more secondary text files 372. The relevance score 374 can correspond to the relevance of each of the two or more secondary text files 372 according to pre-configured keywords. For example, the relevance score 374 may be configured to give a higher relevance to one of the two or more secondary text files 372 if one or more keywords are present, and a lower relevance to one of the two or more secondary text files 372 if one or more alternative keywords are present.

[0048] In step 640, method 400 may include classifying two or more secondary text files 372 into multiple classifications 376. Each of the multiple classifications 376 includes at least one of the two or more secondary text files 372 along with a corresponding relevance score 374. Here, two or more secondary text files 372 having similar relevance scores 374 may be combined into one of the multiple classifications 376. Sort into classifications in this way can help group or collect the most relevant portions of the two or more secondary text files 372. Alternatively, the control system 322 may help group or collect the least relevant portions of the two or more secondary text files 372 and group them into some classifications to exclude one or more of the two or more secondary text files 372 from being analyzed.

[0049] In step 650, method 400 may include removing one or more classifications from a primary text file (e.g., transcribed audio recording 340) whose corresponding relevance score 374 is below a threshold, in order to create a related text file 378. For example, a predetermined threshold may be selected to filter the most relevant portion of the text file. Removing classifications below this threshold can ensure the quality or relevance of the remaining classifications.

[0050] In step 660, method 400 may include annotating the video stream (e.g., video stream 314) with the associated text file 378 by associating the video stream with the associated text file 378 when the first timestamp and the second timestamp match. Here, the control system 322 can annotate only the most relevant images. Annotating only the most relevant images can reduce the computing time and resources required for annotation, and can also reduce the amount of storage required to store the annotated relevant images. Furthermore, annotating the relevant text according to each classification can result in a focused set of annotated figures. For example, the classifications could be classifications of polyp or anomaly types, processes or techniques performed during the procedure, tools used during the procedure, or similar. Thus, the focus of the classifications can further help to focus the input to machine learning to assist in detecting those instances, procedures, or anomalies using neural networks and artificial intelligence.

[0051] Figure 7 is a flowchart further illustrating additional optional operations that may be performed as part of Method 400 from Figure 4, according to an example of the present disclosure. In one example, step 440 of Method 400 from Figure 4 may optionally include processing the video stream with an audio profile 350 by including steps 710 to 740, and generating an audio profile annotation image 348, by a processing circuit of a controller (e.g., a natural language processor 320 or a control system 322).

[0052] For example, step 710 of method 400 may include accessing the voice profile 350 of the physician performing the endoscopic procedure by associating the voice of the physician performing the endoscopic procedure with a voice profile. The voice profile 350 is stored in memory 328 or any other memory of system 300 and can be compared with voices found on each voice recording to find the correct voice profile for the medical professional who has completed the medical procedure. A natural language processor 320 or control system 322 may access the voice profile 350 for the physician performing the endoscopic procedure by associating the voice profile 350 with the voice of the physician performing the endoscopic procedure.

[0053] In step 720, method 400 may include editing one or more voices from a voice recording that do not match voice profile 350 in order to create a voice profile voice recording 354. For example, a natural language processor 320 or a control system 322 may edit one or more voices from a voice recording 318 or a transcribed voice recording 340 that do not match voice profile 350 in order to create a voice profile voice recording 354.

[0054] A voice profile voice recording 354 may contain the voices of people that match one or more voice profiles 350. For example, voice profiles 350 may be maintained only for medical professionals who have appropriate qualifications (e.g., licensed physicians, nurses, medical assistants, or similar) to ensure that the captured words belong to qualified individuals. In another example, each person working around system 300 may have a unique version of voice profile 350, and voice profiles 350 may be tagged with restrictions or confidential information handling permissions as appropriate to match the qualifications of each person from whom voice profile 350 was generated. Thus, a voice profile voice recording 354 may contain tags, marks, or other labels corresponding to the medical licenses or qualifications of the voice profiles 350 contained therein.

[0055] In the example, the audio profile audio record 354 may be stored in memory along with the audio record, the raw transcribed audio record, and the video stream. The audio profile audio record 354 may also include a timestamp that can help the control system 322 synchronize 354 / / with the video stream 314 or one or more images of the video stream 314.

[0056] In step 730, method 400 may include converting the voice profile audio recording 354 into a voice profile text file 356. For example, a natural language processor 320 or a control system 322 may use natural language processing techniques to convert the voice profile audio recording 354 into a voice profile text file 356. Similar to the voice profile audio recording 354, the control system 322 may know one or more voice profiles 350 contained in the voice profile text file 356, which may include tags, marks, or labels corresponding to the medical licenses or qualification certificates of the voice profiles 350 contained therein. The voice profile text file 356 may be stored together with the voice profile audio recording 354, a video stream, or any other files from system 300. The voice profile text file 356 may also include a timestamp to help the control system 322 synchronize the voice profile text file 356 with other files from system 300.

[0057] In step 740, method 400 may include annotating the video stream with the audio profile text file 356 by associating the video stream 314 with the audio profile text file 356, or one or more images, when the first timestamp and the second timestamp match. For example, the natural language processor 320 or control system 322 may annotate one or more images 324 from the video stream 314 with the audio profile text file 356 by associating the video stream 314 with the audio profile text file 356 when the first timestamp 334 and the second timestamp 338 match, thereby generating an audio profile annotation image 348. The audio profile annotation image 348 may include a credential or confidential information handling authorization mark, label, or other instruction for each audio profile 350 contained therein, and may be stored in memory 328, either alone or together with other data of system 300, for future reference. The filtered nature of the voice profiles of 350 medical professionals, or individual physicians, can provide information-rich images that can help focus on reviewing one or more images or focus on input for machine learning.

[0058] Figure 8 is a flowchart further illustrating additional optional operations that may be performed as part of Method 400 from Figure 4, according to an example of the present disclosure. In one example, Method 400 may optionally include steps 810-870 for generating one or more labeled images 332.

[0059] In step 810, method 400 may include identifying that at least one anomaly was found during the endoscopic procedure by analyzing the transcribed text to detect one or more related words indicating at least one anomaly observed during the endoscopic procedure. For example, control system 322 may execute instructions 330 to generate one or more labeled images 332. For example, control system 322 may identify that at least one anomaly 390 was found during the endoscopic procedure by analyzing the transcribed text (e.g., transcribed audio recording 340) to detect one or more related words or keywords (e.g., one or more keywords 360) indicating at least one anomaly 390 observed during the endoscopic procedure.

[0060] In step 820, method 400 may include generating a unique identification label 392 for at least one anomaly 390. The unique identification label 392 may include a second timestamp (e.g., a second timestamp 338) indicating when one or more relevant words were spoken during the endoscopic procedure. The unique identification label 392 can identify the type of polyp and may be used to refer to that particular polyp in future scans or medical procedures. The unique identification label 392 may also be used to track the examination or pathological outcome of the polyp after it has been removed. In another example, the unique identification label 392 may be used to track changes in size, shape, color, texture, or other physical characteristics detected during the medical procedure of the identified polyp.

[0061] In step 830, method 400 may include obtaining one or more images 324 from the video stream 314, each containing a first timestamp 334 that may correspond to a second timestamp 338. In such an example, the second timestamp 338 may indicate when one or more relevant words were spoken during the endoscopic procedure. Thus, the identified polyp is likely to be found in one or more images at or around the time of its corresponding timestamp.

[0062] In step 840, method 400 may include an instruction 330 that configures the control system 322 to record the position of cursor 398 during the performance of a medical procedure. The position of cursor 398 may be the position of the cursor used by the operator of system 300 (e.g., a physician, nurse, or similar) to perform the medical procedure. For example, the position of cursor 398 may include a timestamp (e.g., a first timestamp 334 or a second timestamp 338). The position of cursor 398 may be stored in memory 328 and recalled later for processing or overlay.

[0063] In step 850, method 400 may include creating one or more labeled images 332 by labeling one or more images 324 with a unique identifier label 392, the placement of the cursor 398 in a first timestamp 334 corresponding to a second timestamp 338. For example, the placement of the cursor 398 can be annotated, overlaid, projected onto, or similarly operated on one or more images 324 by associating the placement of the cursor 398 in the first timestamp 334 with one or more images 324 in the second timestamp 338, thereby creating one or more labeled images 332. One or more labeled images 332 may include a unique identifier label 392 and the placement of the cursor 398, and can help direct a review by a physician reviewing the data or help focus machine learning during the execution of a machine learning process.

[0064] In step 860, method 400 may include replacing one or more images from video stream 314 with one or more labeled images 332. In one example, one or more labeled images 332 may then replace one or more images 324 from video stream 314 with one or more labeled images 332, where video stream 314 containing one or more images 324 may be projected onto a display in the operating room. In another example, the original stream of video stream 314 may be displayed on a first display, and video stream 314 containing one or more images 324 may be displayed on another display in the operating room.

[0065] In step 870, method 400 may include storing one or more labeled images 332 and a video stream in non-temporary machine-readable memory. A video stream 314 containing one or more images 324 may be stored separately from the video stream 314 in order to store the video stream 314 containing one or more images 324.

[0066] Figure 9 is a flowchart further illustrating additional optional actions that may be performed as part of Method 400 from Figure 4, according to an example of the present disclosure. In one example, Method 400 may optionally include steps 910–930.

[0067] In step 910, method 400 may include extracting one or more labeled images 332. For example, the control system 322 may extract one or more of the one or more labeled images 332 from the steps of method 400 as described in Figure 8. Thus, one or more labeled images 332 may be separated from one or more images or any unlabeled images of the video stream 314.

[0068] In step 920, method 400 may include storing one or more labeled images separated from the video stream in non-temporary computer-readable memory in order to create an anomaly record 368. For example, the control system 322 may store one or more labeled images 332 (Figure 8), separated from the video stream 314, in memory 328 in order to create an anomaly record 368. The anomaly record 368 may include a record for each anomaly having a unique identification label 392. For example, if several of the one or more labeled images 332 include an image of a single polyp, each of those one or more labeled images 332 having a corresponding unique identification label 392 may be stored together in the anomaly record 368. Thus, each of the unique identification labels 392 may include a unique anomaly record 368 that can be used to track changes, inspections, or other outcomes of the anomaly. Furthermore, anomaly records 368 corresponding to the corresponding group or subset of unique identification labels 392 can be used for further machine learning regarding specific types of polyps or groups of polyp types captured in each of the one or more labeled images 332.

[0069] In step 930, method 400 may include storing an anomaly dataset 386 in an anomaly record 368. The anomaly dataset 386 includes at least one of the following: image quality score 394, the instrument used to manipulate the anomaly 396, the location of the anomaly, or identification of the physician performing the endoscopic procedure 399. The anomaly dataset 386 may be used to determine best practices or potentially to suggest best practices to physicians performing future procedures that encounter one or more anomalies of similar quality.

[0070] The image quality score 394 may be configured to provide an image quality score that can be used to determine a confidence level or to filter out images that contain or are blurred. For example, a higher image quality score may indicate a sharp image with few or no obstructions. A lower image quality score may indicate blurring, the presence of obstructions, or a lack of focus or clarity in the image. In the example, the control system 322 or any other image processor may be configured to run an algorithm that can analyze and determine the image quality score 394.

[0071] The instruments used to manipulate anomaly 396 may be surgical scalpels, blades, suction devices, sutures, stitches, or any other type of instrument that can engage one or more of the anomalies within the body. For example, the instruments used to manipulate anomaly 396 may be captured to help suggest the instruments to a physician performing a future medical procedure when the corresponding anomaly is encountered.

[0072] Identifying the physician performing the endoscopic procedure 399 can be used to question healthcare providers. Furthermore, identifying the physician performing the endoscopic procedure 399 can be used to learn physician preferences so that the system 300 can learn the tools, procedures, or steps each physician prefers when encountering different anomalies. This understanding by the system 300 can help the system 300 recommend the surgical physician's preferred procedures, tools, or steps for future medical procedures. Identifying the physician performing the endoscopic procedure 399 can also help direct the review of anomalies after the examination is complete.

[0073] Figure 10 is a flowchart further illustrating additional optional operations that may be performed as part of Method 400 from Figure 4, according to an example of the present disclosure. In one example, steps 810–830 from Figure 8 of Method 400 may optionally include steps 1010–1040 to enable a natural language processor 320 or a control system 322 to identify the relevant image 358.

[0074] In step 1010, method 400 may include generating a distribution of keywords 352 found in the transcribed text. For example, instruction 330 may configure a processing circuit of a natural language processor 320 or a control system 322 to generate a distribution of keywords 352. Each keyword in the distribution of keywords 352 can be found in the transcribed text (for example, the transcribed audio recording 340).

[0075] In the example, the natural language processor 320 or control system 322 can generate a distribution of keywords 352 by counting the frequencies of one or more keywords 360. Other natural language processing techniques for sorting one or more keywords 360 may be used to generate the distribution of keywords 352. For example, relevance scores, confidence scores, or any other analysis may be completed by the natural language processor 320 or control system 322 to find the relevance of one or more keywords 360. One or more keywords 360 may include words that can indicate an anomaly found during the procedure. For example, one or more keywords 360 may include "polyp," "anomaly," "look here," "just there," any other words that can signal that an anomaly was encountered during the procedure, or similar.

[0076] In step 1020, method 400 may include assigning identifier 362 to one or more of the keywords 360. The natural language processor 320 or control system 322 may also assign identifier 362 to one or more of the keywords 360. Identifier 362 can indicate the type or style of one or more keywords 360 found in the audio recording or text file. For example, if a polyp is detected, identifier 362 may indicate that a polyp or other anomaly has been found.

[0077] In step 1030, the method 400 may include instructions that configure the processing circuit of the natural language processor 320 or the control system 322 to identify one or more associated images 358 by associating identifiers 362 of one or more keywords from the keyword 360 with one or more images 324 of the video stream 314 when a first timestamp 334 and a second timestamp 338 match to generate one or more identified images 364. By matching the first timestamp 334 and the second timestamp 338, the natural language processor 320 or the control system 322 may find one or more images that can contain a visual depiction of an anomaly detected from the physician's speech.

[0078] In step 1040, method 400 may include annotating one or more identified images 364 with one or more keywords 360 to create one or more identified and annotated images. For example, a natural language processor 320 or a control system 322 may annotate one or more identified images 364 with identifiers 362 to one or more keywords 360 to create one or more identified and annotated images 366. One or more identified and annotated images 366 may be displayed on a display in the operating room to assist in further analysis of abnormalities during the performance of a medical procedure. In another example, one or more identified and annotated images 366 may be stored in memory (e.g., memory 328) or in a file directory for later retrieval or analysis.

[0079] Figure 11 is a flowchart that further illustrates additional optional actions that may be performed as part of Method 400 from Figure 4, according to an example of the present disclosure. In one example, Method 400 from Figure 4 may optionally include steps 1110–1130.

[0080] In step 1110, method 400 may include transmitting one or more images, including annotations, from the video stream to the physician after the endoscopic procedure. To confirm the identity and location of the anomaly 370, the control system 322 may transmit one or more identified and annotated images 366 to the physician. For example, the control system 322 may transmit one or more identified and annotated images 366 to the physician via email, charting software, or any other physical or electronic means that allows the physician to analyze the identity and location of the anomaly 370 in one or more identified and annotated images 366. Such a review may be completed in the operating room or on any computer that can later communicate with system 300.

[0081] In step 1120, method 400 may include receiving confirmation from a physician of the identity and location of one or more anomalies on an image. The control system 322 may receive confirmation from a physician of the identity and location of one or more anomalies 370 on an image 324.

[0082] In step 1130, method 400 may include storing the identity and location of the anomaly, as well as one or more images, in a database. After the control system 322 receives confirmation of the identity and location of the anomaly 370, the control system 322 may send the images to a file directory used for training or machine learning. In another example, the control system 322 may store the confirmation in an anomaly dataset, anomaly records, or patient medical records.

[0083] Figure 12 is a flowchart that further illustrates additional optional actions that may be performed as part of Method 400 from Figure 4, according to an example of the present disclosure. In one example, Method 400 from Figure 4 may optionally include steps 1210–1220.

[0084] In step 1210, method 400 may include receiving one or more pathology results 380, each corresponding to a sample associated with an anomaly 382 from one or more images 324 from a video stream 314. The pathology results may provide information on whether the anomaly is diseased or for further diagnosis of the anomaly 382.

[0085] In step 1220, method 400 may include storing one or more pathology results in a database along with one or more corresponding images. For example, control system 322 may receive one or more pathology results 380. One or more pathology results 380 may correspond to a sample associated with an anomaly 382 from one or more images 324 from video stream 314. Control system 322 may then store one or more pathology results 380 in database 384 along with one or more corresponding images 324.

[0086] Figure 13 illustrates a schematic diagram of an example of an annotated image 1300. The annotated image 1300 may be, for example, any of the annotated images described herein and may include image 1310, annotation 1320, marking box 1330, polyp identification box 1340, and process identification box 1350.

[0087] Image 1310 may be an individual frame from a video stream captured by a camera during an endoscopic procedure. Image 1310 may be from a timestamp corresponding to a spoken keyword, or from other indicators of an anomaly found during the procedure. The controller may analyze Image 1310 to ensure that the clearest version of the video stream is used from the timestamp corresponding to the found anomaly. For example, Image 1310 may be an image from the video stream before or after the corresponding timestamp in which the anomaly was found, provided that the image can provide a clearer or better image of the found anomaly.

[0088] Annotation 1320 may be placed on image 1310, as shown in Figure 13. In another example, annotation 1320 may be shifted laterally to 1310, for example, in a polyp identification box 1340, a process identification box 1350, or in any area around image 1310. Annotation 1320 may be a spoken keyword or a unique identifier generated for an anomaly. Annotation 1320 can help identify the location of an anomaly encountered during the performance of a medical procedure.

[0089] The marking box 1330 can be overlaid on 1310 to help identify any anomalies found. For example, the marking box 1330 can help a physician reviewing annotated images to quickly find anomalies to improve anomaly review. In another example, the marking box 1330 can help a machine learning algorithm focus on anomalies to improve the quality of learning.

[0090] The polyp identification box 1340 may include information about anomalies from a medical procedure or from a physician's review after a medical procedure. For example, the polyp identification box 1340 may include annotations of utterances made by a medical professional before and after a timestamp of spoken keywords. In another example, the polyp identification box 1340 may include notes entered by a physician after the physician has reviewed the annotated image 1300. The information provided to the polyp identification box 1340 helps to improve machine learning by providing additional information about the annotated image 1300, which can help to improve the information provided for machine learning by sorting the annotated image 1300 into groups of similar findings.

[0091] The process identification box 1350 may include process information relating to a medical procedure. For example, the process identification box 1350 may include a timestamp of the video stream in which the image was captured, a timestamp when the keyword was recognized, identification of a confidence level or polynomial, and other processing information of the medical procedure that may be useful to know after the procedure is completed. The process identification box 1350 may also include manufacturing information or model number of the equipment used to perform the medical procedure.

[0092] The example of an annotated image 1300 shown in Figure 13 is merely one example of an annotated image 1300. This example, including the information presented therein, is not intended to limit the scope of the present invention in any way. Rather, the information provided is intended to be an example of an annotated image 1300 that the system described herein can generate.

[0093] Figure 14 illustrates a block diagram of an exemplary machine 1400 in which one or more of the techniques (e.g., methods) described herein may be performed. The example may include, or be operated by, logic or a number of components or mechanisms within the machine 1400 as described herein. A circuit (e.g., a processing circuit) is a collection of circuits implemented on a tangible entity of the machine 1400, including hardware (e.g., simple circuits, gates, logic, etc.). Member relationships between circuits may change flexibly over time. A circuit includes members that, individually or in combination, can perform a specified operation when operating. In one example, the hardware of a circuit may be designed immutably to perform a particular operation (e.g., hardwired). In one example, the hardware of a circuit may include variably connected physical components (e.g., execution units, transistors, simple circuits, etc.) that include a machine-readable medium (e.g., a movable arrangement of magnetic, electrical, or invariant mass particles, etc.) that is physically modified to encode instructions for a particular operation. When connecting physical components, the fundamental electrical properties of the hardware components are changed, for example, from insulator to conductor, or from conductor to insulator. Instructions allow embedded hardware (e.g., an execution unit or loading mechanism) to create members of a circuit in the hardware via variable connections in order to perform a specific part of an operation during operation. Thus, in one example, a machine-readable medium element is either part of a circuit or is coupled communicatively to other components of a circuit when the device is operating. In one example, any one of the physical components may be used by multiple members of multiple circuits. For example, during operation, an execution unit may be used at one time in a first circuit of a first set of circuits, and at a different time reused by a second circuit within the first set of circuits, or a third circuit in a second set of circuits. Additional examples of these components relating to machine 1400 are shown below.

[0094] In alternative embodiments, machine 1400 may operate as a standalone device or may be connected to other machines (e.g., network-connected). In a network-connected deployment, machine 1400 may operate as a server machine, a client machine, or both in a server-client network environment. For example, machine 1400 may operate as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Machine 1400 may be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, web appliance, network router, switch or bridge, or any machine capable of executing instructions (sequentially or otherwise) that specify actions to be performed by that machine. Furthermore, although only a single machine is illustrated, the term “machine” or “collection of machines” shall be interpreted to include machines or collections of machines that individually or in conjunction execute a set of instructions (or more sets of instructions) to perform one or more of the methods described herein, such as cloud computing, software as a service (SaaS), and other computer cluster configurations.

[0095] The machine (e.g., a computer system) 1400 includes a hardware processor 1402 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), main memory 1404, static memory (e.g., memory or storage for firmware, microcode, basic input / output (BIOS), Unified Extensible Firmware Interface (UEFI), etc.) 1406, and mass storage device 1408 (e.g., a hard drive, tape drive, flash storage, or other block device), some or all of which may communicate with each other via an interlink (e.g., a bus) 1430. The machine 1400 may further include a display unit 1410, an alphanumeric input device 1412 (e.g., a keyboard), and a user interface (UI) navigation device 1414 (e.g., a mouse). In one example, the display unit 1410, the input device 1412, and the UI navigation device 1414 may be touchscreen displays. The machine 1400 may further include a storage device (e.g., a drive unit) 1408, a signal generating device 1418 (e.g., a speaker), a network interface device 1420, and one or more sensors 1416 such as a Global Positioning System (GPS) sensor, compass, accelerometer, or other sensor. The machine 1400 may also include an output controller 1428 such as a serial (e.g., Universal Serial Bus (USB)), parallel, or other wired or wireless (e.g., infrared (IR), near-field communication (NFC)) connection for communicating with or controlling one or more peripheral devices (e.g., a printer, a card reader, etc.).

[0096] The registers of processor 1402, main memory 1404, static memory 1406, or mass storage device 1408 may be, or include, a machine-readable medium 1422 in which one or more sets of data structures or instructions 1424 (e.g., software) that are embodied or utilized by one or more of the techniques or functions described herein are stored. The instructions 1424 may also reside, fully or at least partially, in any of the registers of processor 1402, main memory 1404, static memory 1406, or mass storage device 1408 during their execution by machine 1400. In one example, one or any combination of the hardware processor 1402, main memory 1404, static memory 1406, or mass storage device 1408 may constitute the machine-readable medium 1422. Although the machine-readable medium 1422 is exemplified as a single medium, the term “machine-readable medium” may include a single or multiple mediums configured to store one or more instructions 1424 (for example, a centralized or distributed database, and / or associated caches and servers).

[0097] The term “machine-readable medium” can include any medium that can store, encode, or carry instructions for execution by machine 1400, which cause machine 1400 to perform one or more of the technologies of the present disclosure, or that can store, encode, or carry data structures used by or associated with such instructions. Examples of non-limiting machine-readable mediums may include solid memory, optical mediums, magnetic mediums, and signals (e.g., radio frequency signals, other photon-based signals, sound signals, etc.). In one example, a non-temporary machine-readable medium includes a machine-readable medium having a plurality of particles having an immutable (e.g., stationary) mass, and is therefore a composition of a substance. Thus, a non-temporary machine-readable medium is a machine-readable medium that does not contain a transient propagating signal. Specific examples of non-temporary machine-readable media may include non-volatile memory such as semiconductor memory devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0098] For example, information stored on machine-readable medium 1422 or otherwise provided in any way may represent instruction 1424 itself, or instruction 1424 in any format from which instruction 1424 can be derived. This format from which instruction 1424 can be derived may include source code, encoded instructions (e.g., in a compressed or encrypted form), packaged instructions (e.g., divided into multiple packages), or similar. Information representing instruction 1424 in machine-readable medium 1422 may be processed by a processing circuit into instructions for performing any of the operations described herein. For example, deriving instruction 1424 from information (e.g., processing by a processing circuit) may include compiling (e.g., from source code, object code, etc.), interpreting, loading, organizing (e.g., dynamically or statically linking), encoding, decrypting, encrypting, decrypting, packaging, unpackaging, or otherwise manipulating the information into instruction 1424.

[0099] For example, the derivation of instruction 1424 may involve assembling, compiling, or interpreting information (e.g., by a processing circuit) to create instruction 1424 from some intermediate or pre-processed format provided by machine-readable medium 1422. When the information is provided in multiple parts, it may be combined, decompressed, and modified to create instruction 1424. For example, the information may take the form of multiple compressed source code packages (or object code, or binary executable code, etc.) on one or more remote servers. The source code packages may be encrypted, decrypted, decompressed if necessary, assembled (e.g., linked), compiled or interpreted on the local machine (e.g., into a library, a standalone executable, etc.) and executed by the local machine as they are transferred over a network.

[0100] Instruction 1424 may also be transmitted or received over a communication network 1426 using a transmission medium via a network interface device 1420 that utilizes any one of a number of transport protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.). Illustrative communication networks may include, but are not limited to, local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), LoRa / LoRaWAN, or satellite communication networks, mobile phone networks (e.g., cellular networks such as those conforming to 3G, 4G LTE / LTE-A, or 5G standards), legacy telephone service (POTS) networks, and wireless data networks (e.g., the IEEE 702.11 family of standards, the IEEE 702.15.4 family of standards, and peer-to-peer (P2P) networks, as well as Wi-Fi®). In one example, the network interface device 1420 may include one or more plug jacks (e.g., Ethernet, coaxial, or telephone plug jacks) or one or more antennas for connecting to the communication network 1426. In one example, the network interface device 1420 may include multiple antennas for wireless communication using at least one of the following technologies: single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO). The term “transmission medium” is to be interpreted as including any intangible medium on which instructions for execution by machine 1400 can be stored, encoded, or carried, including digital or analog communication signals, or other intangible medium for facilitating such software communication. The transmission medium is a machine-readable medium.

[0101] The following are non-limiting examples, but in particular, specific aspects of the subject matter will be described in detail to address the problems described herein and to provide advantages.

[0102] Example 1 is a method for the automatic annotation of individual frames of a procedure video, the method comprising: receiving a video stream captured by an endoscope camera during an endoscopic procedure, the video stream including a first timestamp; receiving an audio recording captured during an endoscopic procedure, the audio recording including a second timestamp; receiving text transcribed from the audio recording, the transcribed text including a second timestamp; and annotating the video stream with the transcribed text by associating the transcribed audio with the video stream when the first timestamp and the second timestamp match.

[0103] In Example 2, the subject matter of Example 1 is expanded to include converting an audio recording into a text file using natural language processing.

[0104] In Example 3, the subject of Example 2 is annotated by the controller's processing circuit processing the video stream together with a first edited text file, which includes converting the audio recording to a text file using natural language processing, determining that a portion of the audio recording contains patient identification information by analyzing the text file, removing the portion of the audio recording containing patient identification information from the text file to generate a first edited text file, and annotating the video stream with the first edited text file by associating the first edited text file with the video stream when the first timestamp and the second timestamp match.

[0105] In Example 4, the subject matter of Examples 1-3 is annotated by the controller's processing circuit processing the video stream together with the audio profile text file, accessing the audio profile for the physician performing the endoscopic procedure by associating the audio profile with the voice of the physician performing the endoscopic procedure, editing one or more voices from the voice recording that do not match the audio profile in order to create an audio profile audio recording, converting the audio profile audio recording into an audio profile text file, and annotating the video stream with the audio profile text file by associating the audio profile text file with the video stream when a first timestamp and a second timestamp match.

[0106] In Example 5, the subject matter of Examples 2-4 is annotated, and the controller processing circuit processes the video stream together with the associated audio text files by splicing the primary text file into two or more secondary text files and generating a relevance score for each of the two or more secondary text files by detecting keywords on each of the two or more secondary text files.

[0107] In Example 6, the subject of Example 5 is annotated by the controller's processing circuit processing a video stream together with associated audio text files, by classifying two or more secondary text files into multiple categories, where each of the multiple categories includes at least one of two or more secondary text files having a corresponding relevance score; by removing one or more of the multiple categories whose corresponding relevance score is below a threshold from the primary text file to create associated text files; and by annotating the video stream with the associated text files by associating the associated text files with the video stream when the first timestamp and the second timestamp match.

[0108] In Example 7, the subject matter of Examples 1 to 6 is carried out by the controller's processing circuit generating one or more labeled images, which includes identifying that at least one abnormality was found during the endoscopic procedure by analyzing transcribed text to detect one or more related words indicating at least one abnormality observed during the endoscopic procedure, and generating a unique identification label for the at least one abnormality, wherein the unique identification label includes a second timestamp indicating when one or more related words were spoken during the endoscopic procedure, and obtaining one or more images from a video stream containing a first timestamp corresponding to the second timestamp.

[0109] In Example 8, the subject of Example 7 is to record the cursor placement in one or more images at a first timestamp corresponding to a second timestamp, where the cursor placement in one or more images indicates the pointer placement operated by the physician during the endoscopic procedure.

[0110] In Example 9, the subject of Example 8 is to create one or more labeled images by labeling one or more images with a unique identification label at the cursor placement in the first timestamp corresponding to the second timestamp.

[0111] In Example 10, the subject of Example 9 is replaced with one or more images from a video stream with one or more labeled images, and the one or more labeled images and the video stream are stored in non-temporary machine-readable memory.

[0112] In Example 11, the subject matter of Examples 9-10 is further expanded to include extracting one or more labeled images, storing one or more labeled images separated from the video stream in non-temporary machine-readable memory to create an anomaly record, and storing an anomaly dataset in the anomaly record, wherein the anomaly dataset includes at least one of the following: image quality score, tools used to manipulate the anomaly, location of the anomaly, or identification of the physician performing the endoscopic procedure.

[0113] Example 12 includes generating a distribution of keywords found in transcribed text, where the subject of Examples 7-11 is to count the frequency of one or more keywords; assigning identifiers to one or more of the keywords; identifying one or more images by associating one or more identifiers of those keywords with one or more images in a video stream when a first timestamp and a second timestamp match; and creating one or more identified and annotated images by annotating the identified one or more images with identifiers to one or more of these keywords.

[0114] In Example 13, the subject matter of Examples 1 to 12 is extended to include transmitting one or more images, including annotations, from a video stream to a physician after an endoscopic procedure, receiving confirmation from the physician of the identity and location of an anomaly on one or more images, and storing the identity and location of the anomaly, as well as one or more images, in a database.

[0115] In Example 14, the subject matter of Examples 1 to 13 is receiving one or more pathology results, the one or more pathology results corresponding to samples related to abnormalities from one or more images from a video stream, and storing the one or more pathology results together with the corresponding one or more images in a database.

[0116] Embodiment 15 is a system for the automatic annotation of individual frames of a procedure video, the system comprising an endoscope having an elongated member including a distal portion, the elongated member having a camera attached to the distal portion, the camera capturing a video stream during the procedure, the video stream including a first timestamp; a microphone configured to capture an audio recording of ambient sounds during the procedure, the audio recording including a second timestamp; a natural language processor configured to receive the audio recording and the transcribed audio recording, the transcribed audio recording including a second timestamp; a memory including instructions; and a controller including a processing circuit, the processing circuit configured, when in operation, to receive a video stream from the camera, receive an audio recording from the microphone, receive transcribed text from the natural language processor, and annotate the video stream with the transcribed text by associating the transcribed audio with the video stream when the first timestamp and the second timestamp match.

[0117] In Example 16, the subject of Example 15 is carried out by a controller processing circuit processing a video stream together with a first edited text file in order to annotate the video stream, by converting the audio recording to a text file using natural language processing, determining that a portion of the audio recording contains patient identification information by analyzing the text file, removing the portion of the audio recording containing patient identification information from the text file to generate a first edited text file, and annotating the video stream with the first edited text file by associating the first edited text file with the video stream when the first timestamp and the second timestamp match.

[0118] In Example 17, the subject matter of Examples 15-16 is carried out by the controller's processing circuit processing the video stream together with an audio profile text file in order to annotate the video stream, by accessing the audio profile for the physician performing the endoscopic procedure by associating the audio profile with the voice of the physician performing the endoscopic procedure, editing one or more voices from the voice recording that do not match the audio profile in order to create an audio profile voice recording, converting the audio profile voice recording into an audio profile text file, and annotating the video stream with the audio profile text file by associating the audio profile text file with the video stream when a first timestamp and a second timestamp match.

[0119] In Example 18, the subject matter of Examples 15-17 is reproduced by the controller's processing circuit generating one or more labeled images for annotating a video stream, by analyzing transcribed text to detect one or more related words indicating at least one anomaly observed during the endoscopic procedure, thereby identifying that at least one anomaly was found during the endoscopic procedure, and by generating a unique identification label for the at least one anomaly, wherein the unique identification label includes a second timestamp indicating when the one or more related words were spoken during the endoscopic procedure, and by obtaining one or more images from the video stream that include a first timestamp corresponding to the second timestamp.

[0120] In Example 19, the subject of Example 18 is configured such that the controller's processing circuit records, by command, the cursor placement in one or more images at a first timestamp corresponding to a second timestamp, where the cursor placement in one or more images indicates the placement of a pointer operated by a physician during an endoscopic procedure, and labels one or more images with a unique identification label at the cursor placement at the first timestamp corresponding to the second timestamp, thereby creating one or more labeled images.

[0121] In Example 20, the subject matter of Examples 15-19 is described by a controller processing circuit that, by instruction, generates a distribution of keywords found in the transcribed text, the keyword distribution is configured to count and generate the frequencies of one or more keywords, assign identifiers to one or more of the keywords, identify one or more images by associating one or more identifiers of those keywords with one or more images in the video stream when a first timestamp and a second timestamp match, and create one or more identified and annotated images by annotating the identified one or more images with identifiers to one or more of these keywords.

[0122] Example 21 is at least one machine-readable medium containing instructions that, when executed by a processing circuit, cause the processing circuit to perform an operation to implement any of Examples 1 to 20.

[0123] Example 22 is an apparatus that includes means for implementing any of Examples 1 to 20.

[0124] Example 23 is a system for implementing any of Examples 1 to 20.

[0125] Example 24 is a method for implementing any of Examples 1 to 20.

[0126] The description detailed above includes references to the accompanying drawings, which form part of the detailed description. The drawings illustrate specific embodiments that may be carried out. These embodiments are also referred to herein as “Examples.” Such embodiments may include elements in addition to those illustrated or described. However, the inventors also intend examples in which only those elements illustrated or described are provided. Furthermore, the inventors also intend examples in which any combination or permutation of those elements illustrated or described (or one or more aspects thereof) is used with respect to a particular embodiment (or one or more aspects thereof) or with respect to other embodiments (or one or more aspects thereof) illustrated or described herein.

[0127] All publications, patents, and patent documents referenced herein are incorporated herein by reference in their entirety, as if they were incorporated individually by reference. In the event of any conflict between the usage herein and those documents incorporated by reference, the usage of the incorporated documents shall be considered supplementary to the usage herein, and in the event of an incompatible conflict, the usage herein shall prevail.

[0128] In this specification, the term "one" includes one or more, regardless of any other examples or uses of "at least one" or "one or more," as is common in patent literature. In this specification, the term "or" is used to refer to non-exclusive "or" such that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise specified. In the appended claims, "including" and "in which" in the English original are used as plain English paraphrases of the phrases "comprising" and "wherein," respectively. Also, in the following claims, "including" and "comprising" in the English original are open-ended; that is, a system, device, article, or process that includes elements in addition to those listed after such words in the claim is still considered to fall within the scope of that claim. Furthermore, in the following claims, terms such as "first," "second," and "third" are used merely as labels and are not intended to impose numerical requirements on the subject.

[0129] As used herein, the term “approximately” means roughly, within a range, about, or around that range. When the term “approximately” is used with a numerical range, it modifies that range by extending the boundary above and below the specified numerical value. Generally, the term “approximately” is used herein to modify a numerical value by a 10% variance above and below the stated number. In one embodiment, the term “approximately” means plus or minus 10% of the numerical value of the number in which it is used. Thus, approximately 50% means within the range of 45% to 55%. Numerical ranges described by endpoints herein include all numbers and fractions that fall within that range (for example, 1 to 5 includes 1, 1.5, 2, 2.75, 3, 3.90, 4, 4.24, and 5). Similarly, numerical ranges described herein by endpoints include subranges that fall within that range (for example, 1 to 5 includes 1 to 1.5, 1.5 to 2, 2 to 2.75, 2.75 to 3, 3 to 3.90, 3.90 to 4, 4 to 4.24, 4.24 to 5, 2 to 5, 3 to 5, 1 to 4, and 2 to 4). It should also be understood that all numbers and fractions are presumed to be modified by the word “approximately”.

[0130] The above description is intended to be illustrative and not restrictive. For example, the embodiments (or one or more aspects thereof) described above may be used in combination with each other. Other embodiments may be used by those skilled in the art who have reviewed the above description. The abstract is submitted with the understanding that it is intended to allow the reader to quickly confirm the nature of the technical disclosure and is not to be used to interpret or limit the claims or their meaning. Also, in the forms for carrying out the above invention, various features may be grouped together in order to streamline the disclosure. This should not be interpreted as meaning that the disclosed features not claimed are essential to any claim. Rather, the subject matter of the invention may be found in fewer features than all the features of a particular disclosed embodiment. Accordingly, the following claims are incorporated into forms for carrying out the invention, such that each claim stands on its own as a distinct embodiment. The scope of the embodiments should be determined with reference to the supplementary claims, along with the entire scope of the equivalents that are the subject of the supplementary claims. [Explanation of symbols]

[0131] 10 Endoscopy Systems 12. Imaging and control systems 14 Endoscopy 16 Control Unit 18 Output Units 20 Input Units 22 Light source units 24 Fluid source 26 Suction pump 28 Insertion Section 30 Functional Sections 32 Handle section 34 Cable Sections 36 Coupler Sections 38 Control knob 40A Port 40B Port 41 Cart 42 Image Processing Unit 44 Treatment Generator 46 Drive Unit 47 Cables 48 Surgical instruments 300 Systems 302 Endoscope 304 Slender member 306 Proximal portion 308 Distal portion 310 Control mechanism 312 Camera 314 video streams 316 Microphone 318 Audio Recordings 320 Natural Language Processors 322 Control System 328 memory 330 command 332 Labeled Images 334 First timestamp 338 Second timestamp 340 audio recordings 342 First edited text file 348 Voice Profile Annotation Images 350 voice profiles 352 keywords 354 Voice Profiles Voice Recording 356 Voice Profile Text Files 358 related images 360 Keywords 362 Identifiers 364 identified images 366 identified and annotated images 368 Anomaly Record 370 Abnormal 372 Secondary text files 374 relevance score 376 Classification 378 Related text files 382 Abnormality 384 Databases 386 Anomaly Datasets 390 Abnormal 392 unique identification labels 394 Image Quality Score 396 Abnormality 398 Cursor 399 Endoscopic Procedures 400 ways 1300 annotated images 1310 images 1320 annotations 1330 Marking Box 1340 Polyp Identification Box 1350 Process Identification Box 1400 machines 1402 Hardware Processors 1404 Main Memory 1408 Storage Devices 1410 Display Unit 1412 Alphanumeric input device 1414 User Interface (UI) Navigation Devices 1416 Sensor 1418 Signal Generating Devices 1420 Network Interface Device 1422 Machine-readable media 1424 Instructions 1426 Communication Network 1428 Output Controller 1430 Interlink

Claims

1. A method for automatic annotation of individual frames of a procedure video, A controller processing circuit receives a video stream captured by an endoscope camera during an endoscopic procedure, wherein the video stream includes a first timestamp. A step of receiving an audio recording captured during the endoscopic procedure, wherein the audio recording includes a second timestamp. A step of receiving text transcribed from the audio recording, wherein the transcribed text includes the second timestamp, The steps include: annotating the video stream with the transcribed text by associating the transcribed audio with the video stream when the first timestamp and the second timestamp match; A method that includes this.

2. The method according to claim 1, further comprising the step of converting the audio recording into a text file using natural language processing.

3. The annotation step involves the processing circuit of the controller processing the video stream together with the first edited text file. Converting the audio recording into a text file using natural language processing, By analyzing the aforementioned text file, it is determined that a portion of the audio recording contains identifying information about the patient. The first edited text file is generated by deleting a portion of the audio recording containing identification information about the patient from the aforementioned text file. When the first timestamp and the second timestamp match, the first edited text file is associated with the video stream, thereby annotating the video stream with the first edited text file. The method according to claim 2, comprising the step of being carried out by

4. The annotation step involves the processing circuit of the controller processing the video stream together with the audio profile text file. Accessing the voice profile for the physician performing the endoscopic procedure by associating the voice profile with the voice of the physician performing the endoscopic procedure, Editing one or more audio clips from the audio recording that do not match the audio profile in order to create an audio profile audio recording, Converting the aforementioned voice profile audio recording into the aforementioned voice profile text file, Annotating the video stream with the audio profile text file by associating the audio profile text file with the video stream when the first timestamp and the second timestamp match. The method according to claim 1, comprising the step of being carried out by

5. The annotation step involves the processing circuit of the controller processing the video stream together with the associated audio text file. The process of splicing a primary text file into two or more secondary text files, To generate a relevance score for each of the two or more secondary text files by detecting keywords in each of the two or more secondary text files, The method according to claim 2, comprising the step of being carried out by

6. The annotation step involves the processing circuit of the controller processing the video stream together with the associated audio text file. Classifying the two or more secondary text files into multiple categories, wherein each of the multiple categories includes at least one of the two or more secondary text files having a corresponding relevance score. The process involves removing one or more of the aforementioned classifications whose corresponding relevance score is below a threshold from the primary text file, and creating a related text file. Annotating the video stream in the associated text file by associating the associated text file with the video stream when the first timestamp and the second timestamp match. The method according to claim 5, comprising the step of being carried out by

7. The processing circuit of the controller generates one or more labeled images. By analyzing the transcribed text, one or more related words indicating at least one abnormality observed during the endoscopic procedure are detected, thereby identifying that at least one abnormality was found during the endoscopic procedure. To generate a unique identification label for the at least one abnormality, wherein the unique identification label includes the second timestamp indicating when one or more related words were spoken during the endoscopic procedure, To obtain one or more images from the video stream, which includes the first timestamp corresponding to the second timestamp; The method according to claim 1, performed by...

8. The method according to claim 7, comprising the step of recording the placement of a cursor in one or more images in the first timestamp corresponding to the second timestamp, wherein the placement of the cursor in one or more images indicates the placement of a pointer operated by a physician during the endoscopic procedure.

9. The method according to claim 8, further comprising the step of creating one or more labeled images by labeling one or more images with the unique identification label in the arrangement of the cursor in the first timestamp corresponding to the second timestamp.

10. The steps of replacing one or more images from the video stream with one or more labeled images, The steps include storing the one or more labeled images and the video stream in non-temporary machine-readable memory. The method according to claim 9, including the method described in claim 9.

11. The steps include extracting one or more labeled images, The steps include: creating an anomaly record, storing one or more labeled images separated from the video stream in non-temporary machine-readable memory; A step of saving an anomaly dataset in the anomaly record, wherein the anomaly dataset includes at least one of the following: image quality score, tools used to manipulate the anomaly, location of the anomaly, or identification of the physician performing the endoscopic procedure. The method according to claim 9, including the method described in claim 9.

12. A step of generating a distribution of keywords found in the transcribed text, wherein the distribution of keywords includes counting the frequencies of one or more keywords. A step of assigning an identifier to one or more of the aforementioned keywords, The steps of identifying one or more images by associating one or more identifiers from the keywords with one or more images of the video stream when the first timestamp and the second timestamp match, The steps include: annotating the identified one or more images with the identifier to one or more of the keywords to create one or more identified and annotated images; The method according to claim 7, including the method described in claim 7.

13. The steps include transmitting one or more images, including annotations, from the video stream to the physician after the endoscopic procedure, The steps include receiving confirmation from the physician regarding the identity and location of the abnormalities in one or more images, The steps include storing the identity and arrangement of the abnormality, as well as one or more images, in a database. The method according to claim 1, including the method described in claim 1.

14. A step of receiving one or more pathology results, wherein the one or more pathology results correspond to samples associated with abnormalities from one or more images from the video stream, The steps include storing the one or more pathological results in a database along with the one or more corresponding images. The method according to claim 1, including the method described in claim 1.

15. A system for the automatic annotation of individual frames of a procedure video, An endoscope comprising an elongated member including a distal portion, wherein the elongated member includes a camera attached to the distal portion, the camera captures a video stream during the procedure, and the video stream includes a first timestamp, A microphone configured to capture an audio recording of ambient sounds during the procedure, wherein the audio recording includes a second timestamp. A natural language processor configured to receive the aforementioned audio recording and the transcribed audio recording, wherein the transcribed audio recording includes the second timestamp, Memory containing instructions, A controller including a processing circuit, wherein the processing circuit, when in operation, responds to a command, Receiving the video stream from the aforementioned camera, Receiving the audio recording from the microphone, Receiving the transcribed text from the natural language processor, Annotating the video stream with the transcribed text by associating the transcribed audio with the video stream when the first timestamp and the second timestamp match. A controller and A system equipped with these features.

16. In order to annotate the video stream, the processing circuit of the controller processes the video stream together with the first edited text file. Converting the audio recording into a text file using natural language processing, By analyzing the aforementioned text file, it is determined that a portion of the audio recording contains identifying information about the patient. The first edited text file is generated by deleting a portion of the audio recording containing identification information about the patient from the aforementioned text file. When the first timestamp and the second timestamp match, the first edited text file is associated with the video stream, thereby annotating the video stream with the first edited text file. The system according to claim 15, performed by...

17. In order to annotate the video stream, the processing circuit of the controller processes the video stream together with the audio profile text file. Accessing the voice profile for the physician performing the endoscopic procedure by associating the voice profile with the voice of the physician performing the endoscopic procedure, Editing one or more audio clips from the audio recording that do not match the audio profile in order to create an audio profile audio recording, Converting the aforementioned voice profile audio recording into the aforementioned voice profile text file, Annotating the video stream with the audio profile text file by associating the audio profile text file with the video stream when the first timestamp and the second timestamp match. The system according to claim 15, performed by...

18. In order to annotate the video stream, the processing circuit of the controller generates one or more labeled images. By analyzing the transcribed text, one or more related words indicating at least one abnormality observed during the endoscopic procedure are detected, thereby identifying that at least one abnormality was found during the endoscopic procedure. To generate a unique identification label for the at least one abnormality, wherein the unique identification label includes the second timestamp indicating when one or more related words were spoken during the endoscopic procedure, To obtain one or more images from the video stream, which includes the first timestamp corresponding to the second timestamp; The system according to claim 15, performed by...

19. The processing circuit of the controller, in accordance with the command, Recording the placement of cursors in one or more images in the first timestamp corresponding to the second timestamp, wherein the placement of cursors in one or more images indicates the placement of a pointer operated by a physician during the endoscopic procedure. In the arrangement of the cursor in the first timestamp corresponding to the second timestamp, one or more images are labeled with the unique identification label to create one or more labeled images, Replacing one or more images from the video stream with one or more labeled images, Extracting one or more labeled images, To create an anomaly record, one or more labeled images separated from the video stream are stored in the memory, The abnormality record includes storing an abnormality dataset, wherein the abnormality dataset includes at least one of the following: image quality score, tools used to manipulate the abnormality, location of the abnormality, or identification of the physician performing the endoscopic procedure. The system according to claim 18, configured to perform the following:

20. The processing circuit of the controller, in accordance with the command, The process involves generating a distribution of keywords found in the transcribed text, wherein the distribution of keywords is generated by counting the frequencies of one or more keywords. Assigning an identifier to one or more of the aforementioned keywords, Identifying one or more images by associating one or more of the identifiers from the keywords with one or more of the images in the video stream when the first timestamp and the second timestamp match, Annotating the identified one or more images with the identifiers of one or more of the keywords to create one or more identified and annotated images, After the endoscopic procedure, one or more images, including annotations, are transmitted from the video stream to the physician. The physician confirms the identity and location of the abnormalities in one or more images, Receiving one or more pathology results, wherein the one or more pathology results correspond to samples related to abnormalities from one or more images from the video stream, The identity and arrangement of the abnormality, the one or more images, and the pathological results are stored in a database along with the corresponding one or more images. The system according to claim 15, configured to perform the following: