Ai-based video documentation of endoscopy

The endoscope system automates video documentation by identifying areas of interest and integrating contextual information to create concise multimedia reports, addressing inefficiencies in existing systems and improving procedural documentation.

WO2026106727A1PCT designated stage Publication Date: 2026-05-21GYRUS ACMI INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
GYRUS ACMI INC
Filing Date
2025-10-02
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing endoscope systems lack efficient automated video documentation capabilities, leading to time-consuming manual review and annotation of lengthy procedures, limited summarization, and insufficient contextual information, which hampers efficient interpretation and documentation of critical procedural moments.

Method used

An endoscope system with an imaging device and controller circuit that identifies areas of interest and integrates contextual information to generate a multimedia procedure report, summarizing selected images or video clips with text or audio overlays.

Benefits of technology

The system reduces the workload of healthcare professionals by providing concise, comprehensive procedure summaries, enhancing efficiency and effectiveness in accessing critical moments, and facilitating improved diagnosis, treatment planning, and patient education.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025049191_21052026_PF_FP_ABST
    Figure US2025049191_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for automated video documentation of an endoscopy procedure are disclosed. An endoscope system (10) includes an endoscope (14) having an imaging device to produce images or video streams of a target anatomy, and a controller circuit (320) to identify from the images or video streams an area of interest (AOI) in the target anatomy. The controller circuit (320) receives contextual information about one or more procedure workflow events such as an event log and / or an operator observation or action log, select or adapt a portion, less than an entirety, of the images or video streams based on the AOI identified and the contextual information received, and generate a multimedia procedure report comprising the selected or adapted image or video portion and corresponding contextual information presented as text or audio overlay. The multimedia procedure report can be presented to a user, or assist in diagnosis and treatment planning.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 AI-BASED VIDEO DOCUMENTATION OF ENDOSCOPYCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of priority to U.S. Provisional Patent Applications Serial No. 63 / 719,203, filed November 12, 2024, the contents of which are incorporated herein by reference in its entirety.FIELD OF THE DISCLOSURE

[0002] The present document relates generally to endoscopy systems, and more particularly to systems and methods for guided maneuverer video documentation of an endoscopy procedure.BACKGROUND

[0003] Endoscopes have been used in a variety of clinical procedures, including, for example, illuminating, imaging, detecting and diagnosing one or more disease states, providing fluid delivery (e.g., saline or other preparations via a fluid channel) toward an anatomical region, providing passage (e.g., via a working channel) of one or more therapeutic devices or biological matter collection devices for sampling or treating an anatomical region, and providing suction passageways for collecting fluids (e.g., saline or other preparations), among other procedures. Examples of such anatomical region can include gastrointestinal tract (e.g., esophagus, stomach, duodenum, pancreaticobiliary duct, intestines, colon, and the like), renal area (e.g., kidney(s), ureter, bladder, urethra) and other internal organs (e.g., reproductive systems, sinus cavities, submucosal regions, respiratory tract), and the like.

[0004] Some endoscopes include a working channel through which an operator can perform suction, placement of diagnostic or therapeutic devices (e.g., a brush, a biopsy needle or forceps, a stent, a basket, or a balloon), or minimally invasive surgeries such as tissue sampling or removal of unwanted tissue (e.g., benign or malignant strictures) or foreign objects (e.g., calculi). Some endoscopes can be used with a laser or plasma system to deliver energy to an anatomical target (e.g., soft or hard tissue or calculi) to achieve desired treatment. For example, laser has been used in applications of tissue ablation, coagulation, vaporization, fragmentation, and lithotripsy to break down calculi inDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 kidney, gallbladder, ureter, among other stone-forming regions, or to ablate large calculi into smaller fragments. One example of endoscopy is colonoscopy, which has been shown a potential to reduce incidence and mortality rate of colorectal cancer. Colonoscopy is typically performed with rapid advance of a colonoscope to the cecum, and an endoscopist can perform thorough inspection to identify any anomalies (e.g., polyps) and to perform necessary treatment (e.g., polypectomy) during the withdrawal of the colonoscope.

[0005] Endoscope systems generally have image or video recording and still photography capabilities that allow manual image capture by an endoscopist. An imaging sensor (e.g., a camera) can be incorporated into an endoscope to take live images or video streams during an endoscopy procedure. The imaging sensor can operate under a preset imaging mode to obtain live images or video streams. The endoscope systems may also allow endoscopists to document the scenes or other findings during an endoscopy procedure, which is an important and valuable feature for diagnosis, treatment planning, patient education, and record-keeping. Some endoscope systems can intergrade with Electronic Medical Record (EMR) systems, providing instant storage directly into patient records.SUMMARY

[0006] Advances in endoscopy have allowed for better detection and management of certain types of pathologies or anomalies during an endoscopy procedure. For example, image enhanced endoscopy (IEE) techniques have allowed for better detection and management of colorectal cancers and other pathologies. However, adoption of such advanced endoscopy technologies can be operator dependent and generally requires specialized training. The high operator-dependency may lead to variability in image interpretation, anomaly diagnosis, and treatment outcomes. Advanced training on such topics can be costly.

[0007] Traditional endoscope systems primarily consist of standard video recording and still photography capabilities, allowing manual image capture by the endoscopist. In some facilities, entire endoscopy procedures (including video or image recordings) are recorded, reviewed and annotated manually by a healthcare professional (HCP), who can also manually identify clinically criticalDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 portions to be archived for future use (e.g., subsequent diagnosis, treatment planning, or for physician training or patient education). Manual video documentation can pose challenges for both HCPs and healthcare facilities. First, reviewing the entirety or an extensive portion of the procedure video recordings can be time and recourse consuming, and costly for the healthcare facilities. Manually navigating through the video recordings to quickly access and identify key moments of the procedure such as a landmark of interest, suspicious anomalies, or a critical procedural event can take a lot of time and effort.Second, existing procedure recording systems generally have limited summarization capabilities, particularly when handling lengthy video recordings of complicated procedures. Accordingly, HCPs often need to view the entire recordings to understand critical moments of a procedure, and make postprocedure annotations, which can be inefficient and even impractical in certain occasions. Furthermore, video documentation recorded by conventional endoscope systems generally lack or have insufficient contextual information. For example, current endoscope systems generally do not provide effective means for integrating into the video documentation invaluable contextual information about workflow events during an endoscopy procedure, such as physician commentary or machine-generated procedure descriptions. The lack of contextualization can make it difficult for HCPs to interpret the image and video recordings accurately and efficiently. For at least the foregoing reasons, the present inventors have recognized an unmet need for an automated video documentation system that can produce a concise yet comprehensive endoscopy procedure summary with improved effectiveness and efficiency.

[0008] The present disclosure describes endoscope systems for automated video documentation of an endoscopy procedure. The automated video documentation involves the process of creating a multimedia procedure report comprising selected or adapted images or video clips and contextual information of workflow events during the procedure summarized and integrated into the images or video clips. An exemplary endoscope system includes an endoscope having an imaging device to produce images or video streams of a target anatomy, and a controller circuit to identify from the images or video streams an area of interest (AOI) in the target anatomy. The controller circuit receives contextual information about one or more procedure workflow eventsDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 such as an event log and / or an operator observation or action log, select or adapt a portion, less than an entirety, of the images or video streams based on the AOI identified and the contextual information received, and generate a multimedia procedure report comprising the selected or adapted image or video portion and corresponding contextual information presented as text or audio overlay. The multimedia procedure report can be presented to a user, or assist in subsequent diagnosis, treatment planning, or for physician training or patient education purposes.

[0009] Example 1 is an endoscope system comprising: an endoscope, including an imaging device to obtain images or video streams of a target anatomy in a patient during an endoscopy procedure; and a controller circuit configured to: analyze the obtained images or video streams to identify an area of interest (AOI) in the target anatomy during the procedure; receive contextual information about one or more workflow events during the procedure; based at least in part on the identification of the AOI and the received contextual information, select or adapt a portion, less than an entirety, of the obtained images or video streams of the target anatomy; generate a multimedia procedure report comprising (i) the selected or adapted portion of the images or video streams, and (ii) a portion of the received contextual information related to the selected or adapted portion of the images or video streams.

[0010] In Example 2, the subject matter of Example 1 optionally includes the AOI that can include a pre-determined anatomical landmark or an anomalous structure in the target anatomy.

[0011] In Example 3, the subject matter of any one or more of Examples 1-2 optionally include the endoscope that can be a colonoscope including the imaging device to obtain images or video streams of one or more colon segments during a colonoscopy procedure, wherein the controller circuit is configured to identify a colon landmark or anomaly based on the images or video streams of one or more colon segments.

[0012] In Example 4, the subject matter of any one or more of Examples 1-3 optionally includes the controller circuit that can be configured to: generate or receive at least one trained machine-learning (ML) model each trained using a training dataset comprising images or video streams of a target anatomy and contextual information obtained during endoscopy procedures on a patientDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 population; and apply one or more of (i) the selected or adapted portion of the images or video streams, or (ii) the related potion of the contextual information, to the trained at least one ML model to perform one or more of identifying the AOI, recognizing one or more workflow events, selecting the portion of the obtained images or video streams, determining respective video playback speeds for the selected portion of video streams, or generating the multimedia procedure report.

[0013] In Example 5, the subject matter of any one or more of Examples 1-4 optionally includes the contextual information that can include an event log containing records of at least one of: a setup, or a change thereof, of the endoscope during the procedure; an operating mode, or a change thereof, of a device used in the procedure; or a patient response or a medical event occurred during the procedure.

[0014] In Example 6, the subject matter of any one or more of Examples 1-5 optionally includes the contextual information that can include an operator observation or action log containing records of at least one of: verbal command or commentary during the procedure; manipulation of the endoscope or a device used in the procedure; or a treatment action taken during the procedure.

[0015] In Example 7, the subject matter of any one or more of Examples 1-6 optionally includes the controller circuit that can be configured to select or adapt the portion of the images or video streams spatially related to the identified AOI.

[0016] In Example 8, the subject matter of any one or more of Examples 1-7 optionally include the controller circuit that can be configured to select or adapt the portion of the images or video streams temporally related to the one or more workflow events.

[0017] In Example 9, the subject matter of any one or more of Examples 1-8 optionally includes the received contextual information that can include an operator observation or action log, wherein the controller circuit is configured to apply natural language processing (NLP) to the operator observation or action log to generate structured observation or action data, and to recognize the one or more workflow events using the structured observation or action data.

[0018] In Example 10, the subject matter of any one or more of Examples 1-9 optionally includes the controller circuit that can be configured toDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 edit the selected or adapted portion of the images or video streams, and to generate the multimedia procedure report using at least the edited images or video stream portion.

[0019] In Example 11, the subject matter of Example 10 optionally includes wherein to edit the selected or adapted portion of the images or video streams includes to trim, crop, stitch, or compress the selected or adapted portion of the images or video streams.

[0020] In Example 12, the subject matter of any one or more of Examples 10-11 optionally include wherein to edit the selected or adapted portion of the images or video streams includes to dynamically adjust a video playback speed based at least in part on informational density or clinical significance.

[0021] In Example 13, the subject matter of any one or more of Examples 1-12 optionally include wherein to generate the multimedia procedure report, the controller circuit is configured to synthesize the related portion of the received contextual information into a text or audio overlay on the selected or adapted portion of the images or video streams.

[0022] In Example 14, the subject matter of any one or more of Examples 1-13 optionally includes the multimedia procedure report that can include one or more themed video summaries each comprising themed images or video segments and contextual information related thereto, the one or more themed video summaries including at least one of: a pre-procedure preparation themed video summary; a procedure phase themed video summary; an anatomical landmark themed video summary; or an anomaly detection themed video summary.

[0023] In Example 15, the subject matter of Example 14 optionally includes the controller circuit that can be configured to rank two or more themed video summaries in a specific order for prioritized presentation to the user, or for prioritized storage or transmission.

[0024] In Example 16, the subject matter of any one or more of Examples 14-15 optionally includes the controller circuit that can be configured to receive a user selection from two or more themed video summaries for presentation to the user, or for storage or transmission.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0025] Example 17 is a method for real-time endoscopy scene analysis and video documentation of an endoscopy procedure performed using an endoscope. The method comprises steps of: obtaining images or video streams of a target anatomy during an endoscopy procedure using an imaging device associated with an the endoscope; analyzing the obtained images or video streams to identify an area of interest (AO I) in the target anatomy during the endoscopy procedure; receiving contextual information about one or more workflow events during the procedure; based at least in part on the identification of the AOI and the received contextual information, selecting or adapting a portion, less than an entirety, of the obtained images or video streams of the target anatomy; generating a multimedia procedure report comprising (i) the selected or adapted portion of the images or video streams, and (ii) a portion of the received contextual information related to the selected or adapted portion of the images or video streams.

[0026] In Example 18, the subject matter of Example 17 optionally includes the AOI that can include a pre-determined anatomical landmark or an anomalous structure in the target anatomy.

[0027] In Example 19, the subject matter of any one or more of Examples 17-18 optionally includes the images or video streams that are related to one or more colon segments obtained during a colonoscopy procedure, and wherein the identified AOI includes a colon landmark or anomaly.

[0028] In Example 20, the subject matter of any one or more of Examples 17-19 optionally includes training at least one machine-learning (ML) model using a training dataset comprising images or video streams of a target anatomy and contextual information obtained during endoscopy procedures on a patient population; and applying one or more of (i) the selected or adapted portion of the images or video streams or (ii) the related potion of the contextual information to the trained at least one ML model to perform one or more of identifying the AOI, recognizing one or more workflow events, selecting the portion of the obtained images or video streams, determining respective video playback speeds for the selected portion of video streams, or generating the multimedia procedure report.

[0029] In Example 21, the subject matter of any one or more of Examples 17-20 optionally includes the contextual information that can includeDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 an event log containing records of at least one of: a setup, or a change thereof, of the endoscope during the procedure; an operating mode, or a change thereof, of a device used in the procedure; or a patient response or a medical event occurred during the procedure.

[0030] In Example 22, the subject matter of any one or more of Examples 17-21 optionally includes the contextual information that can include an operator observation or action log containing records of at least one of: verbal command or commentary during the procedure; manipulation of the endoscope or a device used in the procedure; or a treatment action taken during the procedure.

[0031] In Example 23, the subject matter of any one or more of Examples 17-22 optionally includes selecting or adapting the portion of the images or video streams based on a spatial relationship to the identified AOI or a temporal relationship to the one or more workflow events.

[0032] In Example 24, the subject matter of any one or more of Examples 17-23 optionally includes the received contextual information that can include an operator observation or action log, the method further comprising: applying natural language processing (NLP) to the operator observation or action to generate structured observation or action data; and recognizing the one or more workflow events using the structured observation or action data.

[0033] In Example 25, the subject matter of any one or more of Examples 17-24 optionally includes editing the selected or adapted portion of the images or video streams, and generating the multimedia procedure report using at least the edited images or video stream portion.

[0034] In Example 26, the subject matter of Example 25 optionally includes editing the selected or adapted portion of the images or video streams by trimming, cropping, stitching, or compressing the selected or adapted portion of the images or video streams; or dynamically adjusting a video playback speed based at least in part on informational density or clinical significance.

[0035] In Example 27, the subject matter of any one or more of Examples 17-26 optionally includes generating the multimedia procedure report by synthesizing the related portion of the received contextual information into a text or audio overlay on the selected or adapted portion of the images or video streams.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0036] In Example 28, the subject matter of any one or more of Examples 17-27 optionally includes the multimedia procedure report comprising one or more themed video summaries each comprising themed images or video segments and contextual information related thereto, the one or more themed video summaries including at least one of: a pre-procedure preparation themed video summary; a procedure phase themed video summary; an anatomical landmark themed video summary; or an anomaly detection themed video summary.

[0037] In Example 29, the subject matter of Example 28 optionally includes ranking two or more themed video summaries in a specific order for prioritized presentation to the user, or for prioritized storage or transmission to a storage device.

[0038] In Example 30, the subject matter of any one or more of Examples 28-29 optionally include receiving a user selection from two or more themed video summaries for presentation to the user, or for storage or transmission to a storage device.

[0039] This summary is an overview of some of the teachings of the present application and not intended to be an exclusive or exhaustive treatment of the present subject matter. Further details about the present subject matter are found in the detailed description and appended claims. Other aspects of the disclosure will be apparent to persons skilled in the art upon reading and understanding the following detailed description and viewing the drawings that form a part thereof, each of which are not to be taken in a limiting sense. The scope of the present disclosure is defined by the appended claims and their legal equivalents.BRIEF DESCRIPTION OF THE DRAWING

[0040] FIGS. 1-2 are schematic diagrams illustrating an example of an endoscope system for use in endoscopy procedures, procedures, such as a colonoscopy procedure.

[0041] FIG. 3 illustrates an example of an endoscope system for realtime endoscopy scene analysis and video documentation of an endoscopy procedure.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0042] FIG. 4 illustrates an example process of generating a multimedia procedure report of a colonoscopy procedure.

[0043] FIG. 5 illustrates an example process of selecting and editing images or video segments and using the same to create a multimedia procedure report.

[0044] FIGS. 6A-6B illustrate examples of training a machine learning (ML) model, and generate a multimedia procedure report comprising video documentation with text or audio overlay using the trained ML model.

[0045] FIG. 7 is a flowchart illustrating an example method of real-time endoscopy scene analysis and video documentation of an endoscopy procedure.

[0046] FIG. 8 is a block diagram illustrating an example machine upon which any one or more of the techniques (e.g., methodologies) discussed herein may perform.DETAILED DESCRIPTION

[0047] This document describes systems, devices, and methods for automated video documentation of an endoscopy procedure with selected or adapted images or video segments and procedure-specific contextual information. An endoscope system includes an endoscope having an imaging device to produce images or video streams of a target anatomy, and a controller circuit to identify from the images or video streams an area of interest (AO I) in the target anatomy. The controller circuit receives contextual information about one or more procedure workflow events such as an event log and / or an operator observation or action log, select or adapt a portion, less than an entirety, of the images or video streams based on the AOI identified and the contextual information received, and generate a multimedia procedure report comprising the selected or adapted image or video portion and corresponding contextual information presented as text or audio overlay. The multimedia procedure report can be presented to a user, or assist in subsequent diagnosis, treatment planning, or for physician training or patient education purposes.

[0048] The systems, devices, and methods described herein may be used in various endoscopy procedures to improve real-time inspection and detection and diagnosis of pathologies, and to automate the process of video documentation of an endoscopy procedure with improved quality and efficiency.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 The automated video documentation as described herein can reduce the workload of HCPs in reviewing, navigating, summarizing, and annotating the procedure, help provide quick and comprehensive access to critical moments of the procedure. In accordance with various embodiments, the present system can transform lengthy videos into a concise (e.g., 60-second) multimedia procedure report that incorporates selected images or video segments with contextual information of key procedure workflow events (manually identified and / or automatically detected), including operator (e.g., endoscopist) observations, commentary, and actions during the procedure, patient vital signs and medical events occurred, device settings and usage of consumables, and machinegenerated insights. In certain examples, the multimedia procedure report or an aspect thereof can be created using artificial intelligence (Al) or machinelearning (ML) techniques. The contextual information can be represented by text or audio overlay upon the selected images or video segments, which provides a comprehensive overview of the procedure in a more efficient and user-friendly manner. The multimedia procedure report as describes herein enables more effective and efficient documentation, fast retrieval of critical information for diagnosis and subsequent treatment planning, archival for long-term storage with low footprint, patient education, and physician training.

[0049] In accordance with various embodiments, the endoscope system as described herein may be designed for efficient communication with medical institutions, archiving, and enhancing patient records. For example, the automated video documentation system proposed herein (or a portion thereof) may be integrated into existing endoscopy equipment, such as via an edgecomputing device mounted on the endoscopy tower. The edge-computing device may be connected to an electronic health record (EHR) system, or a hospital integration system. The edge-computing device may capture live video feed through a video port from the endoscope’s video processor (CV), such as a serial digital interface (SDI) connection for transmitting digital video and audio signals using coaxial or fiber optic cables). The recorded procedure video may be processed in real-time or post-procedure, depending on the setup and requirements. The summarized video can then be converted into proper format (e.g., Digital Imaging and Communications in Medicine, or DICOM) and send to the EHR or other hospital integration system.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0050] Although the discussion in this document focuses endoscope systems and methods of using the same in endoscopy procedures, this is meant only by way of example and not limitation. It is within the contemplation of the inventors, and within the scope of this document, that the systems, devices, and methods pertaining to automated video documentation as discussed herein may also be used in other medical examination and diagnostic procedures involving imaging capabilities including, for example, fluoroscopy, ultrasound imaging, magnetic resonance imaging and radiography procedures, among others. For example, during an endobronchial ultrasound (EBUS) procedure, ultrasound image analysis of the lungs and nearby lymph nodes can be used for landmark or nodule detection and characterization, and one or more multimedia procedure reports can be generated using the automated video documentation as described herein.

[0051] FIG. 1 is a schematic diagram of an endoscope system 10 for use in an endoscopy procedure, such as colonoscopy. The system 10 can include an imaging and control system 12 and an endoscope 14. The system 10 is an illustrative example of an endoscope system suitable for use with the systems, devices, and methods described herein, such as a colonoscopy system for use in image-guided colonoscopy with automated video documentation as described in this document.

[0052] The endoscope 14 can be insertable into an anatomical region for imaging or to provide passage of or attachment to (e.g., via tethering) one or more sampling devices for biopsies or therapeutic devices for treatment of a disease state associated with the anatomical region. The endoscope 14 can interface with and connect to imaging and control system 12. The endoscope 14 may be a colonoscope, though other types of endoscopes can be used with the features and teachings of the present disclosure. The imaging and control system 12 can include a control unit 16, an output unit 18, an input unit 20, a light source unit 22, a fluid source 24, and a suction pump 26.

[0053] The imaging and control system 12 can include various ports for coupling with the endoscope system 10. For example, the control unit 16 can include a data input / output port for receiving data from and communicating data to the endoscope 14. The light source unit 22 can include an output port for transmitting light to the endoscope 14, such as via a fiber optic link. The fluidDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 source 24 can include a port for transmitting fluid to the endoscope 14. The fluid source 24 can include, for example, a pump and a tank of fluid or can be connected to an external tank, vessel, or storage unit. The suction pump 26 can include a port to draw a vacuum from the endoscope 14 to generate suction, such as for withdrawing fluid from the anatomical region into which the endoscope 14 is inserted. The output unit 18 and the input unit 20 can be used by an operator of the endoscope system 10 (e.g., an endoscopist) to control functions of the endoscope system 10 and view the output of the endoscope 14. The control unit 16 can also generate signals or other outputs from treating the anatomical region into which the endoscope 14 is inserted. In some examples, the control unit 16 can generate electrical output, acoustic output, fluid output, and the like for treating the anatomical region with, for example, cauterizing, cutting, freezing, and the like.

[0054] The fluid source 24 can be in communication with control unit 16 and can include one or more sources of air, saline, or other fluids, as well as associated fluid pathways (e.g., air channels, irrigation channels, suction channels, or the like) and connectors (barb fittings, fluid seals, valves, or the like). The fluid source 24 can be utilized as an activation energy for a biasing device or a pressure-applying device of the present disclosure. The imaging and control system 12 can also include the drive unit 46, which can include a motorized drive for advancing a distal section of endoscope 14.

[0055] The endoscope 14 can include an insertion section 28, a functional section 30, and a handle section 32, which can be coupled to a cable section 34 and a coupler section 36. The insertion section 28 can extend distally from the handle section 32, and the cable section 34 can extend proximally from the handle section 32. The insertion section 28 can be elongated and can include a bending section and a distal end to which the functional section 30 can be attached. The bending section can be controllable (e.g., by a control knob 38 on the handle section 32) to maneuver the distal end through tortuous anatomical passageways (e.g., stomach, duodenum, kidney, ureter, etc.). The insertion section 28 can also include one or more working channels (e.g., an internal lumen) that can be elongated and can support the insertion of one or more therapeutic tools of the functional section 30. The working channel can extend between the handle section 32 and the functional section 30. AdditionalDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 functionalities, such as fluid passages, guide wires, and pull wires, can also be provided by the insertion section 28 (e.g., via suction or irrigation passageways or the like).

[0056] A coupler section 36 can be connected to the control unit 16 to connect to the endoscope 14 to multiple features of the control unit 16, such as the input unit 20, the light source unit 22, the fluid source 24, and the suction pump 26.

[0057] The handle section 32 can include the knob 38 and the port 40A. The knob 38 can be connected to a pull wire or other actuation mechanisms that can extend through the insertion section 28. The port 40 A, as well as other ports, such as a port 40B (FIG. 2), can be configured to couple various electrical cables, guide wires, auxiliary scopes, tissue collection devices, fluid tubes, and the like to the handle section 32, such as for coupling with the insertion section 28.

[0058] According to examples, the imaging and control system 12 can be provided on a mobile platform (e.g., a cart 41) with shelves for housing the light source unit 22, the suction pump 26, an image processing unit 42 (FIG. 2), etc. Alternatively, several components of the imaging and the control system 12 (shown in FIGS. 1 and 2) can be provided directly on the endoscope 14 to make the endoscope “self-contained.”

[0059] The functional section 30 can include components for treating and diagnosing anatomy of a patient. The functional section 30 can include an imaging device, an illumination device, and an elevator. The functional section 30 can further include optically enhanced biological matter and tissue collection and retrieval devices. For example, the functional section 30 can include one or more electrodes conductively connected to the handle section 32 and functionally connected to the imaging and control system 12 to analyze biological matter in contact with the electrodes based on comparative biological data stored in the imaging and control system 12. In other examples, the functional section 30 can directly incorporate tissue collectors.

[0060] In some examples, the endoscope 14 can be robotically controlled, such as by a robot arm attached thereto. The robot arm can automatically, or semi-automatically (e.g., with certain user manual control or commands), via an actuator, position and navigate the endoscope 14 (e.g., theDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 functional section 30 and / or the insertion section 28) in the target anatomy or position a device at a desired location with desired posture to facilitate an operation of an anatomical target. In accordance with various examples discussed in this document, a controller can generate a control signal to the actuator of the robot arm to facilitate anomaly inspection and diagnosis under the target or recommended imaging mode in a robotically assisted endoscopy procedure.

[0061] FIG. 2 is a schematic diagram of the endoscope system 10 of FIG.1 including the imaging and control system 12 and the endoscope 14. FIG. 2 schematically illustrates components of the imaging and the control system 12 coupled to the endoscope 14, which in the illustrated example includes a colonoscope. The imaging and control system 12 can include the control unit 16, which can include or be coupled to an image processing unit 42, a treatment generator 44, and a drive unit 46, as well as the light source unit 22, the input unit 20, and the output unit 18. The control unit 16 can include, or can be in communication with, an endoscope, a surgical instrument 48, and an endoscope system, which can include a device configured to engage tissue and collect and store a portion of that tissue and through which imaging equipment (e.g., a camera) can view target tissue via inclusion of optically enhanced materials and components. The control unit 16 can be configured to activate a camera to view target tissue distal of the endoscope system. Likewise, the control unit 16 can be configured to activate the light source unit 22 to shine light on the surgical instrument 48, which can include select components configured to reflect light in a particular manner, such as enhanced tissue cutters with reflective particles.

[0062] The coupler section 36 can be connected to the control unit 16 to connect to the endoscope 14 to multiple features of the control unit 16, such as the image processing unit 42 and the treatment generator 44. In examples, the port 40A can be used to insert another surgical instrument 48 or device, such as a daughter scope or auxiliary scope, into the endoscope 14. Such instruments and devices can be independently connected to the control unit 16 via the cable 47. In examples, the port 40B can be used to connect coupler section 36 to various inputs and outputs, such as video, air, light, and electric.

[0063] The image processing unit 42 and light source unit 22 can each interface with the endoscope 14 (e.g., at the functional section 30) by wired orDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 wireless electrical connections. The imaging and control system 12 can accordingly illuminate an anatomical region, collect signals representing the anatomical region, process signals representing the anatomical region, and display images representing the anatomical region on the display unit 18. The imaging and control system 12 can include the light source unit 22 to illuminate the anatomical region using light of desired spectrum (e.g., broadband white light, narrow-band imaging using preferred electromagnetic wavelengths, and the like). The imaging and control system 12 can connect (e.g., via an endoscope connector) to the endoscope 14 for signal transmission (e.g., light output from light source, video signals from imaging system in the distal end, diagnostic and sensor signals from a diagnostic device, and the like).

[0064] The treatment generator 44 can generate a treatment plan, which can be used by the control unit 16 to control the operation of the endoscope 14, or to provide with the operating physician a guidance for maneuvering the endoscope 14, during an endoscopy procedure. In an example, the treatment generator 44 can generate an endoscope navigation plan, including estimated values for one or more cannulation or navigation parameters (e.g., an angle, a force, etc.) for maneuvering the steerable elongate instrument, using patient information including an image of the target anatomy. The endoscope navigation plan can help guide the operating physician to cannulate and navigate the endoscope in the patient anatomy. The endoscope navigation plan may additionally or alternatively be used to robotically adjust the position, angle, force, and / or navigation of the endoscope or other instrument.

[0065] FIG. 3 illustrates an example of an endoscope system 300 for real-time endoscopy scene analysis and automated video documentation (also referred to as “video summarization”) of an endoscopy procedure. The endoscope system 300 may identify critical anatomical and pathologic findings based on substantially real-time scene analysis, and create a multimedia procedure report comprising selected or adapted images or video segments, and contextual information about one or more workflow events during the procedure. By way of example and not limitation, the endoscope system 300 may be used in a colonoscopy procedure. The endoscope system 300 may be implemented as a part of the control unit 16 in FIG. 1.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0066] The endoscope system 300 may include one or more of an endoscope 310, contextual source(s) 315, a controller circuit 320, a user interface 330, and a storage device 340. The endoscope 310 can be an example of the endoscope 14 as described above and shown in FIGS. 1-2. The endoscope 310 may include, among other things, an imaging system 312 and a lighting system 314. The imaging system 312 can include at least one imaging sensor or device (e.g., a camera) configured to obtain images or video streams of a target anatomy of a patient during an endoscopy procedure. The imaging sensor or device may be located at a distal portion or a distal end of the endoscope 310. The imaging system 312 may be controllably adjusted to operate on different settings including zoom settings, contrast settings, exposure levels, or viewing angles toward and around the target anatomy. The lighting system 314 may include one or more light sources to produce illumination on the target anatomy via one or more lighting lenses. The lighting system 314 may be controllably adjusted to provide different lighting or illumination conditions. The imaging system 312 and the lighting system 314 together can define an imaging mode for capturing endoscopic images or video streams of the target anatomy.

[0067] The controller circuit 320 may include circuit sets comprising one or more other circuits or sub-circuits that may, alone or in combination, perform the functions, methods, or techniques described herein. In an example, the controller circuit 320 and the circuits sets therein may be implemented as a part of a microprocessor circuit, which may be a dedicated processor such as a digital signal processor, application specific integrated circuit (ASIC), microprocessor, or other type of processor for processing information including physical activity information. Alternatively, the microprocessor circuit may be a general-purpose processor that may receive and execute a set of instructions of performing the functions, methods, or techniques described herein. In an example, hardware of the circuit set may be immutably designed to carry out a specific operation (e.g., hardwired). In an example, the hardware of the circuit set may include variably connected physical components (e.g., execution units, transistors, simple circuits, etc.) including a computer readable medium physically modified (e.g., magnetically, electrically, moveable placement of invariant massed particles, etc.) to encode instructions of the specific operation. In connecting the physical components, the underlying electrical properties of a hardware constituent areDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 changed, for example, from an insulator to a conductor or vice versa. The instructions enable embedded hardware (e.g., the execution units or a loading mechanism) to create members of the circuit set in hardware via the variable connections to carry out portions of the specific operation when in operation. Accordingly, the computer readable medium is communicatively coupled to the other components of the circuit set member when the device is operating. In an example, any of the physical components may be used in more than one member of more than one circuit set. For example, under operation, execution units may be used in a first circuit of a first circuit set at one point in time and reused by a second circuit in the first circuit set, or by a third circuit in a second circuit set at a different time.

[0068] The controller circuit 320 can create a video documentation of an endoscopy procedure using images or video segments selected or adapted from the endoscopic images or video streams provided by the imaging system 312, and at least a portion of contextual information received from the contextual source(s) 315 during the endoscopy procedure. The contextual source(s) 315 may include an event log containing records of endoscope setup information, operating modes of one or more devices used in the procedure, or patient responses or medical events occurred during the procedure. The contextual source(s) 315 may additionally or alternatively include an operator (e.g., the endoscopist) observation or action log, such as related to recognition of critical anatomical structures or clinically significant events during the procedure.Examples of the operator observations or actions may include verbal command or commentary during the procedure (e.g., voice commands with “wake words” like “Record” and “Stop Recording”, commentary about the findings during the procedure, commentary about patient past and current indications and procedures received, patient demographic and medical information, recognition of landmarks, characterization of anatomical structures, diagnosis of anomalies), manipulation of the endoscope system or one or more devices being used in the procedure (e.g., press of an endo button such as for waterjet or insufflation, actuation of a foot switch, selection of a zooming mode such as for closer mucosal observation, changing of an imaging mode or lighting mode, such as high-definition white light imaging (WLI), chromoendoscopy techniques like dye-based, or virtual CE like narrow-band imaging (NBI), texture and colorDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 enhancement (TXI) imaging, red dichromatic imaging (RDI), etc., operation of tools for tissue section, sampling, or treatment, among others), or treatment actions taken during the procedure (e.g., polypectomy, clipping, or clotting during the procedure).

[0069] The controller circuit 320 can include an image processor 321, an area of interest (AOI) detector 322, a device tracking circuit 323, a procedure summarization circuit 324, and a real-time alert and recommendation circuit 329. The image processor 321 can analyze the images or video streams obtained from the imaging system 312, and generate endoscopic image or video features.Examples of the image or video features include statistical features of pixel values or morphological features, such as comers, edges, blobs, curvatures, speeded up robust features (SURF), or scale-invariant feature transform (SIFT) features, among others. In some examples, the image processor 321 may pre-process the images or video streams, such as filtering, resizing, orienting, or color or grayscale correction, and the endoscopic features may be extracted from the pre-processed images or video streams. In some examples, the image processor 321 may post-process the image features to enhance feature quality, such as edge interpolation or extrapolation to produce continuous and smooth edges.

[0070] The AOI detector 322 may detect an AOI in the target anatomy based at least in part on the endoscopic image or video streams and identify one or more characteristics of the AOI. The AOI may include one or more predetermined anatomical landmarks in the target anatomy. Systemic endoscopic image documentation of such anatomical landmarks serves the purposes of showing crucial anatomic structures at these landmark regions, documenting the extent of the examination, and reflecting the quality of cleansing and mucosal visualization. In an example of upper GI endoscopy, the pre-determined landmarks may include, for example, some or all of proximal esophagus, distal esophagus, Z-line and diaphragm indentation, cardia and fundus on retroflexed view, body (including lesser curvature), body on retroflexed view, Angulus on partial retroflexion, antrum, duodenal bulb, and second part of the duodenum (including the ampulla). In an example of lower GI (colon, in particular) endoscopy, the pre-determined landmarks may include, for example, some or all of lower part of the rectum in retroflexed view, lower part of the rectum, middleDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 part of the sigmoid, descending colon just distal to the splenic flexure, transverse colon just proximal to the splenic flexure, transverse colon just distal to the hepatic flexure, ascending colon just proximal to the hepatic flexure, cecum and ileocecal valve, and cecum and appendiceal orifice. In an example of colonoscopy procedure, the AOI detector 322 may identify one or more predetermined colon landmarks from images or video streams of one or more colon segments, which can be captured during withdrawal of the colonoscope.

[0071] The AOI to be identified from the endoscopic images or video streams may additionally or alternatively include an anomalous structure in the target anatomy. This may include one or more of a presence or absence, a type, a size, a shape, a location, or an amount of pathological tissue or anatomical structures, among other objects in an environment of the target anatomy. In an example of colonoscopy, the anomalous structure may include pathological tissue segments, such as mucosal abnormalities (polyps, inflammatory bowel diseases, Meckel’s diverticulum, lipoma, bleeding, vascularized mucosa etc.), or obstructed mucosa (e.g., segments with bad bowel preparation, distended colon etc.).

[0072] Various techniques may be used to detect and characterize an AOI. In an example, the AOI detector 322 can detect an AOI using a template matching technique involving a comparison of the endoscopic image features (e.g., features characterizing shapes or contours of a structure) to one or more pre-generated templates of known anomalous structure. In another example, the AOI detector 322 can use an Al-driven technique to detect an AOI, where the endoscopic images or video streams (or features extracted therefrom) may be fed into a trained ML model that automatically outputs an AOI identification. The ML model may be trained to establish a correspondence between an endoscopic image or video stream (or features extracted therefrom) and one or more recognizable types or characterizations of landmarks or anomalies. Examples of ML models for recognizing an anatomical landmark include Deep Belief Network, ResNet, DenseNet, Autoencoders, capsule networks, generative adversarial networks, Siamese networks, Convolutional Neural Networks (CNN), deep reinforcement learning, support vector machine (SVM), Bayesian models, decision trees, k-means clustering, among other ML models. Examples of ML models for recognizing anomaly from endoscopic images or videoDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 streams include Convolutional Neural Networks, bi-directional LSTM, Recurrent Neural Networks, Conditional Random Fields, Dictionary Learning, or other machine learning techniques (support vector machine, Bayesian models, decision trees, k-means clustering), among other ML techniques. The trained ML model for automatically recognizing an AOI may be a part of the trained ML modes 360 stored in the storage device 340. Examples of training an ML model and using the same to detect an AOI are discussed below with respect to FIGS. 6A-6B.

[0073] The device tracking circuit 323 may track in substantially real time location and orientation of the endoscope or a portion thereof in a particular segment of the anatomy during an endoscopy procedure. In an example of colonoscopy, the device tracking circuit 323 can continuously monitor and record the location and orientation of the colonoscope, and align the position and orientation information with a standardized colon location template. In an example, the device tracking circuit 323 can track the location and orientation of an imaging device (e.g., a camera) at a distal portion of the endoscope during the procedure. With such real-time tracking information, the AOI detector 322 may associate a detected AOI (e.g., an anatomical landmark or an anomaly) with the segment of the anatomy. In an example, the device tracking can be based on an electromagnetic property of the imaging device, or the distal portion of the endoscope detected during the endoscopy procedure. The device tracking circuit 323 may register a tracked location of the imaging device to a pre-generated template of the target anatomy during the endoscopy procedure, and the location of the imaging device can be determined based on a degree of template matching. In another example, the device tracking can be based on one or more anatomical landmarks identified from the images or video streams (or features extracted therefrom) such as obtained during the insertion phase of endoscopy. Once the endoscope location and orientation are determined, the device tracking circuit 323 may register the substantially real-time endoscope location and orientation to a pre-generated template of the anatomy. Information of the endoscope location and orientation may be presented to the user on the user interface 330.

[0074] The procedure summarization circuit 324 may create a video documentation of the procedure that comprises selected or adapted images orDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 video segments and contextual information about one or more workflow events related to the selected or adapted images or video segments. Real-time location and orientation of the endoscope or the imaging device associated therewith, such as produced by the device tracking circuit 323, may also be used for creating the video documentation. The procedure summarization circuit 324 can comprise subcircuits or functional modules, including for example on or more of an image / video selector 325, an image / video editor 326, a context synthesizer 327, and a multimedia report compiler 328. The image / video selector 325 can select or adapt a portion, less than an entirety, of the images or video streams based at least in part on the identification of the AOI and contextual information about one or more workflow events, such as received from the contextual source(s) 315. In an example, images or video segments spatially related to the identified AOI (e.g., images or video segments containing at least a portion of the AOI) can be selected. Additionally or alternatively, the image / video selector 325 can select or adapt images or video segments that are temporally related to one or more workflow events during the procedure. During the procedure, images and video recordings and the workflow events can both be timestamped. The image / video selector 325 can recognize one or more workflow events from the contextual source(s) 315, such as an event log or an operator observation or action log, and use such recognized workflow events to “flag” clinically important video segments. For example, images or video segments that are time-synchronized with the one or more workflow events can be selected.

[0075] In some examples, the contextual information includes an operator observation or action log, and the image / video selector 325 can include a natural language processing (NLP) script generator to produce structured operator observation or action script, and to recognize one or more workflow events from the structured observation or action script. The NLP may involve one or more of computational linguistic models, machine learning or deep learning models, among other Al-based algorithms to recognize crucial observations or actions satisfying a specified condition. The image / video selector 325 can select images or video segments temporally related to (e.g., time-synchronized with) the operator’s verbal cues indicative of crucial observations or actions.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0076] In some examples, the image / video selector 325 may identify the one or more workflow events using a trained machine-learning (ML) model trained to establish a correspondence between contextual information and characteristics of one or more workflow events. In some examples, the ML model may be further trained to prioritize the recognized workflow events such as based on the procedural context and actions taken by the operator. The trained ML model may be a part of the trained ML modes 360 stored in the storage device 340. Examples of training an ML model and using the same to identify workflow events are discussed below with respect to FIGS. 6A-6B.

[0077] The image / video editor 326 can edit the selected or adapted video segments, including one or more operations of image / video trimming, cropping, stitching, or compression. For example, time-cropping can be applied to each of the selected video segments to trim unnecessary footage, focusing on the core content. The trimmed segments can be stitched together into a continuous stream with smooth transitions between scenes, while maintaining logical flow and visual consistency.

[0078] In an example, the image / video editor 326 can dynamically adjust video playback speed of one or more selected video segments. The dynamic speed adjustment can help fit the selected video segments (along with the associated or time-synchronized contextual information) into a compact video documentation of a specified duration (e.g., approximately 60 seconds) without losing important details. By way of non-limiting example, the video playback speed can vary within a specific range between 0.5x and 4x of a “base” speed (e.g., the video recording speed), depending on an informational density of the video segment and / or its clinical importance. Events with higher informational density or of higher clinical significance (e.g., an identified anatomical landmark or a suspicious anomaly with complex structural details, or associated with operator commentary or actions indicative of higher significance) can trigger a slower playback speed, and vice versa. This allows more time to be allocated to complex or significant portions of the procedure, while less critical portions can be displayed briefly. In some examples, the image / video editor 326 can use an Al-based algorithm, such as a trained ML model, to automatically determine a proper video playback speed for each identified video segment based on the informational density and its clinical importance.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0079] The multimedia report compiler 328 can generate a multimedia procedure report using the selected or adapted images or video segments and at least a portion of the contextual information about the procedure workflow events relevant to the selected or adapted images or video segments. In some examples, a context synthesizer 327 may synthesize multiple contextual sources including, for example, an event log (e.g., Al-generated logs detailing characteristics of the recognized AOI or detected workflow events), operator observation or action log (e.g., tool usage, technique application), ambient verbal commands or commentary (or a transcription via NLP), endoscopy report entries (including, for example, critical insights from the user-generated endoscopy report, diagnostic conclusions, and recommendations), among others. The synthesized contextual information may be presented as a text or audio overlay on the selected images or video segments. In an example where the contextual information or a portion thereof takes the form of text, the context synthesizer 327 may perform text-to-speech synthesis, such as using a large language model (LLM) with appropriate grounding and prompt engineering, to generate timestamped narrative to be overlaid upon the selected images or video segments. In some examples, visual cues or markers may be superimposed on the video segments included in the multimedia procedure report to signify the AOI (e.g., identified landmarks or anomalous areas), thereby guiding the user’s focus on those areas of interest.

[0080] In an example, the multimedia report compiler 328 may feed the selected image or video segments, and the synthesized contextual information related to (e.g., time-synchronized with) the selected images or video segments, to a trained machine-learning (ML) model to generate the multimedia procedure report. The trained ML model may be a part of the trained ML modes 360 stored in the storage device 340. Examples of training an ML model and using the same to generate the multimedia procedure report are discussed below with respect to FIGS. 6A-6B.

[0081] The multimedia report compiler 328 may refine the multimedia procedure report before being presented to a user or archived for future use. The refinement can include, for example, “fine-tuning” the images or video segments included therein (e.g., with additional video trimming and stitching), modifying the text and audio overlay (e.g., with further editing or annotation), or dataDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 compression to reduce the file size of the report for efficient storage and transmission without significantly impacting video quality. The multimedia procedure report may be formatted to comply with a standard, such as a Digital Imaging and Communications in Medicine (DICOM) format, to allow or facilitate interoperability.

[0082] The multimedia procedure report may be provided to a user, such as displayed on the user interface 330. Additionally or alternatively, the multimedia procedure report may be stored in the storage device 340. The storage device 340 can be local to the endoscope system 300, or alternatively a remote storage device, such as an Electronic Health Records (EHR) system, a Picture Archiving and Communication System (PACS), or other integrated patient record systems. In response to a user command or a trigger event, the multimedia procedure report may be transferred to the remote storage device for easy access and review by healthcare professionals. The multimedia procedure report may assist in subsequent patient diagnosis, treatment planning, or for physician training or patient education purposes. In some examples, the remote storage device can be a part of a cloud comprising one or more storage and computing devices (e.g., servers) that provides secure access to cloud-based services including, for example, data storage, computing services, and provisioning of customer services, among others. In some examples, at least some of the data processing and computation by the controller circuit 320 as described above may be performed in a cloud. For example, images or video streams or features extracted therefrom may be streamed to the cloud, get processed therein, and the computation results (e.g., the personalized pathologyspecific imaging modality) may be relayed back to local endoscope system.

[0083] In some examples, the multimedia report compiler 328 may generate one or more themed video summaries each corresponding to a particular theme. Each themed video summary comprises themed images or video segments and contextual information related thereto (e.g., time-synchronized therewith). The themed images or video segments can be selected using the image / video selector 325. The contextual information can be a subset of the time-synchronized, synthesized contextual information produced by the context synthesizer 327. Examples of the themed video summaries may include a pre-procedure preparation themed video summary (e.g., bowel preparationDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 theme for a colonoscopy procedure), a procedure phases themed video summary (e.g., an insertion phase and a withdrawal phase of a colonoscopy procedure, withdrawal speed at various colon segments), an anatomical landmark themed documentation (e.g., one or more colon segments or specific structures therein), or an anomaly detection themed documentation (e.g., polys detected during the colonoscopy procedure). For example, a pre-procedure preparation themed video summary may include selected images or video segments of various colon segments or areas indicative of quality of bowel preparation, as well as operator observations, commentaries, or actions related to patient bowel preparation. In another example, an anomaly themed video summary may include selected images or video segments of identified anomalous structures (such as automatically detected by the AOI detector 322), as well as operator observations, commentaries, and actions related to anomaly recognition. The multimedia report compiler 328 may rank multiple themed video summaries in a specific order for prioritized presentation to the user, or for prioritized storage or transmission. In an example, the order may be based on clinical significance such as determined by computer-aided diagnostic or other anomaly detection and characterisation techniques. For example, a video summary of a cancerous colon polyp can be ranked higher than a video summary of a benign colon polyp. In some examples, anomalous structures (e.g., colon polyps) can be characterized or classified based on their appearance under narrow-band imaging (NBI). This classification helps distinguish between different types of polyps (e.g., hyperplastic polyp, adenomatous polyps, deep submucosal invasive cancer), which is crucial for determining the appropriate treatment and management. In some examples, multiple themed video summaries may be indexed and selectable by the themes. A user may select, via the user interface 330, a themed video summary for presentation, or for storage or transmission. In some examples, a multimedia report template can be created for an individual user, which may include one or more user-preferred themed video summaries. For example, users (e.g., endoscopists) of different levels of experience or focusing on different clinical issues may identify their preferred themed video summaries to be created, and a preferred order to be presented.

[0084] The real-time alert and recommendation circuit 329 can generate an alert of the identified AOI (e.g., an anatomical landmark or an anomaly) orDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 other events detected during the procedure. The alert may be delivered through optical means on the diagnostic monitor, or via auditory means like a warning alarm. In an example, highlighting, flash alerts, audible or haptic feedback may be provided to the user. Images or video streams of AOI (e.g., anomalies) identified by the AOI detector 322 may be presented to the user on the user interface 330.

[0085] The real-time alert and recommendation circuit 329 can provide real-time recommendation to record images or video recordings related to the identified AOI and to include the same in the multimedia procedure report. The user may review the images or video recordings of the identified AOIs, provide adequate contextual information such as commentary on the AOI (e.g., anomaly characterization and diagnosis), and either take the recommendation to make video documentation as presented, or reject the recommendation and ignore the identified anomaly. In some examples, the endoscope system 300 may integrate an iterative feedback mechanism for continuous enhancement in recognizing AOI and workflow events of interest. This can involve collecting feedback from endoscopists on recommendations, such as their explicit feedback (e.g., satisfaction levels with the recommendations, a “like” button etc.). Alternatively, the endoscope system 300 can run in a shadow mode to monitor actions of experienced endoscopists and correct or refine the recommendation based on user interactions. For instance, if a suggestion is ignored or times out, the system 300 can instantly recalibrate its subsequent recommendations. Advanced algorithms like Q-learning or Markov decision processes might be employed, adapting the system to the specific preferences of each institution. Metrics like recommendation acceptance rate can be tracked to further optimize the efficacy of refining recommendations.

[0086] FIG. 4 illustrates an example process of generating a multimedia procedure report of a colonoscopy procedure, such as using the endoscope system 300. The colonoscopy procedure, as illustrated, includes colonoscope insertion and subsequent colonoscope withdrawal phases. During the insertion phase, endoscopic images or video streams can be obtained respectively for each of a plurality of colon segments, including the rectosigmoid segment, sigmoid segment, descending segment, transverse segment, ascending segment, and cecum, such as using the imaging system 312. Various AOIs, including preDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 determined anatomical landmarks and anomalies, may be identified automatically using the AOI detector 322 during the insertion phase or alternatively the withdrawal phase. In an example, at the end of the insertion phase when the distal tip of the colonoscope is determined to be reaching the cecum (such as determined by the device tracking circuit 323), AOI detection process can be be manually or automatically initiated. As described above with respect to FIG. 3, the AOI detector 322 may analyze the scans (i.e., images or video streams) of the endoscopic feed using template-matching or AI / ML based algorithm, and identify unusual or suspicious regions indicative of pathological abnormalities, such as polyps or cancerous tissue. Images or video segments that are spatially related to the identified AOIs can be selected, along with the associated timestamps.

[0087] Contextual information about one or more workflow events, such as an event log or an operator observation or action log, may be collected during the procedure. As illustrated in FIG. 4, the operator observation or action log can include tracked operator actions such as “scope in” 421, “reaching caecum” 422, “polyp detected” 423, “polyp removed” 424, and “scope out” 425. The operator observation or action log can further include operator’s voice commands or commentary on findings during the procedure, such as voice commentaries 431, 432, 433, and 435 as shown in FIG. 4. In some examples, the voice commands or commentary can be processed using a NPL script generator, as described above with respect to FIG. 3. Each of the workflow events, such as the tracked operator actions and / or operator voice commands and commentary, can be timestamped, and synchronized with the images or video segments which are also timestamped during the procedure. In the example illustrated in FIG. 4, image or video segment 412 is temporally related to user action 422 and voice commentary 432, images or video segments 413 and 414 are temporally related to user actions 423 and 424 and voice commentary 433, image or video segment 415 is temporally related to user action 425 and voice commentary 435. Images or video segments that are temporally related to the workflow events can be selected, along with the associated timestamps. In one example, the temporally related images or video segments can be those time-synchronized with the workflow events. In another example, the temporally related images or video segments can be those suggested or implied by the operator’s voice commandsDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 or commentary. For example, if the commentary is made at time T, and the commentary refers to a landmark encountered or an event occurred at a prior time (T-5 seconds) or a prior scene, then the images or video segments taken at that time can be selected.

[0088] As illustrated in FIG. 4, the contextual information may further include user commands to freeze frames and / or intra-procedure reporting, such as commands or reports 441 through 447. Similar to operator actions and voice commentary, these user commands and reports can also be timestamped and used for selecting images or video segments to be included in the video summarization 450.

[0089] Various contextual information as described above may be combined with the selected images or video segments to create a video summarization 450 of the colonoscopy procedure. Various sources of contextual information, including user actions, voice commentary, and user commands to freeze frames, may be synthesized into text or audio overlay on the selected images or video segments, such as an audio voiceover 451 that can be integrated into the video summarization 450 to create a multimedia procedure report 460, which can further be presented to the user or archived in a storage system such as aPACS 470.

[0090] FIG. 5 illustrates an example process of selecting and editing images or video segments obtained during an endoscopy procedure (such as the colonoscopy procedure shown in FIG. 4), and using the same to create a multimedia procedure report, such as using the endoscope system 300. Video segments 412, 413, 414, and 415, which are spatially related to one or more identified AOIs or temporally related to one or more workflow events during a colonoscopy as described above with respect to FIG. 4, can be further selected based on clinical significance, or a user specified theme, as described above with respect to FIG. 3. In the illustrated example, to create an anomaly themed video summary, only video segments 413 (corresponding to user action 423 “polyp detected”) and 414 (corresponding to user action 424 “polyp removed”) are selected. One or more editing operations 510, including image or video strimming, cropping, and dynamic playback speed adjustment as described above with respect to FIG. 3, may be applied to each of the selected video segments 413 and 414. The resulting edited video segments 513 and 514 can beDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 further processed via an integration step 520, including video stitching to produce a continuous stream with adequate logical flow and visual consistency, and video compression to reduce the file size of the summarized video stream 530 without significantly impacting video quality.

[0091] Contextual information 540 about one or more workflow events, such as user actions and observations, reported findings, event logs including event stamps, and Al-generated logs, can be recorded during the procedure, and analyzed using an NLP script generator 550. The resulting scrip 551 can be further synthesized using a text-to-speech synthesizer 560 (an embodiment of the context synthesizer 327 of system 300 in FIG. 3) to produce narrative and voiceover 561. A video mixer 570 can integrate the narrative and voiceover 561 into the summarized video stream 530 to produce a multimedia procedure report, which can be archived in a storage device 580.

[0092] FIGS. 6A-6B illustrate examples of training a machine learning (ML) model, and generate a multimedia procedure report comprising video documentation with text or audio overlay using the trained ML model. In an example, the same ML model may also be trained to accomplish multiple subtasks and output various aspects of the multimedia procedure report, such as AOI detection, workflow events recognition, image or video segments selection, or dynamically determined video segment playback speed. In some examples, multiple ML models may be trained separately to determine respective different aspects of the multimedia procedure report. The multiple, individually trained ML models may have different model structures. Examples of the such individually trained ML models may include one or more of a first ML model trained to recognize an AOI (e.g., an anomaly) from the image or video streams, a second ML model can be trained to recognize one or more workflow events, a third ML mode trained to select or adapt video segments to be included in the procedure report, a fourth ML model trained to determine a video playback speed of a video segment, or a fifth ML model trained to produce a multimedia procedure report such as based on the inference results from other trained ML models.

[0093] By way of example, FIG. 6A illustrates an ML model training phase during which an ML model 630 may be trained to generate a multimedia procedure report, or an aspect thereof as described above. The training phaseDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 includes constructing a training dataset that comprises a plurality of endoscopic images or video streams 610 of a particular anatomical target (e.g., one or more colon segments) and contextual training data 620. The endoscopic images or video streams 610 can be recorded from the same type of procedures (e.g., colonoscopy) performed on a plurality of patients. The contextual training data 620 may include contextual information about workflow events (e.g., event log, and operator observation, commentary, and action log) recorded during respective procedures. Endoscope location information for each of the plurality of endoscopic images or video streams may also be included in the training dataset.

[0094] The ML model 630 may have a neural network structure comprising an input layer, one or more hidden layers, and an output layer. The plurality of endoscopic images or video streams 610 and the contextual training data 620 may be fed into the input layer of the ML model 630, which propagates the input data through one or more hidden layers to the output layer that outputs a recommended pathology-specific imaging modality. The ML model 630 can perform tasks, without explicitly being programmed, by making inferences based on patterns found in the analysis of data. The ML model 630 explores the study and construction of algorithms (e.g., ML algorithms) that may learn from existing data and make predictions about new data. Such algorithms operate by building the ML model 630 from training data in order to make data-driven predictions or decisions expressed as outputs or assessments.

[0095] The ML model 630 may be trained using supervised learning or unsupervised learning. Supervised learning uses prior knowledge (e.g., examples that correlate inputs to outputs or outcomes) to learn the relationships between the inputs and the outputs. The goal of supervised learning is to learn a function that, given some training data, best approximates the relationship between the training inputs and outputs so that the ML model can implement the same relationships when given inputs to generate the corresponding outputs.Unsupervised learning is the training of an ML algorithm using information that is neither classified nor labelled and allowing the algorithm to act on that information without guidance. Unsupervised learning is useful in exploratory analysis because it can automatically identify structure in data.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0096] Common tasks for supervised learning are classification problems and regression problems. Classification problems, also referred to as categorization problems, aim at classifying items into one of several category values. Regression algorithms aim at quantifying some items (for example, by providing a score to the value of some input). Some examples of commonly used supervised-ML algorithms are Logistic Regression (LR), Naive-Bayes, Random Forest (RF), neural networks (NN), deep neural networks (DNN), matrix factorization, and Support Vector Machines (SVM). Examples of DNN include a convolutional neural network (CNN), a recurrent neural network (RNN), a deep belief network (DBN), or a hybrid neural network comprising two or more neural network models of different types or different model configurations. Some common tasks for unsupervised learning include clustering, representation learning, and density estimation. Some examples of commonly used unsupervised learning algorithms are K-means clustering, principal component analysis, and autoencoders.

[0097] Another type of ML is federated learning (also known as collaborative learning) that trains an algorithm across multiple decentralized devices holding local data, without exchanging the data. This approach stands in contrast to traditional centralized machine-learning techniques where all the local datasets are uploaded to one server, as well as to more classical decentralized approaches which often assume that local data samples are identically distributed. Federated learning enables multiple actors to build a common, robust machine learning model without sharing data, thus allowing to address critical issues such as data privacy, data security, data access rights and access to heterogeneous data.

[0098] The training of the ML model 630 may be performed continuously or periodically, or in near real time as additional procedure data are made available. The training process involves algorithmically adjusting one or more ML model parameters (e.g., weights or bias at any particular layer of a neural network model), until the ML model being trained satisfies a specified training convergence criterion. By way of example and not limitation, the ML model 630 may be trained with weighted square loss (for explicit feedback) or with binary cross-entropy loss (for implicit feedback). Other training techniques, such as deep factorization machine, wide and deep learning, deep structuredDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 semantic models, or autoencoder based recommender systems, may be used. The trained ML model 630 can establish a correspondence between the endoscopic images or video streams 610 and an ideal or recommended endoscope maneuvering plan, including a guided endoscope navigation path and / or a recommended imaging mode or setting.

[0099] The trained ML model 630 can establish a correspondence between the training dataset (e.g., the endoscopic images or video streams 610 and the contextual training data 620) and a multimedia procedure report, such as a video documentation with text or audio overlay. In various examples, different ML algorithms may be used to learn different aspects or components of the multimedia procedure report. In an example of anomaly detection and characterization, video stream of the target anatomy (e.g., colon segments) can be subjected to real-time analysis using a trained ML model. Leveraging Convolutional Neural Networks, bi-directional LSTM, Recurrent Neural Networks, and Conditional Random Fields, the trained ML model can effectively pinpoint an anomaly and accordingly select video segments spatially related to the detected anomaly. Other machine learning techniques such as decision trees and k-means clustering, may also be employed for thorough analysis. In some examples, the trained ML model can include a spatio-temporal network trained to not only detect anomalies but also to extract essential features from the video stream. These features can then be utilized to assist in video documentation and procedure recommendation.

[0100] FIG. 6B illustrates an inference phase, where live images or video feed 650 during an endoscopy procedure and associated contextual information 660 can be applied to the trained ML model 630 (obtained from the training phase as illustrated in FIG. 6A) to generate a multimedia procedure report 670. The multimedia procedure report 670 may be presented to a user (e.g., an endoscopist), or archived for future use (e.g., subsequent diagnosis, treatment planning, or for physician training or patient education). Depending on the architecture of the trained ML model 630, in some examples, other aspects of the multimedia procedure report, such as anomaly detection and characterization, workflow events recognition, image or video segment selection, or dynamic video playback speed, may also be determined during the inference phase, and provided to the user or archived as needed.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0101] FIG. 7 is a flowchart illustrating an example method 700 for realtime endoscopy scene analysis and video documentation of an endoscopy procedure. The method 700 may be implemented in the endoscope system 300 to generate a multimedia procedure report comprising selected or adapted images or video segments and relevant contextual information. Although the processes of the method 700 are drawn in one flow chart, they are not required to be performed in a particular order. In various examples, some of the processes can be performed in a different order than that illustrated herein.

[0102] At step 710, images or video streams of distinct segments of the target anatomy may be obtained using an imaging system associated with an endoscope. The images or video streams may be acquired when the imaging system is set to one of a plurality of available imaging modes or settings. An imaging mode refers to one or more of a lighting modality, an optical magnification, or a viewing angle of the imaging device. Examples of lighting modality may include high-definition white light imaging (WLI), chromoendoscopy techniques like dye-based, or virtual CE like narrow-band imaging (NBI), texture and color enhancement (TXI) imaging, red dichromatic imaging (RDI), among others. The optical magnification defines a zoom setting (e.g., zoom in or zoom out of a suspicious anomaly of the target anatomy). The viewing angle of the imaging device, also referred to as a field of view, describes the angular extent of a given scene that is imaged by the imaging device.

[0103] At step 720, the obtained images or video streams may be analyzed to identify an area of interest (AOI) in the target anatomy, such as using the AOI detector 322. One example of the AOI includes one or more predetermined anatomical landmarks in the target anatomy, such as landmarks of the upper GI, or landmarks of the lower GI duct. The AOI may additionally or alternatively include an anomalous structure in the target anatomy, such as one or more of a presence or absence, a type, a size, a shape, a location, or an amount of pathological tissue or anatomical structures, among other objects in an environment of the target anatomy. In some examples, the AOI may include a “blind spot” region of the target anatomy, such as hidden places inadvertently missed for lesion screening during colonoscopy. In colonoscopy for example, because colon has innumerable folds and two flexures, such anatomical characteristics may cause certain. The AOI (e.g., a pre-determined landmark orDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 an anomaly) may be identified using techniques such as a template matching technique, or Al or ML based techniques, as described above with respect to FIG. 3.

[0104] At step 730, contextual information about one or more workflow events during the procedure can be received. The contextual information may include an event log containing records of setup information of the endoscope system, operating modes of one or more devices being used in the procedure, or the patient’s response or medical events occurred during the procedure. The contextual information may additionally or alternatively include an operator observation or action log identifying structures or events of clinical significance during the procedure. Examples of the operator observations or actions may include verbal command or commentary during the procedure, manipulation of the endoscope system or one or more devices being used in the procedure, or treatment actions taken during the procedure (e.g., polypectomy, or clipping or clotting of an anatomical structure during the procedure).

[0105] At step 740, a portion, less than an entirety, of the obtained images or video streams of the target anatomy may be selected or adapted from the receive images or video streams based at least in part on the identification of the AOI and the received contextual information, such as using the image / video selector 325. The selection or adaptation can be based on a spatial relationship to the identified AOI, such that video segments containing the identified AOI can be selected. Additionally or alternatively, selection or adaptation can be based on a temporal relationship to (e.g., time-synchronized with) the recognized one or more workflow events. In some examples, the received contextual information includes an operator observation or action log. A natural language processing (NLP) script generator and generate structured operator observation or action script, and one or more workflow events can then be recognized using the structured observation or action script.

[0106] The selected or adapted video segments can further be edited, such as by apply one or more of image or video trimming, cropping, stitching, or compression. For example, time-cropping can be applied to each of the selected video segments to trim unnecessary footage, focusing on the core content. The trimmed segments can be stitched together into a continuous stream with smooth transitions between scenes, while maintaining logical flow and visualDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 consistency. The editing may additionally or alternatively include dynamically adjusting a video playback speed of the video segments, such that the selected images or video segments (along with the time-synchronized contextual information) can be fitted into a compact video documentation of a specified duration (e.g., approximately 60 seconds) without losing important details. In an example, the video playback speed can be dynamically determined or adjusted based on informational density or clinical significance of event.

[0107] At step 750, a multimedia procedure report can be generated, such as using the multimedia report compiler 328. The multimedia procedure report comprises the selected or adapted portion of the images or video streams, and a portion of the received contextual information related to the selected or adapted portion of the images or video streams. In some examples, multiple sources of contextual information may be synthesized into a text or audio overlay on the selected images or video segments to produce the multimedia procedure report. Examples of the multiple contextual data sources can include an event log, operator observation or action log, ambient verbal commands or commentary, or endoscopy report entries, among others. In an example where the contextual information or a portion thereof takes the form of text, text-to-speech synthesis, such as using a large language model (LLM) with appropriate grounding and prompt engineering, may be used to generate timestamped narrative overlaid upon the selected images or video segments. In some examples, visual cues or markers may be superimposed on the video segments included in the multimedia procedure report.

[0108] The multimedia procedure report may be refined before being released or archived. Examples of the refinement process include “fine-tuning” the video segments, further editing or annotation of the text and audio overlay, data compression of the report, among others. The multimedia procedure report may be formatted to comply with a standard to facilitate interoperability.

[0109] In an example, the multimedia procedure report can include one or more themed video summaries each comprising themed images or video segments and contextual information related thereto. Examples of the themed video summaries may include a pre-procedure preparation themed video summary, a procedure phases themed video summary, an anatomical landmark themed video summary, or an anomaly detection themed documentation. In anDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 example, multiple themed video summaries may be ranked in a specific order (e.g., an order of clinical significance) for prioritized presentation to the user, or for prioritized storage or transmission. In another example, multiple themed video summaries may be indexed and selectable by the themes. A user may select a themed video summary for presentation, or for storage or transmission.

[0110] In some examples, one or more trained machine-learning (ML) models may be used to identify an AOI at step 720, to recognize one or more workflow events from the contextual information received at step 730, to select or adapt a portion of the images or video segments at step 740, or to generate the multimedia procedure report at step 750. The one or more ML models may each be trained using respective training datasets comprising, for example, images or video streams of a target anatomy and contextual information obtained during endoscopy procedures on a patient population. As described above with respect to FIGS. 6A-6B, during the inference phase, one or more of (i) the selected or adapted portion of the images or video streams or (ii) the related potion of the contextual information may be fed into the one or more trained ML models to perform actions such as identifying the AOI, recognizing one or more workflow events, or generating the multimedia procedure report.

[0111] The multimedia procedure report may be provided to a user (e.g., an endoscopist) or a process. More specifically, in one example, at step 762, the multimedia procedure report may be displayed to the user (e.g., an endoscopist). In another example, at step 764, an alert may be generated to notify or warn the user about the identified AOI, and receiving user’s feedback. Real-time recommendation to record images or video recordings related to the identified AOI and to include the same in the multimedia procedure report may also be provided to the user. In yet another example, at step 766, the multimedia procedure report may be archived in a storage device, such as an Electronic Health Records (EHR) system or a Picture Archiving and Communication System (PACS), among other integrated patient record systems. The multimedia procedure report may be used in subsequent diagnosis, treatment planning, or for physician training or patient education, among other uses.

[0112] FIG. 8 illustrates generally a block diagram of an example machine 800 upon which any one or more of the techniques (e.g., methodologies) discussed herein may perform. Portions of this description mayDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 apply to the computing framework of various portions of the endoscope system 300.

[0113] In alternative embodiments, the machine 800 may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, the machine 800 may operate in the capacity of a server machine, a client machine, or both in server-client network environments. In an example, the machine 800 may act as a peer machine in peer-to-peer (P2P) (or other distributed) network environment. The machine 800 may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile telephone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein, such as cloud computing, software as a service (SaaS), other computer cluster configurations.

[0114] Examples, as described herein, may include, or may operate by, logic or a number of components, or mechanisms. Circuit sets are a collection of circuits implemented in tangible entities that include hardware (e.g., simple circuits, gates, logic, etc.). Circuit set membership may be flexible over time and underlying hardware variability. Circuit sets include members that may, alone or in combination, perform specified operations when operating. In an example, hardware of the circuit set may be immutably designed to carry out a specific operation (e.g., hardwired). In an example, the hardware of the circuit set may include variably connected physical components (e.g., execution units, transistors, simple circuits, etc.) including a computer readable medium physically modified (e.g., magnetically, electrically, moveable placement of invariant massed particles, etc.) to encode instructions of the specific operation. In connecting the physical components, the underlying electrical properties of a hardware constituent are changed, for example, from an insulator to a conductor or vice versa. The instructions enable embedded hardware (e.g., the execution units or a loading mechanism) to create members of the circuit set in hardware via the variable connections to carry out portions of the specific operation whenDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 in operation. Accordingly, the computer readable medium is communicatively connected to the other components of the circuit set member when the device is operating. In an example, any of the physical components may be used in more than one member of more than one circuit set. For example, under operation, execution units may be used in a first circuit of a first circuit set at one point in time and reused by a second circuit in the first circuit set, or by a third circuit in a second circuit set at a different time.

[0115] Machine (e.g., computer system) 800 may include a hardware processor 802 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 804 and a static memory 806, some or all of which may communicate with each other via an interlink (e.g., bus) 808. The machine 800 may further include a display unit 810 (e.g., a raster display, vector display, holographic display, etc.), an alphanumeric input device 812 (e.g., a keyboard), and a user interface (UI) navigation device 814 (e.g., a mouse). In an example, the display unit 810, input device 812 and UI navigation device 814 may be a touch screen display. The machine 800 may additionally include a storage device (e.g., drive unit) 816, a signal generation device 818 (e.g., a speaker), a network interface device 820, and one or more sensors 821, such as a global positioning system (GPS) sensor, compass, accelerometer, or other sensors. The machine 800 may include an output controller 828, such as a serial (e.g., universal serial bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection to communicate or control one or more peripheral devices (e.g., a printer, card reader, etc.).

[0116] The storage device 816 may include a machine-readable medium 822 on which is stored one or more sets of data structures or instructions 824 (e.g., software) embodying or utilized by any one or more of the techniques or functions described herein. The instructions 824 may also reside, completely or at least partially, within the main memory 804, within static memory 806, or within the hardware processor 802 during execution thereof by the machine 800. In an example, one or any combination of the hardware processor 802, the main memory 804, the static memory 806, or the storage device 816 may constitute machine-readable media.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0117] While the machine-readable medium 822 is illustrated as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) configured to store the one or more instructions 824.

[0118] The term “machine-readable medium” may include any medium that is capable of storing, encoding, or carrying instructions for execution by the machine 800 and that cause the machine 800 to perform any one or more of the techniques of the present disclosure, or that is capable of storing, encoding or carrying data structures used by or associated with such instructions. Nonlimiting machine-readable medium examples may include solid-state memories, and optical and magnetic media. In an example, a massed machine-readable medium comprises a machine-readable medium with a plurality of particles having invariant (e.g., rest) mass. Accordingly, massed machine-readable media are not transitory propagating signals. Specific examples of massed machine-readable media may include non-volatile memory, such as semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EPSOM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0119] The instructions 824 may further be transmitted or received over a communication network 826 using a transmission medium via the network interface device 820 utilizing any one of a number of transfer protocols (e.g., frame relay, internet protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), etc.). Example communication networks may include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), mobile telephone networks (e.g., cellular networks), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., Institute of Electrical and Electronics Engineers (IEEE) 802.11 family of standards known as WiFi®, IEEE 802.16 family of standards known as WiMax®), IEEE 802.15.4 family of standards, peer-to-peer (P2P) networks, among others. In an example, the network interface device 820 may include one or more physical jacks (e.g., Ethernet, coaxial, or phonejacks) or one or more antennas to connect to the communication network 826. In an example, the network interface device 820 may include a plurality of antennas toDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 wirelessly communicate using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) techniques. The term “transmission medium” shall be taken to include any intangible medium that is capable of storing, encoding or carrying instructions for execution by the machine 800, and includes digital or analog communications signals or other intangible medium to facilitate communication of such software.Additional Notes

[0120] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, specific embodiments in which the invention can be practiced. These embodiments are also referred to herein as “examples.” Such examples can include elements in addition to those shown or described. However, the present inventors also contemplate examples in which only those elements shown or described are provided. Moreover, the present inventors also contemplate examples using any combination or permutation of those elements shown or described (or one or more aspects thereof), either with respect to a particular example (or one or more aspects thereof), or with respect to other examples (or one or more aspects thereof) shown or described herein.

[0121] In this document, the terms “a” or “an” are used, as is common in patent documents, to include one or more than one, independent of any other instances or usages of “at least one” or “one or more.” In this document, the term “or” is used to refer to a nonexclusive or, such that “A or B” includes “A but not B,” “B but not A,” and “A and B,” unless otherwise indicated. In this document, the terms “including” and “in which” are used as the plain-English equivalents of the respective terms “comprising” and “wherein.” Also, in the following claims, the terms “including” and “comprising” are open-ended, that is, a system, device, article, composition, formulation, or process that includes elements in addition to those listed after such a term in a claim are still deemed to fall within the scope of that claim. Moreover, in the following claims, the terms “first,” “second,” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01

[0122] The above description is intended to be illustrative, and not restrictive. For example, the above-described examples (or one or more aspects thereof) may be used in combination with each other. Other embodiments can be used, such as by one of ordinary skill in the art upon reviewing the above description. The Abstract is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Also, in the above Detailed Description, various features may be grouped together to streamline the disclosure. This should not be interpreted as intending that an unclaimed disclosed feature is essential to any claim. Rather, inventive subject matter may lie in less than all features of a particular disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description as examples or embodiments, with each claim standing on its own as a separate embodiment, and it is contemplated that such embodiments can be combined with each other in various combinations or permutations. The scope of the invention should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

Claims

Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 What is claimed is:

1. An endoscope system, comprising:an endoscope, including an imaging device to obtain images or video streams of a target anatomy in a patient during an endoscopy procedure; and a controller circuit configured to:analyze the obtained images or video streams to identify an area of interest (AOI) in the target anatomy during the procedure;receive contextual information about one or more workflow events during the procedure;based at least in part on the identification of the AOI and the received contextual information, select or adapt a portion, less than an entirety, of the obtained images or video streams of the target anatomy; andgenerate a multimedia procedure report comprising (i) the selected or adapted portion of the images or video streams, and (ii) a portion of the received contextual information related to the selected or adapted portion of the images or video streams.

2. The endoscope system of claim 1, wherein the AOI includes a predetermined anatomical landmark or an anomalous structure in the target anatomy.

3. The endoscope system of any of claims 1-2, wherein the endoscope is a colonoscope including the imaging device to obtain images or video streams of one or more colon segments during a colonoscopy procedure,wherein the controller circuit is configured to identify a colon landmark or anomaly based on the images or video streams of one or more colon segments.

4. The endoscope system of any of claims 1-3, wherein the controller circuit is configured to:generate or receive at least one trained machine-learning (ML) model each trained using a training dataset comprising images or video streams of aDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 target anatomy and contextual information obtained during endoscopy procedures on a patient population; andapply one or more of (i) the selected or adapted portion of the images or video streams, or (ii) the related potion of the contextual information, to the trained at least one ML model to perform one or more of identifying the AOI, recognizing one or more workflow events, selecting the portion of the obtained images or video streams, determining respective video playback speeds for the selected portion of video streams, or generating the multimedia procedure report.

5. The endoscope system of any of claims 1-4, wherein the contextual information includes an event log containing records of at least one of:a setup, or a change thereof, of the endoscope during the procedure; an operating mode, or a change thereof, of a device used in the procedure; ora patient response or a medical event occurred during the procedure.

6. The endoscope system of any of claims 1-5, wherein the contextual information includes an operator observation or action log containing records of at least one of:verbal command or commentary during the procedure; manipulation of the endoscope or a device used in the procedure; or a treatment action taken during the procedure.

7. The endoscope system of any of claims 1-6, wherein the controller circuit is configured to select or adapt the portion of the images or video streams spatially related to the identified AOI.

8. The endoscope system of any of claims 1-7, wherein the controller circuit is configured to select or adapt the portion of the images or video streams temporally related to the one or more workflow events.

9. The endoscope system of any of claims 1-8, wherein the received contextual information includes an operator observation or action log,Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 wherein the controller circuit is configured to apply natural language processing (NLP) to the operator observation or action log to generate structured observation or action data, and to recognize the one or more workflow events using the structured observation or action data.

10. The endoscope system of any of claims 1-9, wherein the controller circuit is configured to edit the selected or adapted portion of the images or video streams, and to generate the multimedia procedure report using at least the edited images or video stream portion.

11. The endoscope system of claim 10, wherein to edit the selected or adapted portion of the images or video streams includes to trim, crop, stitch, or compress the selected or adapted portion of the images or video streams.

12. The endoscope system of claim 10, wherein to edit the selected or adapted portion of the images or video streams includes to dynamically adjust a video playback speed based at least in part on informational density or clinical significance.

13. The endoscope system of any of claims 1-12, wherein to generate the multimedia procedure report, the controller circuit is configured to synthesize the related portion of the received contextual information into a text or audio overlay on the selected or adapted portion of the images or video streams.

14. The endoscope system of any of claims 1-13, wherein the multimedia procedure report comprises one or more themed video summaries each comprising themed images or video segments and contextual information related thereto, the one or more themed video summaries including at least one of a pre-procedure preparation themed video summary;a procedure phase themed video summary;an anatomical landmark themed video summary; oran anomaly detection themed video summary.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 15. The endoscope system of claim 14, wherein the controller circuit is configured to rank two or more themed video summaries in a specific order for prioritized presentation to the user, or for prioritized storage or transmission.

16. The endoscope system of claim 14, wherein the controller circuit is configured to receive a user selection from two or more themed video summaries for presentation to the user, or for storage or transmission.

17. A method for real-time endoscopy scene analysis and video documentation of an endoscopy procedure performed using an endoscope, the method comprising:obtaining images or video streams of a target anatomy during an endoscopy procedure using an imaging device associated with the endoscope; analyzing the obtained images or video streams to identify an area of interest (AOI) in the target anatomy during the endoscopy procedure;receiving contextual information about one or more workflow events during the procedure;based at least in part on the identification of the AOI and the received contextual information, selecting or adapting a portion, less than an entirety, of the obtained images or video streams of the target anatomy; andgenerating a multimedia procedure report comprising (i) the selected or adapted portion of the images or video streams, and (ii) a portion of the received contextual information related to the selected or adapted portion of the images or video streams.

18. The method of claim 17, wherein the AOI includes a pre-determined anatomical landmark or an anomalous structure in the target anatomy.

19. The method of any of claims 17-18, wherein the images or video streams are related to one or more colon segments obtained during a colonoscopy procedure, and wherein the identified AOI includes a colon landmark or anomaly.

20. The method of any of claims 17-19, comprising:Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 training at least one machine-learning (ML) model using a training dataset comprising images or video streams of a target anatomy and contextual information obtained during endoscopy procedures on a patient population; and applying one or more of (i) the selected or adapted portion of the images or video streams or (ii) the related potion of the contextual information to the trained at least one ML model to perform one or more of identifying the AOI, recognizing one or more workflow events, selecting the portion of the obtained images or video streams, determining respective video playback speeds for the selected portion of video streams, or generating the multimedia procedure report.

21. The method of any of claims 17-20, wherein the contextual information includes an event log containing records of at least one of:a setup, or a change thereof, of the endoscope during the procedure; an operating mode, or a change thereof, of a device used in the procedure; ora patient response or a medical event occurred during the procedure.

22. The method of any of claims 17-21, wherein the contextual information includes an operator observation or action log containing records of at least one of:verbal command or commentary during the procedure; manipulation of the endoscope or a device used in the procedure; or a treatment action taken during the procedure.

23. The method of any of claims 17-22, wherein selecting or adapting the portion of the images or video streams is based on a spatial relationship to the identified AOI or a temporal relationship to the one or more workflow events.

24. The method of any of claims 17-23, wherein the received contextual information includes an operator observation or action log, the method further comprising:applying natural language processing (NLP) to the operator observation or action to generate structured observation or action data; andDocket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 recognizing the one or more workflow events using the structured observation or action data.

25. The method of any of claims 17-24, further comprising editing the selected or adapted portion of the images or video streams, and generating the multimedia procedure report using at least the edited images or video stream portion.

26. The method of claim 25, wherein editing the selected or adapted portion of the images or video streams includes:trimming, cropping, stitching, or compressing the selected or adapted portion of the images or video streams; ordynamically adjusting a video playback speed based at least in part on informational density or clinical significance.

27. The method of any of claims 17-26, wherein generating the multimedia procedure report includes synthesizing the related portion of the received contextual information into a text or audio overlay on the selected or adapted portion of the images or video streams.

28. The method of any of claims 17-27, wherein the multimedia procedure report comprises one or more themed video summaries each comprising themed images or video segments and contextual information related thereto, the one or more themed video summaries including at least one of:a pre-procedure preparation themed video summary;a procedure phase themed video summary;an anatomical landmark themed video summary; oran anomaly detection themed video summary.

29. The method of claim 28, comprising ranking two or more themed video summaries in a specific order for prioritized presentation to the user, or for prioritized storage or transmission to a storage device.Docket No.: 5409.970W01Client Ref No.: GAP24090-DUNV-W01 30. The method of any of claims 28-29, comprising receiving a user selection from two or more themed video summaries for presentation to the user, or for storage or transmission to a storage device.