Systems and methods for processing colon images and videos
The GUI system enhances colonoscopy by dynamically tracking polyps and providing real-time feedback, addressing missed polyp detection and removal issues through vector-based and 3D reconstruction techniques, ensuring thorough colon visualization and improved procedural efficiency.
Patent Information
- Application Number
- JP2025113027
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-04
- Filing Date
- 2025-07-03
- Publication Date
- 2025-09-25
AI Technical Summary
Existing colonoscopy techniques struggle with high rates of missed polyp detection and removal due to operator fatigue, inattention, and limitations in visualizing the entire colon surface, leading to potential missed adenomas and cancers.
A system and method for generating a graphical user interface (GUI) that dynamically tracks polyps in real-time using vector representations and 3D reconstructions, providing real-time feedback on polyp locations, camera movement, and colon coverage, enabling accurate polyp detection and removal.
Improves polyp detection and removal rates by ensuring complete visualization of the colon surface and providing real-time guidance on polyp locations, reducing the risk of missed lesions and enhancing the efficiency of colonoscopy procedures.
Smart Images

Figure 2025138833000001_ABST
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims the benefit of priority to U.S. Provisional Patent Application No. 16 / 430,461, filed June 4, 2019, the contents of which are incorporated herein by reference in their entirety. [Background technology]
[0002] The present invention, in some embodiments thereof, relates to colonoscopy, and more particularly, but not exclusively, to systems and methods for processing colon images and videos and / or processing colon polyps automatically detected during a colonoscopy procedure.
[0003] Colonoscopy is the gold standard for detecting colon polyps. During a colonoscopy, a long, flexible tube called a colonoscope is advanced through the colon. A video camera at the end of the colonoscope captures images, which are then shown to the physician on a display. The physician examines the lining of the colon to check for polyps. Any polyps identified are removed using the colonoscope's instruments. Early removal of cancerous polyps can eliminate or reduce the risk of colon cancer. Summary of the Invention [Means for solving the problem]
[0004] According to a first aspect, a method for generating instructions to present a graphical user interface (GUI) for dynamically tracking at least one polyp in multiple endoscopic images of a patient's colon includes repeating for multiple endoscopic images: tracking the position of a region depicting at least one polyp in each endoscopic image relative to at least one previous endoscopic image; if the position of the region is outside the respective endoscopic image, calculating a vector from within the respective endoscopic image to the position of the region outside the respective endoscopic image; creating an augmented endoscopic image by augmenting each endoscopic image with a representation of the vector; and generating instructions for displaying the augmented endoscopic image in the GUI.
[0005] According to a second aspect, a method for generating instructions for presenting a GUI for dynamically tracking 3D movement of an endoscopic camera capturing multiple 2D endoscopic images within a patient's colon includes, for each endoscopic image of the multiple endoscopic images, iteratively: feeding each 2D endoscopic image to a 3D reconstruction neural network; outputting a 3D reconstruction of each 2D endoscopic image by the 3D reconstruction neural network, wherein pixels of the 2D endoscopic image are assigned 3D coordinates; calculating a current 3D position of the endoscopic camera within the colon according to the 3D reconstruction; and generating instructions for presenting the current 3D position of the endoscopic camera on a colon map in the GUI.
[0006] According to a third aspect, a method for calculating a three-dimensional volume of a polyp based on at least one two-dimensional (2D) image includes receiving at least one 2D image of an inner surface of the colon captured by an endoscopic camera positioned within a lumen of the colon, receiving a representation of a region of the at least one 2D image depicting at least one polyp, supplying the at least one 2D image to a 3D reconstruction neural network, outputting a 3D reconstruction of the at least one 2D image by the 3D reconstruction neural network, wherein pixels of the 2D image are assigned 3D coordinates, and calculating an estimated 3D volume of at least one polyp within the region of the at least one 2D image according to an analysis of the 3D coordinates of the pixels of the region of the at least one 2D image.
[0007] In a further embodiment of the first aspect, the vector representation depicts a direction and / or orientation for adjusting the endoscopic camera to capture at least one other endoscopic image depicting a region of the at least one endoscopic image.
[0008] In a further embodiment of the first aspect, when the location of a region depicting at least one polyp appears in the respective endoscopic image, an augmented endoscopic image is generated by augmenting the respective endoscopic image with the location of the region, and a representation of the vector is excluded from the augmented endoscopic image.
[0009] A further embodiment of the first aspect further includes calculating the location of an area depicting at least one polyp in the patient's colon, creating a colon map by marking a schematic representation of the patient's colon with an indication indicating the location of the area depicting the at least one polyp, and generating instructions for presenting the colon map in a GUI, wherein the colon map is dynamically updated with the location of newly detected polyps.
[0010] In a further embodiment of the first aspect, the method further includes translating and / or rotating at least one endoscopic image of the consecutive subset of the plurality of endoscopic images, including each endoscopic image to generate a processed consecutive subset of the plurality of endoscopic images, wherein a region depicting at least one polyp is in the same approximate position in all of the images of the consecutive subset of the plurality of endoscopic images; providing the processed consecutive subset of the plurality of endoscopic images to a detection neural network; outputting, by the detection neural network, a current region depicting at least one polyp for each endoscopic image; creating an augmented image of each endoscopic image by augmenting each endoscopic image with the current region; and generating instructions for presenting the augmented image in a GUI.
[0011] In a further embodiment of the first aspect, an output of the neural network is provided for a previous endoscopic image that is successively earlier than the respective endoscopic image, and when the tracked position of a region depicting at least one polyp in the respective endoscopic image is in a different position from the region output by the neural network for the previous endoscopic image, an augmented image is generated for the respective endoscopic image based on the tracked position.
[0012] A further embodiment of the second aspect further includes tracking a 3D position of the endoscopic camera and plotting the tracked 3D position of the endoscopic camera within the colon map GUI.
[0013] In a further embodiment of the second aspect, the forward tracking 3D position of the endoscopic camera is marked on the colon map presented in the GUI with a mark indicating that the forward direction of the endoscopic camera is going deeper into the colon, and the reverse tracking 3D position of the endoscopic camera presented on the colon map is marked with another mark indicating the reverse direction of the endoscopic camera being removed from the colon.
[0014] A further embodiment of the second aspect further includes providing each endoscopic image to a detection neural network, outputting by the detection neural network a representation of a region of the endoscopic image depicting at least one polyp, and calculating an estimated 3D position of the at least one polyp within the region of the endoscopic image according to the 3D reconstruction, and generating instructions for presenting the 3D position of the at least one polyp on a colon map within the GUI.
[0015] A further embodiment of the second aspect further comprises receiving an indication of surgical removal of at least one polyp from the colon and marking a 3D location of the at least one polyp on the colon map with the indication of removal of the at least one polyp.
[0016] A further implementation of the second aspect further includes tracking a 3D position of an endoscopic camera capturing a plurality of endoscopic images, calculating an estimated distance from the current 3D position of the endoscopic camera to a 3D position of at least one polyp previously identified using previously acquired endoscopic images, and generating instructions to present a display within the GUI if the estimated distance is below a threshold.
[0017] A further embodiment of the second aspect further includes analyzing each 3D reconstruction to estimate a portion of the inner surface of the colon depicted in each endoscopic image; tracking cumulative portions of the inner surface of the colon depicted in successive endoscopic images during a spiral scanning operation of the endoscopic camera during the colonoscopy procedure; and generating instructions for presenting in a GUI at least one of an estimate of remaining portions of the inner surface not yet depicted in previously captured endoscopic images and an estimate of total coverage of the area of the inner surface, wherein the analyzing, tracking, and generating are repeated during the spiral scanning operation.
[0018] In a further embodiment of the second aspect, each portion corresponds to a time window having an interval corresponding to the amount of time to cover the respective portion during the spiral scanning operation, and an indication of adequate coverage is generated when at least one image that largely depicts the respective portion is captured during the time window, and / or another indication of inadequate coverage is generated when no image that largely depicts the respective portion is captured during the time window.
[0019] In a further embodiment of the second aspect, an indication of the total amount of inner surface depicted in the image relative to the amount of inner surface not depicted is calculated by aggregating the portion covered during the spiral scanning movement relative to the portion not covered during the spiral scanning movement.
[0020] In a further embodiment of the second aspect, the 3D reconstruction neural network is trained with a training data set of 2D endoscopic image pairs defining input images that correspond to 3D coordinate values calculated for pixels of the 2D endoscopic images calculated by the 3D reconstruction process that defines the ground truth.
[0021] A further embodiment of the second aspect further includes receiving a representation of at least one anatomical landmark of the colon, the at least one anatomical landmark dividing the colon into multiple portions; tracking a 3D position of the endoscopic camera relative to the at least one anatomical landmark; calculating an amount of time spent by the endoscopic camera at each of the multiple portions of the colon; and generating instructions for presenting the amount of time spent by the endoscopic camera at each of the multiple portions of the colon in a GUI.
[0022] In a further embodiment of the third aspect, a representation of a region of the at least one 2D image depicting at least one polyp is output by a detection neural network trained to segment polyps in the 2D image.
[0023] In a further embodiment of the third aspect, the 3D reconstruction of the at least one 2D image is fed to a detection neural network in combination with the at least one 2D image to output a representation of a region depicting the at least one polyp.
[0024] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention applies. Although methods and materials similar or equivalent to those described herein may be used in the practice or testing of embodiments of the present invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. Furthermore, the materials, methods, and examples are illustrative only and are not intended to be necessarily limiting. [Brief explanation of the drawings]
[0025] Some embodiments of the present invention are described herein, by way of example only, with reference to the accompanying drawings, in which it is emphasized that the details shown, with particular reference to the details of the drawings, are by way of example and for illustrative discussion of embodiments of the present invention, and in this regard, the description taken together with the drawings will make apparent to those skilled in the art how embodiments of the present invention may be practiced.
[0026] [Figure 1] 1 is a flowchart of a method for processing images acquired by a camera positioned on an endoscope in the colon of a target patient, according to some embodiments of the present invention. [Figure 2] 1 is a block diagram of components of a system for processing images acquired by a camera placed on an endoscope in the colon of a target patient, according to some embodiments of the present invention. [Figure 3] 10 is a flowchart of a process for tracking the location of a region depicting one or more polyps in each endoscopic image relative to one or more previous endoscopic images, according to some embodiments of the present invention. [Figure 4]FIG. 2 is a schematic diagram illustrating an example of feature-based KD tree matching between two consecutive images according to some embodiments of the present invention; [Figure 5] 5 is a geometric transformation matrix between the frames of FIG. 4 according to some embodiments of the present invention. [Figure 6] 1A-1C are sequentially subsequent schematic diagrams showing a bounding box ROI indicating a tracked detected polyp according to some embodiments of the present invention, in which the tracked bounding box is no longer depicted in the image, and an arrow is added to indicate the direction in which the camera should be moved to re-depict the polyp ROI in the captured image. [Figure 7] 1A-1C are schematic diagrams depicting a detected polyp and successively later frames depicting a tracked polyp calculated using the transformation matrices described herein, in accordance with some embodiments of the present invention. [Figure 8] 1 is a flowchart of a process for 3D reconstruction of 2D images captured by a camera inside a patient's colon, according to some embodiments of the present invention. [Figure 9] 1 is a flowchart illustrating an exemplary 3D tracking process for tracking 3D movement of a camera, according to some embodiments of the present invention. [Figure 10] 1 is an example of a 3D rigid transformation matrix for tracking 3D movement of a colonoscopy camera, according to some embodiments of the present invention. [Figure 11] 1 is a schematic diagram of a mosaic and / or panoramic image of the colon being constructed for each pixel in the combined 2D images, where a 3D position in a single 3D coordinate system is calculated, according to some embodiments of the present invention; FIG. [Figure 12] 1A-1C are schematic diagrams depicting respective images presented within respective quarters of the inner surface of the colon depicted therein, according to some embodiments of the present invention. [Figure 13] 1 is a flowchart of a method for calculating a quarter represented by a frame according to some embodiments of the present invention. [Figure 14]1A-1C are schematic diagrams illustrating a process for calculating polyp volume from a 3D reconstructed surface calculated from 2D images, according to some embodiments of the present invention. [Figure 15] 1 is a flowchart of an exemplary process for calculating polyp volume from 2D images, according to some embodiments of the present invention. [Figure 16] 10A-10C are schematic diagrams illustrating a series of augmented images presented in a GUI according to generated instructions for tracking a polyp ROI, according to some embodiments of the present invention. [Figure 17] 1A-1C are schematic diagrams of colon maps presented in a GUI showing the trajectory of endoscope movement, the locations of detected polyps, removed polyps, and anatomical landmarks, according to some embodiments of the present invention. [Figure 18] 1A-1C are schematic diagrams depicting quadrants of the inner surface of the colon that are depicted in one or more images and / or quadrants of the inner surface of the colon that are not yet depicted in the images, according to some embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0027] The present invention, in some embodiments thereof, relates to colonoscopy, and more particularly, but not exclusively, to systems and methods for processing colon images and videos and / or processing colon polyps automatically detected during a colonoscopy procedure.
[0028] As used herein, the terms "image" and "frame" are sometimes interchangeable. Images captured by a camera of a colonoscope may be individual frames of a video captured by the camera.
[0029] As used herein, the terms "endoscope" and "colonoscope" are sometimes used interchangeably.
[0030] Some embodiments of the present invention relate to systems, methods, devices, and / or code instructions (i.e., stored on a memory and executable by one or more hardware processors) for generating instructions to present a graphical user interface (GUI) for dynamically tracking one or more polyps in two-dimensional (2D), optionally color, endoscopic images of a patient's colon captured by a camera of an endoscope positioned within the lumen of the colon, e.g., during a colonoscopy procedure. The location of a region depicting one or more polyps (e.g., a region of interest (ROI)) in a current endoscopic image is optionally tracked relative to one or more previous endoscopic images based on matching visual features between the current and previous images, e.g., features extracted based on a speed robust feature (SURF) process. Once the location of the ROI depicting the polyp is determined for the current image so that it is located outside the image boundary, a vector is calculated. The vector points from the location in the current image to the location of the ROI located outside the current image. The location in the current image may be, for example, the on-screen location of the ROI in a previous image in which the ROI was located. The augmented endoscopic image is created by augmenting a current endoscopic image with a representation of the vectors, for example, by introducing the representation of the vectors as a GUI element and / or as an overlay on the endoscopic image. Instructions are generated for presenting the augmented endoscopic image on a display within the GUI.
[0031] The vector representation depicts a direction and / or orientation for adjusting the endoscopic camera to capture another endoscopic image depicting the ROI of the polyp. The augmented endoscopic image is augmented with the vector representation for display within the GUI.
[0032] The vector representation may be an arrow pointing to the location of the ROI outside the image. Moving the camera in the direction of the arrow will restore the polyp in the image.
[0033] This process is repeated for each captured image to assist the operator in maintaining the polyp in the image. As the camera moves and the polyp no longer appears in the current image, the display of the vector instructs the operator as to how to manipulate the camera to recapture the polyp in the image.
[0034] Some embodiments of the present invention relate to a system, method, device, and / or code instructions (i.e., stored in a memory and executable by one or more hardware processors) that generate instructions for dynamically tracking the 3D movement of an endoscopic camera capturing endoscopic images within a patient's colon. The captured 2D images (e.g., each image, or every few images, e.g., every third, fourth, or other number of images) are fed to a 3D reconstruction neural network, optionally a convolutional neural network (CNN). The 3D reconstruction neural network outputs a 3D reconstruction of each 2D endoscopic image. Pixels of the 2D endoscopic images are assigned 3D coordinates. The current 3D position of the endoscopic camera (i.e., endoscope) within the colon is calculated according to the 3D reconstruction. For example, the 3D position of the endoscope is determined based on the values of the 3D coordinates of the current image. Instructions are generated to present the current 3D position of the endoscopic camera on a colon map within a GUI. The colon map depicts a virtual map of the patient's colon.
[0035] The 3D position of the endoscope can be tracked and plotted as a trajectory on a map of the colon, for example, to track the path of the endoscope within the colon during a colonoscopy procedure.
[0036] The forward and reverse directions of the endoscope may be marked, for example, by arrows and / or color coding.
[0037] The 3D location of the detected polyp can be marked on the colon map. The 3D location of the detected polyp can be tracked relative to the 3D location of the camera. If the distance between the camera and the polyp is less than a threshold, instructions can be generated to present an indication within the GUI. The indication can be, for example, a marking of the polyp if it is present in the image, an arrow pointing to the location of the polyp if it is not depicted in the image, and / or a message that the camera is in proximity to the polyp, optionally within instructions on how to move the camera to capture an image depicting the polyp.
[0038] Polyps surgically removed from the colon may be marked on a colon map.
[0039] Optionally, the portion of the inner surface of the colon depicted in the image is analyzed, for example, based on a virtual division of the inner surface into quarters. When a colonoscope is used to visually scan the inner wall of the colon, the extent of the portion is optionally tracked cumulatively, for example, in a spiral motion as the colonoscope is withdrawn from the colon (or moved forward within the colon). The spiral motion is performed, for example, by a clockwise (or counterclockwise) orientation of the camera as the colonoscope is slowly withdrawn (or pushed in). Alternatively, the inner portion of the colon is imaged incrementally, for example, by pulling back (or pushing forward) the colonoscope a certain distance, stopping the backward (or forward) movement of the camera, and imaging the periphery by pointing the camera in a circular (or transverse) pattern, where pulling back (or pushing forward), stopping, and imaging are repeated along the length of the colon. Optionally, when a colonoscope is used to visually scan the inner wall of the colon, each portion (e.g., quarter) is mostly represented by one or more images. Estimates of the depicted inner surface and / or the remaining inner surface (e.g., quarter) may be generated, and instructions for presentation within the GUI may be generated. The estimation may be performed in real time, for example, quarter by quarter, and / or as a global estimate for the entire colon (or most of it) based on the aggregation of coverage of each portion (e.g., quarter). Previously covered and / or remaining coverage of the inner surface of the colon helps ensure that the entire inner surface of the colon has been captured in the images, reducing the risk of missing polyps.
[0040] Optionally, based on 3D tracking of the colonoscope, an amount of time spent by the colonoscope in one or more defined portions of the colon is calculated. Instructions for presentation of time may be generated for presentation in the GUI, for example, the amount of time spent in each portion of the colon is presented on a corresponding portion of the colon map.
[0041] Aspects of some embodiments of the present invention relate to systems, methods, devices, and / or code instructions (i.e., stored on a memory and executable by one or more hardware processors) that generate instructions for calculating dimensions (e.g., size) of polyps. The dimensions may be 2D dimensions, e.g., area and / or radius of a flat polyp, and / or 3D dimensions, e.g., volume and / or radius of a protruding polyp. A representation of a region of a 2D image depicting a polyp may be manually drawn by an operator (e.g., using a GUI) and / or output by a detection neural network that is fed the 2D image and trained to segment polyps in the 2D image. The 2D image is fed to a 3D reconstruction neural network that outputs 3D coordinates for pixels of the 2D image. The dimensions of the polyp are calculated according to an analysis of the 3D coordinates of pixels of a region of interest (ROI) in the 2D image depicting the polyp.
[0042] Optionally, instructions are generated to present a warning in the GUI when a polyp size exceeds a threshold. The threshold may define a minimum size for a polyp to be removed. Polyps with sizes below the threshold may be left in place.
[0043] At least some implementations of the systems, methods, devices, and / or code instructions described herein relate specifically to the medical problem of treating patients for the identification and removal of polyps in their colons. Using standard colonoscopy techniques, adenomas may be missed in up to 20% of cases, and cancers may be missed in approximately 0.6% of cases, as evidenced by the eventual detection of these missed lesions during interval colonoscopies. Adenoma detection rates (ADRs) vary and depend on patient risk factors, physician performance, and instrumental limitations. The patient's individual anatomy and the quality of bowel preparation are important determinants of a high-quality colonoscopy. A physician's performance of a high-quality colonoscopy depends on factors such as successful cecal intubation, careful visual inspection during extended retraction times, and overall endoscopic experience. While endoscopist fatigue and inattention are risk factors for physicians missing polyps, an earlier start time in a session correlates with better outcomes. Of note, although adenoma detection rates (ADRs) increased in facilities that implemented quality improvement programs, awareness of monitoring or simply being observed had a more positive effect on ADRs.
[0044] At least some implementations of the systems, methods, devices, and / or code instructions described herein improve the rate of polyp detection and / or removal during colonoscopy procedures. The improvements are facilitated, at least in part, by the GUI described herein, which helps instruct the operator to (i) display previously identified polyps that have disappeared from the currently captured image by an arrow pointing in the direction of operation of the colonoscope camera to recapture an image of the polyp, (ii) present and update a colon map displaying the 2D and / or 3D locations of identified polyps to help ensure that all identified polyps are evaluated and / or removed, (iii) trace portions of the inner circumferential surface of the colon to identify portions of the inner surface that have not been captured by the image and therefore not analyzed to identify polyps to ensure that no portions of the colon are not imaged and polyps are missed, (iv) calculate the volume of polyps to help provide data to aid in which polyps to remove and / or diagnose cancer, and / or (v) calculate the amount of time spent by the colonoscope in each portion of the colon. The GUI is presented and adapted in real time to the images captured during the colonoscopy procedure, allowing real-time feedback to the physician operator in helping improve polyp detection and / or removal rates.
[0045] At least some implementations of the systems, methods, devices, and / or code instructions described herein address the technical challenge of devices that improve polyp detection and / or removal rates. Specifically, at least some implementations of the systems, methods, devices, and / or code instructions described herein improve image processing and / or GUI techniques, such as code that analyzes captured images and / or a GUI used by an operator, to help increase polyp identification and / or detection rates, compared to standard approaches. For example, optics that achieve a wider field of view and improve image resolution, and distal colonoscope attachments such as balloon caps or rings to improve visualization behind mucosal folds, are passive and rely on the skill of the operator to track identified polyps. In contrast, the GUI described herein automatically tracks identified polyps.
[0046] At least some implementations of the systems, methods, devices, and / or code instructions described herein address the technical problem of calculating polyp volume. Based on standard practice, polyp size is measured only after the polyp has been removed from the patient, as described, for example, with reference to Kume, Keiichiro, et al., "Endoscopic measurement of poly size using a novel calibrated hood," Gastroenterology Research and Practice 2014 (2014). The importance of measuring polyp size is described, for example, with reference to Summers, Ronald M., "Polyp size measurement at CT colonography: What do we know and what do we need to know?" Radiology 255.3 (2010):707-720. In contrast, at least some of the systems, methods, devices, and / or code instructions described herein calculate polyp size in vivo while the polyp is attached to the colon wall, before the polyp is removed. Calculating polyp volume before the polyp is removed can provide several advantages, for example, polyps above a threshold volume are targeted for removal and / or polyps below a threshold volume are left in the patient. The calculated polyp volume before removal may be compared to the post-removal volume, for example, to determine whether the entire polyp has been removed, and / or to compare the surface polyp volume with the unseen portion of the polyp below the surface as a cancer risk, and / or to help grade the polyp and / or cancer risk.
[0047] At least some implementations of the systems, methods, devices, and / or code instructions described herein address the technical challenge of neural network processing, which is slower than the rate of images captured in video by a colonoscope camera. The process described herein for tracking polyps (and / or associated ROIs) based on extracted features compensates for delays in the process for detecting polyps using a neural network, which outputs data for the image, optionally an indication of the detected polyp, and / or its location. The neural network-based detection process is computationally more expensive than the feature extraction and tracking process (e.g., 25 milliseconds (ms)-40 ms on a typical personal computer (PC) with an I7 Intel processor and an Nvidia GTK 1080 TI GPU, although the typical time difference between consecutive frames is in the range of 20 ms-40 ms). Therefore, a delay scenario can be created in which the detection result output by the neural network for a frame number indicated by i is ready only if a subsequent frame (e.g., a number indicated by i+2, or a later frame) has already been presented. Such a delay scenario can result in odd situations when an initial frame depicts a polyp but a later frame does not (e.g., the camera is shifted so that the polyp is not captured in the image), and delays in the neural network that detects polyps are only available when the polyp is no longer depicted, creating a situation where a display of a detected polyp is provided when the presented image does not depict the polyp. Note that when frame number i+2 is available for presentation, frames must be presented as soon as they are available (e.g., in real time and / or immediately) because delays in the presentation of frames are unacceptable from a clinical and / or regulatory perspective and could result in losses, for example, in attempts to remove the imaged polyp.Computationally efficient feature-based tracking, which results in rapid processing compared to neural network-based processing (e.g., less than 10 ms on a typical PC with an I7 Intel processor), is used to transform a ROI depicting a polyp detected in frame i to a position in frame i+2. The transformed position is the position presented on the display at frame number i+2. Optionally, the bounding box of the ROI (e.g., only the bounding box of the ROI) is transformed to the i+2 frame, since the contour transformation may be less accurate due to the effects of different 3D positions that are not necessarily taken into account in the 2D transformation. Note that the i+2 frame is one example, and other examples, such as, but not limited to, i+1, i+3, i+4, i+5, and more, may be used.
[0048] Before describing at least one embodiment of the present invention in detail, it is to be understood that the invention is not necessarily limited in its application to the details of construction and arrangement of components and / or methods set forth in the following description and / or illustrated in the drawings and / or examples. The invention is capable of other embodiments or of being practiced or carried out in various ways.
[0049] The present invention may be a system, a method, and / or a computer program product, which may include a computer-readable storage medium having computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0050] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPRM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, punch card, or mechanically encoded devices such as raised structures in grooves in which instructions are recorded, and any suitable combination of the above. As used herein, a computer-readable storage medium should not be interpreted as a transitory signal itself, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal transmitted through a wire.
[0051] The computer-readable program instructions described herein can be downloaded to each computing / processing device from a computer-readable storage medium or to an external computer or storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0052] Computer-readable program instructions for carrying out the operations of the present invention may be source or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the present invention.
[0053] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0054] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus for manufacturing a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other apparatus to function in a particular manner, and a computer-readable storage medium having stored instructions includes an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0055] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device to cause the computer, other programmable data processing apparatus, or other device to perform a series of operational steps to generate a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0056] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of an instruction, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative embodiments, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a special-purpose hardware-based system that performs predetermined functions or acts or executes a combination of special-purpose hardware and computer instructions.
[0057] Reference is now made to Figure 1, which is a flowchart of a method for processing images acquired by a camera located on an endoscope within a target patient's colon, tracking polyps, tracking camera movement, mapping the location of polyps, calculating coverage of the inner surface of the colon, tracking the amount of time spent in different portions of the colon, and / or calculating polyp volume, in accordance with some embodiments of the present invention. Reference is also made to Figure 2, which is a block diagram of components of a system 200 for processing images acquired by a camera located on an endoscope within a target patient's colon, tracking polyps, tracking camera movement, mapping the location of polyps, calculating coverage of the inner surface of the colon, tracking the amount of time spent in different portions of the colon, and / or calculating polyp volume, in accordance with some embodiments of the present invention. System 200 can optionally perform the operations of the method described with reference to Figure 1 via a hardware processor 202 of a computing device 204 executing code instructions stored in memory 206.
[0058] A camera disposed on the imaging probe 212, e.g., a colonoscope, captures images of the inside of the patient's colon, e.g., obtained during a colonoscopy procedure. The colon images are optionally 2D images and optionally color images. The colon images may be acquired as streamed video and / or a sequence of still images. The captured images may be processed in real time and / or offline (e.g., after the procedure is completed).
[0059] The captured images may be stored in an image repository 214, which may optionally be implemented as an image server, such as a Picture Archiving and Communication System (PACS) server and / or an Electronic Health Record (EHR) server. The image repository may be in communication with the network 210.
[0060] The computing device 204 receives captured images, for example, in real time directly from the imaging probe 212 and / or from the image repository 214 (e.g., in real time or offline). The real-time images may be received during the colonoscopy procedure to guide the operator, as described herein. The captured images may be received by the computing device 204 via one or more imaging interfaces 220, such as a wire connection (e.g., an output from the imaging probe 212 plugged into the imaging interface via a connecting wire), a wire connection (e.g., an antenna), a local bus, a port for connecting a data storage device, a network interface card, other physical interface implementation, and / or a virtual interface (e.g., a software interface, a virtual private network (VPN) connection, an application programming interface (API), a software development kit (SDK)).
[0061] The computing device 204 analyzes the captured image as described herein and generates instructions to dynamically adjust the graphical user interface presented on the user interface (e.g., display) 226, e.g., elements of the GUI are injected as an overlay onto the captured image and presented on the display as described herein.
[0062] Computing device 204 may be embodied as, for example, a dedicated device, a client terminal, a server, a virtual server, a colonoscopy workstation, a gastrointestinal epidemiology workstation, a virtual machine, a computing cloud, a mobile device, a desktop computer, a thin client, a smartphone, a tablet computer, a laptop computer, a wearable computer, a glasses computer, and a watch computer. Computing 204 may include gastrointestinal epidemiology and / or colonoscopy workstations and / or other devices for allowing an operator to view GUIs created from the processing of colonoscopy images, for example, real-time presentations that point arrows toward currently unseen polyps on the image, and / or colon maps that present 2D and / or 3D locations of polyps, and / or other features described herein.
[0063] The computing device 204 may perform one or more operations depicted with reference to FIG. 1 and / or operate as one or more servers (e.g., network server, web server, computing cloud, virtual server) and may include locally stored software that provides services (e.g., one or more of the operations described with reference to FIG. 1) to one or more client terminals 208 (e.g., client terminals used by users to view colonoscopy images, e.g., colonoscopy workstations including a display presenting images captured by the colonoscope 212, remotely located colonoscopy workstations, PACS servers, remote EHR servers, remotely located displays for remote viewing of the procedure by trainees, etc.). The service may provide software as a service (SaaS) to the client terminal 208 over the network 210, for example, provide applications and / or colonoscopy applications to the client terminal 208 for local download as add-ons to a web browser, and / or provide functionality to the client terminal 208 using a remote access session, such as via a web browser, an application programming interface (API), and / or a software development kit (SDK), for example, for injecting GUI elements into colonoscopy images and / or presenting colonoscopy images within a GUI.
[0064] Different architectures of the system 200 may be implemented, for example: *The computing device 204 is connected between the imaging probe 212 and the display 226 and is, for example, a component of a colonoscopy workstation. Such an implementation may be used for real-time processing of images captured by the colonoscope during a colonoscopy procedure and real-time presentation of a GUI described herein on the display 226, e.g., for injecting GUI elements and / or presenting images within the GUI, e.g., presenting directional arrows to currently unseen polyps and / or presenting a colon map showing the 2D and / or 3D locations of polyps, and / or other features described herein. In such an embodiment, a computing device 204 may be installed for each colonoscopy workstation (e.g., including the imaging probe 212 and / or display 226). *The computing device 204 acts as a central server, servicing multiple colonoscopy workstations, such as client terminals 208 (e.g., including imaging probes 212 and / or displays 226) over the network 210. In such an embodiment, a single computing device 204 may be installed to service multiple colonoscopy workstations. *The computing device 204 can be installed as code on an existing device such as a server 218 (e.g., a PACS server, an EHR server) to provide local offline processing for each device, e.g., offline analysis of colonoscopy videos captured by different operators and stored on the PACS and / or EHR server. The computing device 204 can be installed on an external device that communicates with the server 218 (e.g., a PACS server, an EHR server) via the network 210 to provide local offline processing for multiple devices.
[0065] The client terminal 208 may be implemented as a colonoscopy workstation, which may include, for example, an imaging probe 212 and a display 226, a desktop computer (e.g., running a viewer application for viewing colonoscopy images), a mobile device (e.g., laptop, smartphone, eyeglasses, wearable device), and a remote station server for remotely viewing colonoscopy images.
[0066] The hardware processor 202 may be implemented as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), a digital signal processor (DSP), and an application-specific integrated circuit (ASIC). The processor 202 may include one or more processors (homogeneous or heterogeneous) that may be arranged for parallel processing, as a cluster, and / or as one or more multi-core processing units.
[0067] Memory 206 (also referred to herein as program storage and / or data storage) stores code instructions for execution by hardware processor 202, e.g., random access memory, read-only memory, and / or storage devices, e.g., non-volatile memory, magnetic media, semiconductor memory devices, hard drives, removable storage, and optical media (e.g., DVD, CD-ROM). For example, memory 206 may store code 206A that implements one or more operations and / or features of the method described with reference to FIG. 1 and / or GUI code 206B that generates instructions for presenting within a GUI and / or presents a GUI described herein based on the instructions (e.g., injecting GUI elements into a colonoscopy image, overlaying GUI elements on an image, presenting a colonoscopy image within a GUI, and / or presenting and dynamically updating a colon map described herein).
[0068] Computing device 204 may include data storage 222 for storing data, such as received colonoscopy images, colon maps, and / or processed colonoscopy images presented in the GUI. Data storage 222 may be embodied, for example, as memory, a local hard drive, a removable storage device, an optical disk, a storage device, and / or a remote server and / or a computing cloud (e.g., accessed via network 210).
[0069] The computing device 204 may include one or more of a data interface 224, and optionally a network interface, such as a network interface card, a wireless interface for connecting to a wireless network, a physical interface for connecting to a cable for network connectivity, a virtual interface implemented in software, network communication software providing a higher layer of network connectivity, and / or other implementations, for connecting to the network 210. The computing device 204 may access one or more remote servers 218 using the network 210, for example, to download updated imaging processing code, updated GUI code, and / or to obtain images for offline processing.
[0070] It is noted that the imaging interface 220 and the data interface 224 may be implemented as a single interface (e.g., a network interface, a single software interface), and / or as two independent interfaces, such as a software interface (e.g., an API, a network port) and / or a hardware interface (e.g., two network interfaces), and / or a combination thereof (e.g., a single network interface and two software interfaces, two virtual interfaces on a common physical interface, a virtual network on a common network port). The term / component imaging interface 220 may sometimes be interchanged with the term / data interface 224.
[0071] The computing device 204 may communicate with one or more of the server 218, the imaging probe 212, the image repository 214, and / or the client terminal 208 using the network 210 (or another communication channel), such as via a direct link (e.g., cable, wireless) and / or an indirect link (e.g., via an intermediate computing device such as a server and / or via a storage device), for example, according to different architecture implementations described herein.
[0072] The imaging probe 212 and / or computing device 204 and / or client terminal 208 and / or server 218 include or communicate with a user interface 226 that includes mechanisms designed for a user to input data (e.g., marking polyps for removal) and / or view a GUI including colonoscopy images, directional arrows, and / or a colon map. Exemplary user interfaces 226 include, for example, one or more of a touchscreen, a display, a keyboard, a mouse, augmented reality glasses, and voice-activated software using a speaker and microphone.
[0073] At 100, the endoscope is inserted and / or moved into the patient's colon, for example, by advancing the endoscope forward (i.e., from the rectum to the cecum), retracting it (e.g., from the cecum to the rectum), and / or adjusting the orientation of at least the camera of the endoscope (e.g., up, down, left, right), and / or leaving the endoscope in place.
[0074] It should be noted that the endoscope may be adjusted based on the GUI, e.g., manually by an operator and / or automatically by a user, e.g., the user may adjust the endoscope camera according to the presented arrows to recapture a polyp that has moved out of the image.
[0075] At 102, an image is captured by the camera of the endoscope. The image is a 2D image, optionally in color. The image depicts the inside of the colon and may or may not depict polyps.
[0076] The images can be captured as a video stream, and individual frames of the video stream can be analyzed.
[0077] The images may be analyzed individually and / or as a set of consecutive images as described herein. Each image in the sequence may be analyzed, or some intermediate images may be ignored, optionally a predetermined number, e.g., every third image, is analyzed and the middle two images are ignored.
[0078] Optionally, one or more polyps depicted in the image are treated. The polyps may be treated via an endoscope. The polyps may be treated by their surgical removal, for example, for sending to a pathology laboratory. The polyps may be treated by ablation.
[0079] Optionally, treated polyps are marked, for example, manually by a physician (e.g., by making a selection using the GUI, by pressing a "polyp removal" icon) and / or automatically by a code (e.g., detecting movement of a surgical resection device). As described herein, marked treated polyps may be tracked and / or presented on a colon map presented in the GUI.
[0080] At 104, the image is provided to a detection neural network that outputs an indication of whether a polyp is depicted in the image (or not). The detection neural network may include a segmentation process that identifies the location of detected polyps within the image, for example, by generating bounding boxes and / or other contours that depict the polyps within the 2D frame.
[0081] An exemplary neural network-based process for polyp detection is the Automated Polyp Detection System (APDS) described with reference to International Patent Application Publication No. WO2017 / 042812, entitled "A SYSTEM AND METHOD FOR DETECTION OF SUSPICIOUS TISSUE REGIONS IN AN ENDOSCOPIC PROCEDURE," by the same inventors as the present application.
[0082] The automated polyp detection process performed by the detection neural network may be executed in parallel and / or independently of features 106-114, for example, on the same computing device and / or processor and / or on another real-time connected computing device and / or platform connected to the computing device that executes the features described with reference to 106-114.
[0083] The output of the neural network is calculated and provided for each endoscopic image, and if the tracked position (i.e., as described with reference to 106) of the region depicting the polyp in each endoscopic image is in a different position from the region output by the neural network for the previous endoscopic image, an augmented image is generated for each endoscopic image based on the calculated tracked position of the polyp. This situation occurs when the image frame rate is faster than the processing rate of the detection neural network. The detection neural network completes image processing after one or more consecutive images are captured. In such a case, if the results of the detection neural network are used, the polyp position calculated for the older image may not necessarily reflect the polyp position for the current image.
[0084] Optionally, one or more endoscopic images of a successive subset of endoscopic images, including the respective endoscopic image and one or more images located consecutively before the respective endoscopic image (e.g., captured before the respective endoscopic image) depicting the tracked ROI (as described with reference to 106), are provided to the detection neural network. The successive subset of images can be provided to the detection neural network in parallel with the tracking process (as described with reference to 106). Alternatively, the subset of images is first processed by the tracking process, as described with reference to 106. One or more of the post-processed images may be translated and / or rotated to generate a subset of endoscopic images in which a region depicting a polyp (e.g., ROI) detected by the tracking process is in the same approximate position in all images, e.g., at the same pixel location on the display for all images. The processed successive subset of images is provided to the detection neural network to output a current region depicting the polyp. The images may be augmented with the region detected by the detection neural network.
[0085] Alternatively or additionally, the calculated tracking positions of the regions depicting the polyp in the image and / or the output of the tracking process described with reference to 106 (e.g., a 2D transformation matrix between successive frames) are provided to a detection neural network. The tracked positions and / or the 2D transformation matrix may be provided to the neural network when the tracked positions of the polyp are within the image or when the tracked positions of the polyp are located outside the image. The tracked positions may be provided to the neural network alone or in addition to one or more images (e.g., the current image and / or a previous image). The output of the tracking process (e.g., the tracked positions and / or the 2D transformation matrix) may be used by the neural network process, for example, to improve the accuracy of correlation between polyp detections in successive frames. The output of the tracking process may increase the reliability of the neural network-based polyp detection process (e.g., when a tracked polyp was detected in a previous frame) and / or reduce cases of false positive detections.
[0086] Optionally, in case of a discrepancy between the tracked position of the polyp (e.g., ROI depicting the polyp) calculated based on 106 and the output of the detection neural network, the position of the neural network is used. The position output by the neural network is used to generate instructions for generating an augmented image augmented with an indication of the position of the polyp. The position of the polyp output by the neural network may be considered more reliable than the tracked position calculated as described with reference to 106, but the position calculated by tracking may be computationally more efficient and / or may take less time than processing by the neural network.
[0087] Optionally, the 3D reconstruction of the 2D image described with reference to 108 is fed, alone and / or in combination with the 2D image, to a detection neural network to output a representation of a region depicting at least one polyp.
[0088] Referring back to FIG. 1, at 106, the location of a region (eg, ROI) depicting one or more polyps is tracked within the current endoscopic image relative to one or more previous endoscopic images.
[0089] Note that camera motion (e.g., facing, forward, backward) is indirectly tracked by tracking the movement of the ROI between images, since the camera is moving while the polyps remain stationary in their positions within the colon. Note that some movement of the ROI between frames may be due to peristalsis and / or other natural movements of the colon itself, regardless of whether the camera is stationary or moving.
[0090] Optionally, the location is tracked in 2D. Polyps can be tracked by tracking a ROI that outlines the polyp. The ROI and / or polyp may be detected in one or more previous images by a feature 104 detection neural network.
[0091] The location of the polyp is tracked even if the polyp is not depicted in the current image, for example, the camera is positioned such that the polyp is no longer present in the images captured by the camera.
[0092] Optionally, a vector is calculated from a position in the current image to the position of the polyp and / or ROI located outside the current image. The vector may be calculated, for example, from the position of the ROI on the last (or previous) image that depicted the ROI, from the center of the screen, from the center of the quadrant of the screen closest to the position of the outer ROI, and / or from another region of the image closest to the position of the outer ROI (e.g., a predetermined distance away from a position at the border of the image closest to the position of the outer ROI).
[0093] The vector representation may depict a direction and / or orientation for adjustment of the endoscopic camera to capture another endoscopic image depicting a region of the image.
[0094] Optionally, the tracking algorithm is feature-based. Features can be extracted from an analysis of endoscopic images. Features may be extracted based on the speed-robust feature (SURF) extraction approach described with reference to Bay, Herbert, Tinne Tuytelaars, and Luc Van Gool, "Surf: Speeded up robust features," European conference on computer vision, Springer, Berlin, Heidelberg, 2006. Tracking of extracted features may be performed in two dimensions (2D) between consecutive images (note that one or more intermediate images between analyzed images may be skipped, i.e., ignored). Features can be matched based on, for example, the KD-tree approach described with reference to Silpa-Anan, Chanop, and Richard Hartley, "Optimized KD-trees for fast image descriptor matching (2008):1-8." The KD-tree approach may be selected, for example, based on the observation that the dominant motion in a colonoscopy procedure is the movement of the endoscopic camera through the colon. Features may be matched between successive images according to their descriptors.The best homography is estimated by using, for example, the Random Sample Consensus (RANSAC) method described with reference to Vincent, Etienne, and Robert Laganiere, "Detecting planar homographies in an image pair. ISPA 2001. Proceedings of the 2nd International Symposium on Image and Signal Processing and Analysis. In conjunction with the 23rd International Conference on Information Technology Interfaces (IEEE Cat. IEEE, 2001)," or by computing the 2D affine transformation matrix of the camera motion (and its closest 2D geometric transformation matrix) from frame to frame, as described with reference to Agarwal, Anubhav, C.V., Jawahar, and P.J. Narayanan, "A survey of planar homography estimation techniques." Centre for Visual Information Technology, Tech. Rep. IIIT / TR / 2005 / 12 (2005)."
[0095] A particular region of interest (ROI) can be tracked, and optionally a bounding box indicating the location of one or more polyps therein. The position of the ROI is tracked while the ROI moves out of the image (e.g., video) frame. The ROI may be tracked until the ROI returns (i.e., is redrawn) in the current frame. The ROI may be tracked continuously across image frame boundaries, for example, based on a tracking coordinate system defined externally to the image frame.
[0096] During the time interval between when a polyp is detected (e.g., automatically by the code and / or manually by the operator) and when camera movement is stopped (e.g., the operator focuses their attention on the detected polyp), the polyp may move out of the frame. Instructions may be generated to present an indication of where the polyp (or ROI associated with the polyp) is located outside the currently presented image, for example, by augmenting the current image with the presentation of a directional arrow. The arrow indicates where the operator should move the camera (i.e., the endoscope tip) to bring the polyp (or polyp ROI) back, depicted again in the new frame.
[0097] Optionally, the arrow (or other indication) may be presented until the camera is too far from the tracked polyp (e.g., greater than a defined threshold). The predetermined threshold may be defined, for example, as a distance of the polyp's new position from the center of the current frame that is greater than three times the diagonal length of the frame (in pixels), or as some other value.
[0098] Referring now to FIG. 3, FIG. 3 is a flowchart of a process for tracking the location of a region depicting one or more polyps in each endoscopic image relative to one or more previous endoscopic images, according to some embodiments of the present invention.
[0099] At 302, a new (i.e., current) image, optionally a new frame of video captured by the colonoscope camera, is received. The image may or may not depict one or more polyps being tracked.
[0100] At 304, irrelevant or misleading portions are removed from the image, for example, peripheries that are not part of the colon because they rely on a light source that moves with the camera, the lumen which is a dark area in the colon image where light does not return to the camera (e.g., representing the far central portion of the colon), and / or inconsistent reflections in successive frames.
[0101] At 306, a contrast-limited adaptive histogram equalization (CLAHE) process (e.g., as described with reference to Reza, Ali M., "Realization of the contrast-limited adaptive histogram equalization (CLAHE) for real-time image enhancement. Journal of VLSI signal processing systems for signal, image and video technology 38.1 (2004):35-44) and / or other processes are implemented to enhance the image. A speed-robust feature (SURF) technique is optionally implemented to extract features immediately after the CLAHE (e.g., without significant delay).
[0102] At 308, if the current frame is the first frame in the sequence and / or the first frame in which a polyp is detected, 302 is repeated to obtain the next frame. If the current frame is not the first frame, 310 is performed to process the current frame using the previous frame processed in the previous iteration.
[0103] At 310, a geometric matrix is calculated for the camera movement between two consecutive frames (e.g., frame number denoted i and frame number denoted i+n, i.e., a jump of n frames between the two frames that are processed together), optionally based on a homography.
[0104] At 312, if a transformation is not found, the next image (e.g., frame) in the sequence is processed with the previous (i.e., current) frame (i.e., when frames i and i+n do not match and an attempt is made to match frames i and i+n+1) (by repeating 302), otherwise 314 is executed.
[0105] If 314 cannot find a transformation between two or more previous frames (i.e., the position of the tracked object (i.e., polyp) in the last frame is not calculated and / or cannot be calculated), then 318 is performed, otherwise, if the position of the polyp is calculated, then 316 is performed.
[0106] At 316, once an object (i.e., a polyp) is identified in the image for tracking, the tracked region is fed forward into the current frame according to the transformation found at 310. 322 is implemented to generate instructions for depicting the object (i.e., a polyp) in the current frame, for example, by a visual marking such as a box.
[0107] Alternatively, at 318, the transform of the intermediate frame for which no transform was found is interpolated according to the last transform found between the two frames before and after the intermediate frame.
[0108] At 320, if a tracked object (i.e., a polyp) is present, the tracked region is fed forward to intermediate frames according to the interpolation transformation found at 318. 322 is implemented to generate instructions for rendering the object (i.e., the polyp) in these frames. 316 may also be implemented to render the object (i.e., the polyp) on the current frame.
[0109] The process is repeated to track when the object (i.e., polyp) is outside the frame at 324. The next frame to process may be a jump step, for example, ignoring two or more frames until a new frame is processed, or until each frame is processed.
[0110] Referring now to FIG. 4, FIG. 4 is a schematic diagram illustrating an example of KD tree matching between two consecutive images 402A-B matched together by corresponding features (features shown as plus signs, one labeled 404, matching features marked using lines, one shown as 406) for tracking polyps within regions 408A-B, according to some embodiments of the present invention. Image 402A can be denoted as frame number i. Image 402B can be denoted as frame number i+3. Region E08A indicates the original bounding box of the region, and region E08B indicates the transformation of E08A according to a transformation matrix calculated by a homography of the match between keypoints, as described herein.
[0111] Reference is now made to FIG. 5, which is a geometric transformation matrix 502 between frames 402A-B of FIG. 4, according to some embodiments of the present invention.
[0112] Referring now to FIG. 6, FIG. 6 shows a schematic diagram 602 depicting a bounding box ROI 604 indicating a detected polyp being tracked, in accordance with some embodiments of the present invention, and a successively later schematic diagram 606 (e.g., two frames later) in which the tracked bounding box is no longer depicted in the image, augmented with the presentation of an arrow 608 indicating the direction in which the camera should be moved to re-depict the polyp ROI in the captured image.
[0113] Referring now to FIG. 7, FIG. 7 shows a schematic diagram 702 illustrating a detected polyp 704 and a successively later frame (i.e., three frames later in time) 706 illustrating a tracked polyp 708 calculated using the transformation matrix described herein, in accordance with some embodiments of the present invention.
[0114] Referring back to Figure 1, a 3D reconstruction is calculated for the current 2D image at 108. The 3D reconstruction may define 3D coordinate values for the pixels of the 2D image.
[0115] Optionally, the 3D reconstruction neural network is trained to output a 3D image from an input of a 2D image. The trained 3D reconstruction neural network generates a more accurate 3D image from the 2D image compared to a 3D reconstruction process (e.g., a standard 3D reconstruction process) alone. The 3D reconstruction process is a process different from the 3D reconstruction neural network. The 3D reconstruction process may be based on a standard 3D reconstruction process that calculates a 3D reconstruction using only a single 2D image, for example. The 3D reconstruction neural network is trained using a training dataset of 2D images captured by a colonoscopy camera and corresponding 3D images created from the 2D images by the 3D reconstruction process. The 2D image is designated as an input, and the reconstructed 3D image is designated as ground truth.
[0116] Optionally, the process of computing a 3D reconstruction of the current 2D image is based on a 3D neural network that outputs the 3D reconstruction. The 3D reconstruction neural network may be trained using a pair of training datasets of 2D endoscopic images that define the input image and corresponding 3D coordinate values calculated for pixels of the 2D endoscopic images calculated by the 3D reconstruction process that define the ground truth. A neural network that may be trained on a large training dataset (e.g., on the order of 10,000-100,000 or 100,000-1,000,000) may provide increased accuracy compared to the 3D reconstruction process alone.
[0117] The neural network may be trained using pairs of in-vivo images of the colon captured by an endoscopic camera (denoted as input) and corresponding 3D values estimated by a 3D reconstruction process (denoted as ground truth). The 3D values may be 3D coordinates calculated for 2D pixels of the image. The 3D reconstruction process may be trained using a large dataset (e.g., at least 100,000 images of the colon from at least 100 different colonoscopy videos, or other smaller or larger values).
[0118] An exemplary process for 3D reconstruction of 2D colon images is described. The 3D geometry of a partial internal surface region of the colon may be reconstructed for each frame using a "shape-from-shading (SfS)" process, as described, for example, in Zhang, Ruo, et al., "Shape-from-shading: a survey," IEEE transactions on pattern analysis and machine intelligence 21.8 (1999):690-706, and / or Prasath, V.B. Surya, et al., "Mucosal region detection and 3D reconstruction in wireless capsule endoscopy videos using active contours," 2012 Annual International Conference of the IEEE Engineering in Medicine and Biology Society IEEE, 2012. The camera motion parameters can be calculated based on the Shape from Motion (SfM) process described, for example, in Szeliski, Richard, and Sing Bing Kang, "Recovering 3D shape and motion from image streams using nonlinear least squares," Journal of Visual Communication and Image Representation 5.1 (1994):10-28. Optionally, the 3D positions of one or more feature points can be calculated to integrate the subsurface reconstructed by the SfS algorithm. The SfS algorithm handles moving local light and light attenuation. Real-world conditions during endoscopic examinations within human organs may be mimicked. One exemplary advantage of the 3D reconstruction process described herein compared to other SfS-based processes is that the 3D reconstruction process described herein can calculate an unambiguous reconstruction surface for each frame.The 3D reconstruction process can use an intensity threshold to remove non-Lambertian (e.g., more specular) regions, allowing the SfS process to work on other regions. The SfS process implemented by the inventors (described with reference to Prados, E., Faugeras, O., et al., "Shape from shading: a well-posed problem?" IEEE Conference on Computer Vision and Pattern Recognition, pp. 870-877 (2005)) produces a well-defined surface. A well-defined surface has a resolution of 1 / r. 2 can be calculated by considering the light attenuation of λ and / or by assuming that a spot light source is mounted at the center of the camera's projection, and therefore the image brightness is
number
[0119] The described 3D reconstruction process was applied to colonoscopy video frames obtained according to Kaufman A., Wang J. (2008) 3D Surface Reconstruction from Endoscopic Videos. In: Linsen L., Hagen H., Hamann B. (eds) Visualization in Medicine and Life Sciences. Mathematics and Visualization. Springer, Berlin, Heidelberg, with an average reprojection error of 0.066 pixels for selected feature points.
[0120] The standard 3D process described herein can be used to calculate an estimate of the 3D reconstruction of the colon for each 2D frame. The 2D frames and the corresponding calculated ground truth (i.e., the 3D reconstruction 3D values of the 2D frames) can be divided to create a training dataset (e.g., 70% of the data) and a testing dataset (e.g., 30% of the data). For example, a convolutional neural network (CNN) based on an implementation similar to an encoder-decoder architecture (e.g., as described with reference to Shan, Hongming, et al., "3-D convolutional encoder-decoder network for low-dose CT via transfer learning from a 2-D trained network," IEEE Transactions on Medical Imaging 37.6(2018): 1522-1534) can be trained and tested with training and validation datasets to be robust enough to predict the 3D reconstruction for each 2D frame.
[0121] In an exemplary architecture, the encoder and decoder portions of the 3D reconstruction CNN may be constructed from 2D convolutional layers. The final layer may be multiplied three times to predict the 3D coordinate value for each pixel in the input image. The 3D reconstruction CNN can be used in real time to predict the 3D value of each pixel in each 2D frame of the colon surface in a colonoscopy video. The CNN is trained on the data output by the 3D reconstruction process, optionally using a large number of processed 2D images (e.g., on the order of 100,000, or more or less), but may be more robust and / or accurate than the 3D reconstruction process alone.
[0122] Referring now to Figure 8, Figure 8 is a flowchart of a process for 3D reconstruction of 2D images captured by a camera within a patient's colon, according to some embodiments of the present invention. The 3D reconstruction may be defined as ground truth values associated with the 2D images defined as input values to create a training data set for training a 3D reconstruction neural network, as described herein.
[0123] At 802, the images are optionally provided as a video stream of frames.
[0124] The current frame, denoted n, is processed at 804. Optionally, an SfS process for non-specular regions is used to compute a 3D reconstruction of frame n.
[0125] At 806, multiple images are processed, including images acquired before and after the current image. The frames may be denoted as n-2 through n+2, or other numbers may be used. An SfM process may be used to compute a generic 3D reconstruction of frame #n based on frames #n-2 through #n+2, and to compute ongoing updated features based on estimation of intrinsic camera parameters.
[0126] At 808, the 3D reconstructions calculated in I04 and I06 are integrated for frame n.
[0127] At 810, a 3D reconstruction is provided for each pixel in the 2D frame #n representing the observed colon region. Each pixel in the 2D frame #n is assigned an (x, y, z) value in a 3D coordinate system. The origin of the 3D coordinate system can be correlated to the center point of the frame.
[0128] Next, we describe an exemplary data flow for a 3D reconstruction CNN. In the first phase shown, consecutive 2D frames (i.e., the number of consecutive input frames, represented by n, an odd integer of at least 3) are fed into the 3D CNN and passed through layers (n × 3 × 3) of 3D convolution kernels. The first layer is designed to process all three color channels of the 2D frame by replicating them three times. The data flow continues through batch normalization and rerouting layers, followed by a max-pooling layer, which reduces the size of the output by two in each frame dimension. The data then passes through an additional four layer groups, but using 2D convolution kernels, representing the encoder portion of the 3D CNN. The data flow continues through an additional four layer groups with 2D convolution kernels, but using an upsampling layer instead of a max-pooling layer, and then through a fifth layer, replicated three times, representing the encoder portion of the CNN. The output is at the same resolution as the input 2D frame. For each 2D pixel, three values are output: the 2D pixel's (x, y, z) 3D coordinate values.
[0129] Optionally, the calculated 3D values of the pixels of the 2D frame may be fed to a polyp detection process (e.g., an APDS detection process) described with reference to 104. The calculated 3D values provide additional input (e.g., to a neural network) for detecting polyps in the current frame and / or determining the location of polyps in the current frame. For example, the 3D values may be information indicative of the flatness and / or protrusion of a colon tissue region suspected of being a polyp relative to its surroundings.
[0130] Referring back to Figure 1, the current 3D position of the polyp and / or endoscope is calculated and / or tracked at 110. The 3D position may be calculated based on the output of the 3D reconstruction output by the 3D reconstruction neural network.
[0131] The 3D position may be calculated based on a 3D rigid body transformation matrix calculated between successive 2D endoscopic images based on the 3D coordinates of matched features extracted from successive 2D endoscopic images and / or from 3D reconstructions of the successive 2D endoscopic images. In other words, the 3D transformation matrix is calculated for the matched features using the 3D coordinates of the pixels corresponding to the features, similar to the process described herein, for example, using the 3D coordinates of the features rather than the 2D coordinates to track 2D images based on corresponding features between 2D images. The 3D body transformation matrix indicates the movement of the endoscopic camera between successive 2D endoscopic images. The current 3D position of the endoscopic camera is calculated according to the 3D rigid body transformation matrix taking each 2D endoscopic image into account. A trajectory of the camera movement may be calculated based on the successive 3D transformation matrices. The trajectory of the camera movement may be presented on a colon map, as described herein.
[0132] Optionally, the 3D position of the detected polyp and / or endoscope is iteratively tracked according to the 3D positions calculated for multiple successive images.
[0133] The 3D tracking process may be based on the output of the 2D tracking process described with reference to 106 and / or the 3D reconstruction process described with reference to 108. The feature extraction and / or matching can be performed as described with reference to 2D tracking on 2D consecutive frames of the colonoscopy video. However, note that the homography is calculated for the 3D values (x, y, z) of the extracted features (keypoints) rather than for their 2D values. The 3D values are reconstructed by the 3D reconstruction process described herein to find the best-fit 3D affine transformation. The 3D rigid transformation that is closest to the calculated 3D affine transformation is calculated using, for example, the process described with reference to Yuan, Jie et al., "Application of Feature Point Detection and Matching in 3D Objects Reconstruction," PATTERNS 2011: The Third International Conference on Pervasive Patterns and Applications, pp. 19-24. The 3D rigid transformation matrix describes the 3D movement of the camera within the colon from frame to frame.
[0134] Reference is now made to FIG. 9, which is a flowchart illustrating an exemplary 3D tracking process for tracking 3D movement of a camera, according to some embodiments of the present invention.
[0135] At 1102, matched features between frame n and the next analyzed consecutive frame n+i are provided, for example, as an output of process execution 510 described with reference to FIG.
[0136] At 1104, the calculated 3D coordinate values of the matched features (keypoints) between frames n and n+i are optionally provided from a 3D reconstruction process (e.g., output by a 3D reconstruction neural network) as described herein.
[0137] At 1106, a 3D holography is calculated according to the 3D coordinate values (1104) of the matched features (1102).
[0138] At 1108, the closest 3D rigid transformation matrix is found.
[0139] At 1110, an affine 3D transformation matrix is calculated.
[0140] At 1112, a 3D rigid transformation matrix is calculated that describes the camera motion from frame n to frame n+i.
[0141] Reference is now made to FIG. 10, which is an example of a 3D rigid transformation matrix 1202 for tracking 3D movement of a colonoscopy camera, according to some embodiments of the present invention.
[0142] Referring back to FIG. 1, a 3D tracking algorithm allows for the construction of a 3D trajectory of camera movement (e.g., during a colonoscopy procedure) and is calculated based on 3D tracking of the camera's position. The 3D trajectory can be defined according to a 3D coordinate system, e.g., the origin correlates to the camera position when tracking began (e.g., when the endoscope just entered the colon, e.g., point [0,0,0] in the 3D coordinate system from which the 3D values of the first frame in the colon were reconstructed). The trajectory can be constructed by calculating the camera's 3D position (e.g., its principal point) for each new frame (and / or multiple new frames, e.g., 1-25 new frames) relative to the camera's 3D position for the last frame for which the 3D position was calculated. The calculation of the new 3D position can be derived (e.g., immediately, without significant delay, during the time it takes for the current frame to be presented on the display) from a 3D rigid body transformation matrix between two frames (e.g., by multiplying the newly calculated transformation matrix by the camera's previously calculated 3D position).
[0143] Optionally, the 3D movement of the camera is tracked relative to the 3D position(s) of anatomical landmarks and / or the 3D position(s) of previously detected polyps. The anatomical landmarks may be predefined and / or set by the operator, and may be, for example, the position of the anus, the position of the cecum, anatomical abnormalities, and / or portions of the colon (e.g., transverse, ascending). Polyps are detected automatically and / or manually during camera advancement and removed during camera retraction.
[0144] At 112, the portions of the inner surface of the colon depicted in the image are calculated, with each portion defined as, for example, one-third, one-quarter, one-eighth (or other division) of the circumference of the inner surface of the colon. The coverage of each portion may be dynamically calculated in real time, for example, regardless of whether the currently captured image depicts (e.g., mostly depicts) the respective portion. For example, coverage can be calculated for the entire colon (or portion thereof) as the collection of coverages of the individual portions, providing a result such as approximately 94% of the inner surface of the colon being covered and / or approximately 6% of the inner surface of the colon not being covered.
[0145] Optionally, the representation of the total amount of inner surface depicted in the image relative to the amount of inner surface not depicted is calculated by aggregating the portion covered during the spiral scanning motion relative to the portion not covered during the spiral scanning motion.
[0146] The portion of the inner surface of the colon depicted in an endoscopic image may be calculated based on, for example, an analysis of a 3D reconstruction of the image and / or based on an analysis of the image itself, for example, following identification of the location of the lumen within the image. The lumen may be identified, for example, as a region of pixels having intensity values below a threshold indicative of darkness (i.e., a lumen that does not reflect a light source to the camera).
[0147] Optionally, multiple sequential images are aggregated to generate a single mosaic image. Portions of the inner surface of the colon may be calculated for the single mosaic image and / or for each individual image. Multiple sequential images may be taken at different camera orientations, for example, when the camera is oriented using a clockwise, counterclockwise, x-pattern, or other motion. The camera may remain stationary at the same position along the longitudinal axis of the colon, or may be moved simultaneously, for example, when the camera is pushed forward and / or retracted, for example, while being oriented in a spiral pattern.
[0148] The cumulative portion of the inner surface of the colon depicted in successive endoscopic images may be tracked for each position. The portion of the inner surface of the colon may be calculated for each position within the colon for different orientations of the camera without displacing the camera forward or backward (e.g., when the displacement is zero or less than a predetermined threshold). When the camera is moved, the portion of the inner surface of the colon depicted in the captured images may be recalculated for each new position, for example, as the colonoscope is withdrawn from the colon. For example, Q1, Q2, Q3, Q4 is a set for the current location, and another set of Q1, Q2, Q3, Q4 is a set for the new location. Alternatively, the cumulative portion of the inner surface of the colon depicted in successive endoscopic images is tracked during spiral motion of the camera. The portion of the inner surface of the colon depicted in the captured image may be recalculated in a spiral pattern as the camera is moved, for example, quarters are individually repeated for a spiral motion, e.g., Q1, Q2, Q3, Q4, Q1, Q2, Q3, Q4, Q1, Q2, Q3, Q4.
[0149] Optionally, each portion corresponds to a time window having an interval corresponding to the amount of time to cover the entire inner circumference during the spiral scan operation, i.e., returning to the same arc position after a single spiral scan, i.e., completing approximately 360 degrees, or returning to an arc range defining the same quarter, i.e., completing Q1, then completing Q2, Q3, and Q4, and returning to Q1. The indication of coverage for each portion may be updated relative to the time window. For example, the operator has properly covered all portions when all portions are continuously associated with an indication of coverage (e.g., when all portions are colored red or another color for proper coverage). Each portion may be associated with an indication of proper coverage for a time interval corresponding to the time window, i.e., the portion is properly covered up to the end of the current spiral and should be properly depicted in the new spiral. The operator can use the indication of proper coverage of all portions as a guideline during successive spiral scans to properly capture images of the colon. When one of the representations of one of the portions changes to a representation of inadequate coverage, the operator can capture the respective portion in the image.
[0150] The time window may be, for example, about 2.5 seconds per quarter, about 5 seconds per quarter, or other values per section, or about 10 or 20 seconds per circumferential spiral scan, or other values, based on, for example, clinical guidelines and / or physician practice. The time window may be dynamically calculated and adjusted based on real-time measurements of the spiral scan motion being performed by the operator. For example, if the operator stops the spiral scan to focus on a polyp, the time window is stopped and restarted when the operator resumes the spiral scan. If the operator slows the spiral scan speed, for example, by having the same operator or a student (e.g., a resident) perform the scan, the time window increases accordingly. The real-time speed of the spiral scan motion may be measured, for example, by analysis of the captured images (e.g., in terms of frame capture rate, tracking distance between matching extracted features), and / or by a sensor(s) sensing the movement of the femoral endoscope. An indication of adequate coverage is generated when one or more images that largely depict the respective portion are captured during the time window, and / or another indication of inadequate coverage is generated when no images that largely depict the respective portion are captured during the time window.
[0151] An indication of adequate coverage may be generated, for example, when at least 50%, 60%, 70%, 80%, or 90%, or some other intermediate or greater value of each portion is depicted in each image. The threshold for determining the amount of portion required in each image may be set based, for example, on the lens and the area of the inner surface of the colon depicted in the image.
[0152] Optionally, the 3D values of pixels in the tracked frames are combined into a single 3D coordinate system. A 3D panoramic (mosaic) image of the colon surface can be generated based on a single 3D coordinate system, for example, based on the process described with reference to Morimoto, Carlos, and Rama Chellappa et al., "Fast 3D stabilization and mosaic construction," cvpr. IEEE, 1997. The mosaic image may be constructed continuously (e.g., as the endoscope is withdrawn from the colon).
[0153] For each frame included in the mosaic image, the luminal region (e.g., the center of the colonic tube) may be detected by identifying dark pixels, for example, as pixels having intensity values below a threshold (e.g., below 20 or other value if the intensity range is 0-255), where the dark pixels have correlated 3D positions (or 3D positions correlated to pixels in their immediate environment) that are farthest from the camera position when the frame was captured.
[0154] Optionally, a perimeter position of the inner surface of the colon covered by the current frame is calculated. Optionally, the inner surface of the colon is divided, for example, into four quarters. The quarter of the inner surface of the colon represented by the current frame can be calculated, for example, the first, second, third, or fourth quarter of the colon. The captured images may be aggregated into a mosaic image to incrementally cover quarters, possibly until the mosaic image depicts all four quarters, indicating that the entire circumference of the inner surface of the colon at the current position is depicted in the image.
[0155] The quarter(s) of the inner surface of the colon depicted in the individual images and / or mosaic image may be calculated, for example, based on the detected lumen area. For each new frame (e.g., each individual frame presented and / or added to the mosaic image), the 3D position of the center pixel and / or the orientation of the 3D position of the center pixel relative to the detected 3D position of the lumen is calculated. The 3D position of the lumen may be estimated from the 3D positions of pixels in its nearby environment. The quarter covered by the current frame may be calculated.
[0156] The estimated colon quarter depicted by the currently presented frame and / or mosaic image may be displayed to the user. When all four quarters are covered, an indication may be output indicating that the camera may be moved to a new position, and / or another indication may be generated indicating a lack of adequate imaging of the local colon region when one or more quarters are not covered and the camera is moved. Using the indication, the physician performing the colonoscopy procedure can determine whether the four quarters of the current local colon region are sufficiently covered in consecutive frames. The quarters that are sufficiently covered in the frames may be displayed to the user (e.g., for at least 5 seconds) so that the user can view the combination of the four displayed quarters together on the screen as an indication that the inner surface of the colon is sufficiently covered during the colon scan. The covered quarters may be indicated, for example, by their number appearing and / or flashing on the screen.
[0157] Optionally, while the endoscope is advancing and / or retracting (e.g., on its way out of the colon, such as by being withdrawn from the cecum, until the endoscope is completely removed from the colon), the sequence of covered quarters is dynamically aggregated and / or tagged with their respective 3D positions. The calculated trajectory of the endoscope (e.g., as the endoscope is advancing and / or retracting) may be scanned in windows of a predetermined length (e.g., about 2-3 cm) and / or a predetermined stride length (e.g., about 0.5-1 cm). If a quarter is not covered at all in a number of consecutive scanning windows (denoted by Ncs, e.g., having a value of 3), it is registered as a missed quarter. The percentage of the colon covered by the endoscopic camera during advancement and / or withdrawal is a mathematical relationship:
number
[0158] Reference is now made to Figure 12, which is a schematic diagram illustrating individual images 1402 displayed within each quarter of the interior surface of the colon depicted therein, according to some embodiments of the present invention. Portion 1404A represents quarter 1, portion 1404B represents quarter 2, portion 1404C represents quarter 3, and portion 1404D represents quarter 4.
[0159] Reference is now made to Figure 13, which is a flowchart of a method for calculating the quarter represented by a frame according to some embodiments of the present invention, where the selected quarter represents the quarter that is most covered by the frame since a frame may overlap between two or more quarters.
[0160] At 1502, the current frame (denoted by #n) is provided.
[0161] At 1504, frame #n is registered to the currently constructed 3D mosaic image.
[0162] At 1506, a lumen may be detected in the current frame. The 3D position of the lumen may be estimated.
[0163] At 1508, if a lumen (eg, a dark area in the center of the colonic tract) has not already been detected in the last frame of the current mosaic (eg, in the last 10 seconds of the video), R02-R08 are repeated.
[0164] If a lumen is detected, the 3D location of the central pixel of the image is calculated at 1510. If the central pixel is determined to be within the lumen region, the 3D location is calculated according to pixels in the nearby environment.
[0165] At 1512, the relative direction (eg, 3D vector direction) between the 3D position of the center pixel of the current frame and the estimated 3D position of the last detected lumen region is calculated.
[0166] At 1514, based on the assumption that the vector whose direction was calculated at 1512 originates from the center of the colonic tract (i.e., the center of the lumen), the quarter of the colon that covers the majority of the area of the current frame is determined according to the direction of the vector. The selected quarter is the quarter that is most covered by the current frame (if two or more quarters are equally covered and the uncovered quarters are not displayed).
[0167] The quarter shown in the current frame is provided at 1516. Instructions may be generated to present an indication of the determined quarter in a GUI, as described herein.
[0168] Referring back to FIG. 1 , at 114, dimensions of the detected polyp are calculated. The dimensions may be 2D and / or 3D dimensions, e.g., the volume of the polyp, the radius of a sphere representing the polyp, the surface area of the polyp (e.g., a flat polyp), and / or the radius of a 2D circle representing the polyp. The dimensions may be calculated by performing a best fit of the 3D sphere and / or 2D circle to pixels of a 3D image showing the polyp (created from the 2D image as described herein, optionally created by a 3D CNN). The dimensions are obtained from the best-fit 3D sphere and / or 2D circle, e.g., the radius of the best-fit 3D sphere and / or the radius of the best-fit 2D circle.
[0169] The dimensions of the polyp are calculated based on the calculated polyp-delineating ROI and / or the 3D coordinates of the pixels of the 3D image (i.e., the 3D reconstruction of the 2D image) output by the 3D reconstruction neural network as described with reference to 108 and / or 110. The polyp-delineating ROI may be set manually by an operator (e.g., using a GUI) and / or may be automatically output by the detection network described with reference to 104 when supplied with (one or more) 2D images.
[0170] The processes described herein allow for the calculation of the 3D volume and / or 2D surface area dimensions of a polyp using (optionally only using) 3D coordinate values (x, y, z) calculated for pixels in a 2D image (e.g., output by a 3D CNN) as described herein. In contrast, other processes that calculate 3D and / or 2D dimensions using a 2D image require knowledge of camera characteristics (e.g., the camera's pose relative to the polyp), which may be difficult to obtain.
[0171] The size of the polyp can be calculated based on the radius of the circle that best fits the polyp slice.
[0172] The volume of a polyp can be calculated automatically by taking the calculated 3D values of the 2D pixels inside the outline and / or bounding box that delineates the polyp and finding a best-fit 3D sphere (or circle, if the polyp is flat) of the exposed 3D surface created by interpolating between the 3D values of the pixels. The radius of the sphere (or circle) is the polyp size.
[0173] Note that the 3D volume can be calculated from the calculated radius of a sphere that correlates to the polyp, as described herein.
[0174] The estimated dimensions (e.g., 3D volume) of the polyp(s) may be calculated, for example, by the following process: Calculate a best-fit 3D surface for the 3D coordinates of pixels of a region of at least one 2D image; Calculate multiple normal vectors proximate a 3D position correlating to the centroid of the region (i.e., ROI) depicting the polyp; Determine the side of the 3D surface that is concave according to the relative directions of the normal vectors; Calculate a plane containing the normal vectors of the multiple normal vectors correlating to the 3D position correlating to the centroid; Calculate a vector correlating to the centroid that is tangent to the best-fit 3D surface at the 3D position; Calculate a first curvature as the radius of a tangent parabola of the contour that intersects between the best-fit 3D surface and the calculated plane; Calculate a second curvature as the radius of a tangent parabola of the contour that intersects between the best-fit 3D surface and an orthogonal plane that is orthogonal to the calculated plane, the orthogonal plane containing the normal vectors. The 3D radius of the 3D volume of the polyp is then calculated as the average of the first curvature and the second curvature.
[0175] Referring now to FIG. 14, FIG. 14 includes a schematic diagram illustrating a process for calculating polyp volume from a 3D reconstructed surface calculated from a 2D image, according to some embodiments of the present invention. Schematic 1602 shows a 2D colonoscopy frame with a detected polyp 1604 obtained from the web address site(dot)google(dot)com(slash)site(slash)suryaiit(slash)research(slash)endoscopy(slash)sfs. Schematic 1606 shows a 3D reconstructed surface of frame 1602. Polyp 1608 is a 3D reconstruction of polyp 1604 in 2D image 1602. Image 1610 illustrates a process for calculating polyp volume by finding a best-fit (tangent) 3D sphere 1612 on the concave side of the surface interpolated from the 3D values of pixels within the polyp region, as described herein. The volume is the radius (R) of 3D sphere 1612.
[0176] Reference is now made to FIG. 15, which is a flowchart of an exemplary process for calculating polyp volume from 2D images, according to some embodiments of the present invention.
[0177] A 2D frame denoted #n having a detected polyp is received at 1702. Multiple frames before and after the current frame are received.
[0178] At 1704, a 3D reconstruction of frame #n is calculated as described herein.
[0179] At 1706, pixels within the bounding contour and / or bounding box of the polyp (as available) are identified and a best fit surface of their 3D values is calculated.
[0180] At 1708, multiple normal vectors are calculated for a 3D point neighborhood that correlates to the centroid of the polyp's bounding box (if there is a delineation contour and the centroid falls outside the delineation contour, the centroid is replaced by the nearest pixel within the delineation contour).
[0181] At 1710, the concave side of the surface is identified according to the calculated normal vectors and their relative directions.
[0182] At 1712, an infinite plane is calculated that contains the normal vector relative to the centroid (or its permutation) above and the vector at this 3D point (relative to the centroid) that is tangent to the surface.
[0183] In 1714, the curvature of the contour (the radius of the tangent parabola), which is the intersection between the polyp surface and the calculated plane, is calculated based on the approach described, for example, in Har'el, Zvi et al., "Curvature of curves and surfaces—a parabolic approach," Department of Mathematics, Technion-Israel Institute of Technology (1995).
[0184] At 1716, the last step is repeated for a plane that is orthogonal to the previous plane and contains the normal vectors from 1712 and 1714 (to get an additional estimate for curvature).
[0185] At 1718, the average of the last two calculated curvatures is calculated to provide the 3D radius of the polyp.
[0186] If the 3D radius is greater than a predetermined threshold (e.g., 7 millimeters, 7 centimeters, or other value) at 1720, the polyp size is indicated as flat at 1722. A best-fit 2D circle is calculated. The polyp size is defined according to the size of the 2D circle. If the 3D radius is less than a predetermined threshold at 1726, the polyp size is defined according to the 3D radius.
[0187] Referring back to FIG. 1 , instructions are generated at 116 to update the presentation of the GUI. The image may be augmented. The instructions may be for introducing GUI elements into the image, for example, to present an overlay on top of the image and / or to present the image in one portion of the GUI and other graphical elements in another portion of the GUI, e.g., the covered quadrant and / or the colon map, as described herein.
[0188] The instructions are generated based on the output of one or more of the features described with reference to 104-114. *Instructions are generated to augment the image at the location of the detected polyp as described with reference to 104, e.g., marking an ROI (e.g., a bounding box) depicting the detected polyp, and / or color-coding the polyp and / or bounding box, and / or an arrow pointing to the polyp. *As described with reference to feature 106, instructions are generated based on a representation of a vector pointing from the current image toward the ROI to depict polyps located outside the current image. The instructions are for generating an augmented endoscopic image by augmenting each endoscopic image with the representation of the vector.
[0189] The vectors may be presented as arrows, pointing in the direction the camera should be moved to recapture the polyp in the image. The arrows and images may be presented in 2D.
[0190] Optionally, when the polyp is recaptured in the image (e.g., after moving the camera in the direction of the arrow), instructions are generated to augment the image with a ROI depicting the polyp. The ROI may be remarked on the image without necessarily performing feature 104, i.e., based on tracking only. *Optionally, instructions are generated for presenting a colon map within the colon based on the output of the features described with reference to 106 and / or 110. The 2D and / or 3D locations of polyps (e.g., ROIs depicting the polyps) are output, for example, by the detection neural network and / or the 3D reconstruction neural network and / or manually marked by an operator (e.g., using a GUI, such as by pressing a "polyp" icon). The locations of polyps are marked on a schematic representation of the patient's colon, for example, with indicators such as Xs and / or circles, and optionally color-coded according to whether polyps were identified during colonoscope insertion and / or colonoscope removal. The colon map is dynamically updated with newly detected polyps. *Optionally, instructions are generated for marking polyps presented on the colon map with an indication of treatment (e.g., surgical removal, resection). For example, polyps presented as ellipses in the colon map that have been removed are marked with an X. Ellipses in the map that have not been removed are not marked with an X. Polyps are marked as treated based on an indication of polyp removal, for example, manually provided by the user (e.g., pressing a "polyp removal" icon on the GUI) and / or automatically detected by the code (e.g., based on detecting movement of a surgical tool in the colonoscope). *Optionally, instructions are generated to plot the tracked 3D position of the endoscope on a colon map presented within the GUI, e.g., as a trajectory and / or curve and / or dotted line. The trajectory and / or curve and / or dotted line are dynamically updated as the endoscope moves (e.g., forward and / or backward) within the colon. Optionally, the forward tracking 3D position is marked on the colon map with markings, e.g., a forward arrow, and / or using color coding, indicating the forward direction of the endoscopic camera penetrating deeper into the colon (e.g., from the rectum to the cecum). The reverse tracking 3D position presented on the colon map is marked with another marking, e.g., a reverse arrow and / or a different color, indicating the reverse direction of the endoscopic camera being removed from the colon (e.g., from the cecum to the rectum).
[0191] *Optionally, instructions are generated for presenting calculated dimensions of the detected polyps. For example, the calculated dimensions are presented as numerical values, optionally having units that are close to the ROI marked on the image, for example, a calculated volume value in cubic millimeters. In another example, the calculated dimensions are presented as numerical values having units that are close to the representation of the polyps on the colon map. *Optionally, instructions are generated to present a warning in the GUI indicating a suggestion to remove a polyp when a dimension exceeds a threshold, for example, when the radius of a sphere defining the 3D volume of the polyp exceeds 2 millimeters. The warning may be generated, for example, as coloring of the polyp and / or ROI boundary depicting the polyp on the image, and / or coloring of the polyp representation on the colon map, using a unique color indicating the suggestion to remove (e.g., red for remove and green for leave in), and / or as an audio message in the GUI, and / or as a pop-up text message. *Optionally, instructions are generated to present an estimate of the remaining portion of the inner surface area not yet depicted in a previously captured endoscopic image. Alternatively or additionally, instructions are generated to present an estimate of the portion of the inner surface area that has already been depicted. For example, a circular GUI element divided into four quadrants (or other number of divisions, e.g., as slices) is presented. Quadrants depicted in the image are indicated with one marking, e.g., green. Quadrants not yet depicted are indicated with a different marking, e.g., red. The operator can see the markings (e.g., colored circles) and can point the camera toward quadrants not yet imaged. When the GUI element indicates that all quadrants have been imaged, the operator can displace the colonoscope (e.g., forward or backward).
[0192] A quadrant can depict a mostly imaged area, for example, more than 50%, more than 70%, or 80% of the surface of the quadrant depicted in one or more images. Other divisions can be selected according to the imaging capabilities of the camera's lens, for example, a larger number of divisions can be used for narrow-angle lenses. *Optionally, instructions for issuing a warning are generated according to a predefined system definition (which may be set, for example, through a setup menu) when the camera's current position is close to an anatomical landmark and / or polyp. For example, the definition can be set to generate a warning when the camera's position is close to the landmark's position, and the distance is less than a predetermined threshold. The distance between the camera and the landmark and / or polyp can be calculated, for example, by L2 meters and / or a neighborhood small enough to trigger an alert (e.g., approximately 3 centimeters). The threshold distance value may take into account aggregate calculation errors in constructing the trajectory and / or the fact that the colon itself is not completely stable in the abdomen, i.e., not stable in the coordinate system in which the trajectory is constructed.
[0193] Optionally, the amount of time the endoscope spends at each defined portion of the colon can be calculated and presented. Each portion may be defined, for example, according to the transition points between anatomical landmarks that divide the colon into portions, for example, the ascending colon and the transverse colon (e.g., between the transverse colon and the ascending colon). The anatomical landmarks may be detected manually by the user (e.g., the user marks the landmarks using a GUI) and / or automatically by code (e.g., based on analysis of the image), and the 3D position of the endoscopic camera relative to the anatomical landmarks is tracked. The amount of time spent by the endoscopic camera at each portion of the colon is calculated. Instructions are generated for presenting the amount of time spent by the endoscopic camera at each portion of the colon in the GUI. The time may be presented, for example, as markings on respective portions of a colon map corresponding to portions of the patient's colon.
[0194] The generated instructions are executed at 118. The GUI is updated according to the instructions.
[0195] Referring now to FIG. 16 , FIG. 16 is a schematic diagram illustrating a sequence of augmented images for tracking a polyp ROI presented in a GUI according to generated instructions, in accordance with some embodiments of the present invention. The sequence illustrates 2D tracking of a polyp and the generation of an arrow in the direction of the polyp when the polyp is located outside the image, as described herein. Image 1902 is a first image augmented with an ROI 1904 depicting an automatically detected polyp. In the second image of sequence 1906, the ROI 1904 has moved toward the left of the image due to camera reorientation. In the third image of sequence 1908, the ROI 1904 has moved further to the left of the image. In the fourth image of sequence 1910, the ROI is located outside the image and is therefore not shown in image 1910. An arrow 1912 is shown pointing in the direction of the externally located ROI. In the fifth image of sequence 1914, ROI 1904 has reappeared after the operator reoriented the camera in the direction indicated by arrow 1912 in image 1910. In the sixth image of sequence 1916, ROI 1904 has moved further to the right due to camera movement by the operator.
[0196] 17, which is a schematic illustration of a colon map 2002 presented within a GUI in accordance with some embodiments of the present invention, showing the trajectory of endoscope movement, the locations of detected polyps, removed polyps, and anatomical landmarks. Colon map 2002 may be presented, for example, as an overlap presented in a corner of the image, or within a designated window of the GUI with another designated window for presenting the image. Trajectory 2004 shows the tracked 3D movement of the endoscope during forward movement into the colon, for example, color-coded. Trajectory 2006 shows the tracked 3D movement of the endoscope during reverse movement out of the colon, for example, color-coded with a different color. Ellipse 2008 shows polyps automatically detected during forward movement of the colonoscope, for example, colored with the same color used for forward trajectory 2004. Ellipse 2010 shows polyps automatically detected during reverse movement of the colonoscope, for example, colored with the same color used for reverse trajectory 2006. Circle 2012 indicates polyps manually located by the user and / or automatically detected polyps manually marked by the user. X marking 2014 indicates polyps removed by the operator. Arrowhead 2016 indicates the current 3D position of the distal end (e.g., tip) of the endoscope, optionally the camera. Box 2018 indicates manually specified landmarks, e.g., hemorrhoids, manually entered by the user (e.g., via a GUI). Circle 2020 indicates automatically identified landmarks, e.g., detected by analysis of the 3D reconstructed image. Colon map 2002 dynamically updates as the colonoscope is moved, as polyps are detected and / or as polyps are removed.
[0197] Referring to FIG. 18, FIG. 18 illustrates a schematic 2102 showing quadrants of the colon's inner surface depicted in one or more images 2104A-C and / or quadrants of the colon's inner surface not yet depicted in image 2106, according to some embodiments of the present invention. The schematic 2102 is created for each new position of the camera as the camera is displaced forward and / or backward. Alternatively or additionally, the schematic 2102 is calculated for a recently defined time window (e.g., 5 seconds, or 10 seconds, or other value). The time window may be short enough to eliminate forward and / or backward movement of the camera, but long enough to allow complete imaging of the periphery of the colon's inner wall. The quadrants are updated as new images are acquired at the same location by reorienting the camera. The quadrants 2104A-C and 2106 may be color-coded. The schematic 2102 is dynamically updated as the operator moves the camera to image the remaining portions of the colon's inner surface.
[0198] At 120, one or more of the features described with reference to 100-118 are iterated. The iteration can, for example, dynamically update the GUI to dynamically track polyps, dynamically generate arrows pointing to ROIs depicting polyps located outside the current image, update the colon map with new 2D and / or 3D positions of the polyps, update the camera trajectory of the colon map with camera movement, update the GUI to depict coverage of the inner surface (e.g., quadrants) of the colon, and / or update the calculated 2D and / or 3D dimensions of the detected polyps.
[0199] The updated GUI may be used by the operator, for example, to operate the camera to capture images of polyps, to determine which polyps should be removed, to operate the camera to ensure complete coverage of the inner surface of the colon, and / or to track the location of the camera and / or detected polyps within the colon.
[0200] The description of various embodiments of the present invention is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, practical applications or technical improvements to technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0201] It is anticipated that many related anatomical imaging technologies will be developed during the life of the patent starting with this application, and the scope of the term anatomical imaging is intended to include, a priori, all such new technologies.
[0202] As used herein, the term "about" refers to ±10%.
[0203] The terms "comprises," "comprising," "includes," "including," "having," and their conjugations mean "including but not limited to." This term encompasses the terms "consisting of" and "consisting essentially of."
[0204] The phrase "consisting essentially of" means that a composition or method may include additional components and / or steps, but only if the additional components and / or steps do not materially alter the basic and novel characteristics of the claimed composition or method.
[0205] As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" can include a plurality of compounds, including mixtures thereof.
[0206] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments and / or as excluding the incorporation of features from other embodiments.
[0207] The word "optionally" is used herein to mean "provided in some embodiments and not provided in other embodiments." Any particular embodiment of the present invention may include multiple "optional" features unless such features are inconsistent.
[0208] Throughout this application, various embodiments of the present invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the present invention. Accordingly, the description of a range should be considered to have specifically disclosed not only each individual numerical value within that range, but also all possible subranges. For example, description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as each individual numerical value within that range, e.g., 1, 2, 3, 4, 5, 6, etc. This applies regardless of the breadth of the range.
[0209] Whenever a numerical range is given herein, it is meant to include any numerical value (decimal or integer) recited within the given range. The phrases "ranging from / between" a first designated number and a second designated number, and "ranging from / between" a first designated number to a second designated number are used interchangeably herein and are meant to include the first and second designated numbers, and all decimals and integers therebetween.
[0210] It will be understood that certain features of the invention that are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention that are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as appropriate in other described embodiments of the invention. Certain features described in the context of various embodiments are not considered essential features of those embodiments, unless the embodiment is inoperable without those elements.
[0211] While the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, the present invention is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.
[0212] All publications, patents, and patent applications mentioned herein are incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Furthermore, citation or identification of a reference in this application shall not be construed as an admission that such reference is available as prior art to the present invention. To the extent section headings are used, they should not be construed as necessarily limiting. Additionally, the priority documents of this application are incorporated herein by reference in their entireties.
Claims
1. 1. A system for generating instructions to present a GUI based on dynamically tracking 3D movement of an endoscopic camera capturing multiple 2D endoscopic images within a patient's colon, comprising: feeding each 2D endoscopic image to a 3D reconstruction neural network; outputting, by the 3D reconstruction neural network, a 3D reconstruction of each of the 2D endoscopic images, wherein pixels of the 2D endoscopic images are assigned 3D coordinates; calculating a current 3D position of the endoscopic camera within the colon according to the 3D reconstruction; generating instructions for presenting the current 3D position of the endoscopic camera on a colon map within the GUI; analyzing each of the 3D reconstructions to estimate a portion of the inner surface of the colon depicted in each endoscopic image; tracking cumulative portions of the inner surface of the colon depicted in successive endoscopic images during a spiral scanning motion of the endoscopic camera during a colonoscopy procedure; and generating instructions for presenting in a GUI at least one of an estimate of remaining portions of the inner surface not yet depicted in previously captured endoscopic images and an estimate of total coverage of an area of the inner surface, wherein the analyzing, tracking, and generating are repeated during the spiral scanning motion; for each endoscopic image of a plurality of endoscopic images.
2. tracking the 3D position of the endoscopic camera; 10. The system of claim 1, further comprising code for: plotting the tracked 3D position of the endoscopic camera within a colon map GUI.
3. 2. The system of claim 1, wherein the forward tracking 3D position of the endoscopic camera is marked on the colon map presented on the GUI with a symbol indicating that the forward direction of the endoscopic camera goes deeper into the colon, and the reverse tracking 3D position of the endoscopic camera presented on the colon map is marked with another symbol indicating the reverse direction of the endoscopic camera being removed from the colon.
4. providing each of the endoscopic images to a detection neural network; outputting, by the detection neural network, a representation of a region of the endoscopic image depicting at least one polyp; 2. The system of claim 1, further comprising code for: generating instructions for calculating an estimated 3D position of the at least one polyp within the region of the endoscopic image according to the 3D reconstruction; and presenting the 3D position of the at least one polyp on a colon map within the GUI.
5. receiving an indication for surgical removal of the at least one polyp from the colon; 5. The system of claim 4, further comprising code for performing: marking the 3D location of the at least one polyp on the colon map with an indication of removal of the at least one polyp.
6. tracking a 3D position of an endoscopic camera capturing the plurality of endoscopic images; calculating an estimated distance from a current 3D position of the endoscopic camera to a 3D position of at least one polyp previously identified using previously acquired endoscopic images; The system of claim 4 , further comprising code for executing: generating instructions for presenting a display within the GUI if the estimated distance is below a threshold.
7. 10. The system of claim 1, further comprising code for calculating the percentage of the colon covered by the endoscopic camera during advancement and / or withdrawal, wherein the percentage is calculated using the mathematical relationship (1-Tms / (Nsw / Ncs*4))*100, where Tms indicates the total number of missed quarters, Ncs indicates the number of consecutive scan windows, and Nsw indicates the total number of windows scanned.
8. 2. The system of claim 1, wherein each portion corresponds to a time window having an interval corresponding to an amount of time to cover the respective portion during the spiral scanning operation, and wherein an indication of adequate coverage is generated when at least one image that largely depicts the respective portion is captured during the time window, and / or another indication of inadequate coverage is generated when no image that largely depicts the respective portion is captured during the time window.
9. The system of claim 8, wherein an indication of the total amount of inner surfaces depicted in the image relative to the amount of undepicted inner surfaces is calculated by aggregating the portion covered during the spiral scanning operation relative to the portion not covered during the spiral scanning operation.
10. 2. The system of claim 1, wherein the 3D reconstruction neural network is trained with a training dataset of pairs of 2D endoscopic images that define input images and corresponding 3D coordinate values calculated for pixels of the 2D endoscopic images calculated by a 3D reconstruction process that define ground truth.
11. receiving an indication of at least one anatomical landmark of the colon, the at least one anatomical landmark dividing the colon into a plurality of portions; tracking a 3D position of the endoscopic camera relative to the at least one anatomical landmark; calculating an amount of time spent by the endoscopic camera in each of the plurality of portions of the colon; 10. The system of claim 1, further comprising code for executing: generating instructions for presenting within the GUI the amount of time spent by the endoscopic camera on each of the plurality of portions of the colon.
12. The system of claim 1 , further comprising code for calculating a three-dimensional (3D) volume of a polyp based on at least one 2D endoscopic image of the plurality of 2D endoscopic images.
13. receiving the at least one 2D endoscopic image depicting an interior surface of the colon captured by the endoscopic camera positioned within the lumen of the colon; receiving a representation of a region of the at least one 2D endoscopic image depicting at least one polyp; feeding the at least one 2D endoscopic image to a 3D reconstruction neural network; outputting, by the 3D reconstruction neural network, a 3D reconstruction of the at least one 2D endoscopic image, wherein pixels of the 2D endoscopic image are assigned 3D coordinates; calculating a 3D volume of the at least one polyp within the region of the at least one 2D endoscopic image according to the analysis of the 3D coordinates of pixels of the region of the at least one 2D endoscopic image; The system of claim 12 , further comprising code for calculating a 3D volume of the polyp by:
14. 14. The system of claim 13, wherein the representation of a region of the at least one 2D endoscopic image depicting at least one polyp is output by a detection neural network trained to segment polyps in 2D endoscopic images.
15. 14. The system of claim 13, wherein the 3D reconstruction of the at least one 2D endoscopic image is provided to the detection neural network in combination with the at least one 2D endoscopic image to output the representation of a region depicting the at least one polyp.