Systems and methods for processing colon images and videos
The system enhances colonoscopy by dynamically tracking polyps in 2D and 3D to improve detection and removal rates through real-time GUI feedback, addressing the limitations of current image processing in colonoscopy.
Patent Information
- Application Number
- JP2021572304
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-04
- Filing Date
- 2020-05-06
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2040-05-06
AI Technical Summary
Colonoscopy procedures face challenges in accurately detecting and removing polyps due to limitations in image processing technology, leading to missed lesions and suboptimal adenoma detection rates, which can impact the effectiveness of early cancer prevention.
A system and method for processing colon images and videos using graphical user interfaces (GUIs) that dynamically track polyps in 2D and 3D, providing real-time feedback to adjust the endoscope camera, map polyp positions, and calculate polyp volumes, enhancing detection and removal efficiency.
Improves the detection and removal rate of polyps by ensuring complete evaluation of the colon surface, reducing missed lesions, and providing real-time guidance for optimal polyp identification and surgical intervention.
Smart Images

Figure 0007711946000003 
Figure 0007711946000004 
Figure 0007711946000005
Abstract
Description
Technical Field
[0001] [Related Applications] This application claims the benefit of priority of U.S. Provisional Patent Application No. 16 / 430,461, filed on June 4, 2019, the content of which is hereby incorporated by reference in its entirety.
Background Art
[0002] In some embodiments, the present invention relates to colonoscopy, and more specifically, to systems and methods for processing colon images and videos and / or for processing colon polyps automatically detected during a colonoscopy procedure, but is not limited thereto.
[0003] Colonoscopy is the gold standard for the detection of colon polyps. In colonoscopy, a long flexible tube called a colonoscope is advanced into the colon. A video camera at the end of the colonoscope captures images, which are presented to a physician on a display. The physician examines the inner surface of the colon to check for the presence of polyps. Identified polyps are removed using the instruments of the colonoscope. Removing cancerous polyps early can eliminate or reduce the risk of colon cancer.
Summary of the Invention
Means for Solving the Problems
[0004] According to a first aspect, a method for generating instructions to present a graphical user interface (GUI) for dynamically tracking at least one polyp in a plurality of endoscopic images of a patient's colon includes tracking the position of a region depicting at least one polyp in each endoscopic image relative to at least one previous endoscopic image, calculating a vector from the position of the region within each endoscopic image to the position of the region outside the respective endoscopic image when the position of the region is outside the respective endoscopic image, creating an augmented endoscopic image by augmenting each endoscopic image with the display of the vector, and generating instructions for displaying the augmented endoscopic image within the GUI, and repeating for a plurality of endoscopic images.
[0005] According to a second aspect, a method for generating instructions to present a GUI for dynamically tracking the 3D movement of an endoscope camera that captures a plurality of 2D endoscopic images within a patient's colon includes, for each of the plurality of endoscopic images, supplying each 2D endoscopic image to a 3D reconstruction neural network, outputting, by the 3D reconstruction neural network, a 3D reconstruction of each 2D endoscopic image, where 3D coordinates are assigned to the pixels of the 2D endoscopic image, calculating the current 3D position of the endoscope camera within the colon according to the 3D reconstruction, and generating instructions for presenting the current 3D position of the endoscope camera on a colon map within the GUI, and repeating.
[0006] According to a third aspect, a method for calculating a three-dimensional volume of a polyp based on at least one two-dimensional (2D) image includes receiving at least one 2D image of the inner surface of the colon captured by an endoscopic camera disposed within the lumen of the colon, receiving a display of a region of the at least one 2D image depicting at least one polyp, supplying the at least one 2D image to a 3D reconstruction neural network, outputting, by the 3D reconstruction neural network, a 3D reconstruction of the at least one 2D image, wherein 3D coordinates are assigned to pixels of the 2D image, and calculating an estimated 3D volume of at least one polyp within the region of the at least one 2D image according to an analysis of 3D coordinates of pixels of the region of the at least one 2D image.
[0007] In a further embodiment of the first aspect, the display of the vector depicts a direction and / or orientation for adjustment of the endoscopic camera for capturing at least one other endoscopic image depicting a region of the at least one endoscopic image.
[0008] In a further embodiment of the first aspect, when the position of the region depicting at least one polyp appears in each endoscopic image, an augmented endoscopic image is generated by augmenting each endoscopic image at the position of the region, and the display of the vector is excluded from the augmented endoscopic image.
[0009] In a further embodiment of the first aspect, further includes calculating the position of a region depicting at least one polyp within the patient's colon, creating a colon map by marking a schematic diagram representing the patient's colon using a display indicating the position of the region depicting at least one polyp, and generating an instruction to present the colon map within a GUI, wherein the colon map is dynamically updated at the position of newly detected polyps.
[0010] In a further embodiment of the first aspect, translating and / or rotating at least one endoscopic image among a continuous subset of a plurality of endoscopic images, each endoscopic image for generating a processed continuous subset of the plurality of endoscopic images, such that a region depicting at least one polyp is at the same approximate position in all of the images of the continuous subset of the plurality of endoscopic images; supplying the processed continuous subset of the plurality of endoscopic images to a detection neural network; outputting, by the detection neural network, a current region depicting at least one polyp for each endoscopic image; creating an augmented image of each endoscopic image by augmenting each endoscopic image with the current region; and generating an instruction for presenting the augmented image in the GUI.
[0011] In a further embodiment of the first aspect, when the output of the neural network is provided for a previous endoscopic image that is serially earlier than each endoscopic image and the tracked position of the region depicting at least one polyp in each endoscopic image is at a position different from the region output by the neural network for the previous endoscopic image, an augmented image for each endoscopic image is generated based on the tracked position.
[0012] In a further embodiment of the second aspect, further including tracking the 3D position of the endoscopic camera and plotting the tracked 3D position of the endoscopic camera in the colon map GUI.
[0013] In a further embodiment of the second aspect, the forward tracked 3D position of the endoscopic camera is marked on the colon map presented in the GUI with a mark indicating that the forward direction of the endoscopic camera penetrates deeper into the colon, and the reverse tracked 3D position of the endoscopic camera presented on the colon map is marked with another mark indicating the reverse direction of the endoscopic camera being removed from the colon.
[0014] In a further embodiment of the second aspect, supplying each endoscopic image to a detection neural network, outputting, by the detection neural network, a display of a region of the endoscopic image depicting at least one polyp, calculating, according to the 3D reconstruction, an estimated 3D position of at least one polyp within the region of the endoscopic image, and generating instructions for presenting the 3D position of at least one polyp on a colon map within a GUI.
[0015] In a further embodiment of the second aspect, receiving a display of surgical removal of at least one polyp from the colon, and marking the 3D position of at least one polyp on the colon map with the display of the removal of at least one polyp.
[0016] In a further embodiment of the second aspect, tracking the 3D position of an endoscopic camera that captures a plurality of endoscopic images, calculating, from the current 3D position of the endoscopic camera, an estimated distance from the previously acquired endoscopic images to the 3D position of at least one polyp previously identified using the previously acquired endoscopic images, and generating instructions for presenting a display within the GUI when the estimated distance is below a threshold.
[0017] In a further embodiment of the second aspect, analyzing each 3D reconstruction to estimate a part of the inner surface of the colon depicted in each endoscopic image, tracking, during a spiral scanning operation of the endoscopic camera during a colonoscopy procedure, a cumulative part of the inner surface of the colon depicted in successive endoscopic images, and generating instructions for presenting at least one of an estimation of the remaining part of the inner surface not yet depicted in the previously captured endoscopic images and an estimation of the total coverage of the region of the inner surface within the GUI, wherein analyzing, tracking, and generating are repeated during the spiral scanning operation.
[0018] In a further embodiment of the second aspect, each part corresponds to a time window having an interval corresponding to the amount of time for covering each part during the spiral scanning operation, and when at least one image depicting each part is captured during the time window, a display of appropriate coverage is generated, and / or when an image depicting each part is not captured during the time window, another display of inappropriate coverage is generated.
[0019] In a further embodiment of the second aspect, for the parts not covered during the spiral scanning motion, by aggregating the parts covered during the spiral scanning motion, a display of the total amount of the inner surface depicted in the image relative to the amount of the inner surface not depicted is calculated.
[0020] In a further embodiment of the second aspect, the 3D reconstruction neural network is trained by a training dataset of pairs of 2D endoscopic images that define an input image corresponding to 3D coordinate values calculated for the pixels of the 2D endoscopic images calculated by a 3D reconstruction process that defines the ground truth.
[0021] In a further embodiment of the second aspect, receiving a display of at least one anatomical landmark of the colon, wherein the at least one anatomical landmark divides the colon into a plurality of parts, further includes receiving, tracking the 3D position of the endoscopic camera with respect to the at least one anatomical landmark, calculating the amount of time spent by the endoscopic camera in each of the plurality of parts of the colon, and generating instructions for presenting the amount of time spent by the endoscopic camera in each of the plurality of parts of the colon within the GUI.
[0022] In a further embodiment of the third aspect, the display of the region of at least one 2D image depicting at least one polyp is output by a detection neural network trained to segment the polyp within the 2D image.
[0023] In a further embodiment of the third aspect, the 3D reconstruction of at least one 2D image is provided to a detection neural network in combination with at least one 2D image in order to output a display of an area depicting at least one polyp.
[0024] Unless defined otherwise, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of the present invention, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. Further, the materials, methods, and examples are illustrative only and not intended to be limiting necessarily.
Brief Description of the Drawings
[0025] Some embodiments of the present invention are described herein by way of example only, with reference to the accompanying drawings. It is emphasized that the details shown with specific reference to the drawings are for purposes of illustration and for an exemplary discussion of embodiments of the present invention. In this regard, the description together with the drawings will make apparent to those skilled in the art how embodiments of the present invention can be practiced.
[0026]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
[0027] In some embodiments, the present invention relates to colonoscopy, and more specifically, to systems and methods for processing colon images and videos and / or for processing colon polyps automatically detected during a colonoscopy procedure, but is not limited thereto.
[0028] As used herein, the terms "image" and "frame" are sometimes interchangeable. An image captured by a colonoscope camera may be an individual frame of a video captured by the camera.
[0029] As used herein, the terms "endoscope" and "colonoscope" are sometimes interchangeable.
[0030] Aspects of some embodiments of the present invention relate to a system, method, apparatus, and / or code instructions (i.e., stored on a memory and executable by one or more hardware processors) for generating instructions to present a graphical user interface (GUI) for dynamically tracking one or more polyps in a two-dimensional (2D), optionally colored, endoscopic image of a patient's colon captured by a camera of an endoscope positioned within the lumen of the colon, for example, during a colonoscopy procedure. The position of the region depicting one or more polyps (e.g., region of interest (ROI)) in the current endoscopic image is optionally tracked with respect to one or more previous endoscopic images based on matching visual features between the current image and the previous images, such as features extracted based on a speeded-up robust features (SURF) process. For the current image, when the position of the ROI depicting the polyp is determined to be located outside the boundaries of the image, a vector is calculated. The vector points from a position within the current image to the position of the ROI located outside the current image. The position within the current image may be, for example, the on-screen position of the ROI in a previous image where the ROI was positioned within the image. An augmented endoscopic image is created by augmenting the current endoscopic image with a display of the vector, for example, by introducing the display of the vector as a GUI element onto the endoscopic image and / or as an overlay. Instructions are generated to present the augmented endoscopic image on a display within the GUI.
[0031] The display of the vector depicts a direction and / or orientation for adjustment of the endoscope camera to capture another endoscopic image depicting the ROI of the polyp. The augmented endoscopic image is augmented with the display of the vector for display within the GUI.
[0032] The display of the vector may be an arrow pointing to the position of the ROI outside the image. Moving the camera in the direction of the arrow restores the polyp within the image.
[0033] This process is repeated for the captured images to assist the operator in keeping the polyp within the image. When the camera moves and the polyp no longer appears in the current image, the vector display instructs the operator on how to manipulate the camera to recapture the polyp within the image.
[0034] Aspects of some embodiments of the present invention relate to a system, method, apparatus, and / or code instructions (i.e., stored on a memory and executable by one or more hardware processors) that generate instructions for dynamically tracking the 3D movement of an endoscope camera that captures endoscopic images within a patient's colon. The captured 2D images (e.g., each image, or every few images, e.g., every 3 images, every 4 images, or every other number of images) are fed into a 3D reconstruction neural network, optionally a convolutional neural network (CNN). The 3D reconstruction neural network outputs a 3D reconstruction of each 2D endoscopic image. 3D coordinates are assigned to the pixels of the 2D endoscopic images. The current 3D position of the endoscope camera (i.e., the endoscope) within the colon is calculated according to the 3D reconstruction. For example, the 3D position of the endoscope is determined based on the values of the 3D coordinates of the current image. Instructions are generated to present the current 3D position of the endoscope camera on a colon map within the GUI. The colon map depicts a virtual map of the patient's colon.
[0035] The 3D position of the endoscope is tracked and plotted as a trajectory on the colon map, for example, to track the path of the endoscope within the colon during a colonoscopy procedure.
[0036] The forward and reverse directions of the endoscope may be marked, for example, by arrows and / or color coding.
[0037] The 3D position of the detected polyp can be marked on the colon map. The 3D position of the detected polyp can be tracked relative to the 3D position of the camera. If the distance between the camera and the polyp is less than a threshold value, instructions may be generated to present a display within the GUI. The display may be, for example, the marking of the polyp if the polyp is present in the image, an arrow indicating the position of the polyp if the polyp is not depicted in the image, and / or optionally, a message within the instructions on how to move the camera to capture an image depicting the polyp, indicating that the camera is in proximity to the polyp.
[0038] Polyps surgically removed from the colon may be marked on the colon map.
[0039] Optionally, a portion of the inner surface of the colon depicted within the image is analyzed. For example, it is based on virtually dividing the inner surface into quarters. When a colonoscope is used to visually scan the inner wall of the colon, optionally, for example, when the colonoscope is being withdrawn from the colon (or being moved forward within the colon), in a helical motion, the extent of the portion is cumulatively tracked. The helical motion is performed, for example, by the clockwise (or counterclockwise) orientation of the camera when the colonoscope is being slowly withdrawn (or pushed in). Alternatively, the inner portion of the colon is imaged stepwise, for example, by pulling back (or pushing forward) the colonoscope by a certain distance, stopping the reverse (or forward) movement of the camera, and imaging the surroundings by turning the camera in a circular pattern (or transverse pattern), where pulling back (or pushing forward), stopping, and imaging are repeated over the length of the colon. Optionally, when a colonoscope is used to visually scan the inner wall of the colon, each portion (e.g., quarter) is mostly represented by one or more images. Estimates of the depicted inner surface and / or the remaining inner surface (e.g., quarter) may be generated, and instructions for presentation within the GUI may be generated. The estimation may be performed in real-time, for example, as a global estimate for the entire colon (or a majority thereof) based on each quarter and / or the set of coverages of each respective portion (e.g., quarter), and the previously covered and / or remaining coverage of the inner surface of the colon helps ensure that the entire inner surface of the colon is captured in the image and reduces the risk of a polyp being missed.
[0040] Optionally, based on 3D tracking of the colonoscope, the amount of time spent by the colonoscope in one or more defined portions of the colon is calculated. Instructions for time presentation may be generated for presentation within the GUI, for example, the amount of time spent in each portion of the colon is presented on the corresponding portion of the colon map.
[0041] Aspects of some embodiments of the present invention relate to a system, method, apparatus, and / or code instructions (i.e., stored on a memory and executable by one or more hardware processors) that generate instructions for calculating the dimensions (e.g., size) of a polyp. The dimensions may be 2D dimensions, such as the area and / or radius of a flat polyp, and / or 3D dimensions, such as the volume and / or radius of a raised polyp. The display of the area of the 2D image depicting the polyp is manually depicted by an operator (e.g., using a GUI) and / or output by a detection neural network trained to segment the polyp in the 2D image and supplied with the 2D image. The 2D image is supplied to a 3D reconstruction neural network that outputs 3D coordinates for the pixels of the 2D image. The dimensions of the polyp are calculated according to an analysis of the 3D coordinates of the pixels of the ROI of the 2D image depicting the polyp.
[0042] Optionally, instructions for presenting a warning within the GUI are generated when the dimensions of the polyp exceed a threshold. The threshold may define the minimum dimension of the polyp to be removed. Polyps having dimensions less than the threshold may be left in place.
[0043] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein relate to the medical problem of treating patients, particularly for identifying and removing polyps within a patient's colon. As is apparent from the fact that these missed lesions are ultimately detected in an interval colonoscopy using standard colonoscopy techniques, adenomas can be missed in up to 20% of cases and cancer can be missed in approximately 0.6% of cases. The adenoma detection rate (ADR) varies and depends on patient risk factors, physician performance, and instrumental limitations. The individual anatomical structure of the patient and the quality of bowel preparation are important determinants of a high-quality colonoscopy. The performance of a high-quality colonoscopy by a physician is influenced by factors such as successful cecal intubation, careful visualization during an extended withdrawal time, and overall endoscopic experience. Fatigue and inattention of the endoscopist are risk factors for the physician to miss polyps, while an earlier procedure start time in a session correlates with better outcomes. It should be noted that the adenoma detection rate (ADR) increased in facilities that implemented quality improvement programs, but it was the awareness of monitoring or simply being observed that had a better impact on the adenoma detection rate (ADR).
[0044] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein improve the detection and / or removal rate of polyps during a colonoscopy procedure. The improvement is at least partially facilitated by the GUI described herein, which (i) displays a previously identified polyp that has disappeared from the currently captured image by an arrow indicating the operating direction of the colonoscopy camera to recapture an image of the polyp, by an arrow indicating the direction to recapture an image of the polyp, (ii) helps to present and update a colon map that displays the 2D and / or 3D positions of the identified polyps, ensuring that all identified polyps are evaluated and / or removed, (iii) tracks portions of the inner circumference of the colon to identify portions of the inner surface that have not been captured by an image and thus have not been analyzed to identify polyps, ensuring that there are no portions of the colon that have not been imaged and where polyps may have been missed, (iv) calculates the volume of a polyp to assist in providing data to help determine which polyps to remove and / or assist in a cancer diagnosis, and / or (v) calculates the amount of time spent by the colonoscope in each part of the colon, helping to instruct the operator to do so. The GUI is presented and adapted in real time with respect to the images captured during the colonoscopy procedure and can provide real-time feedback to assist a physician operator in improving the detection and / or removal rate of polyps.
[0045] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein address the technical problem of an apparatus for improving the detection and / or removal rate of polyps. Specifically, at least some implementations of the systems, methods, apparatuses, and / or code instructions described herein analyze the captured images to help increase the identification and / or detection rate of polyps, and / or improve the image processing technology and / or GUI technology by the code used and / or the GUI used by the operator. For example, as compared with a standard approach. For example, an optical system that achieves a wider field of view and improves image resolution, and a distal colon endoscope attachment such as a balloon cap or ring for improving visualization behind the mucosal folds. Such optical devices and attachment devices are passive and rely on the operator's skills when tracking the identified polyps. In contrast, the GUI described herein automatically tracks the identified polyps.
[0046] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein address the technical problem of calculating the volume of a polyp. Based on standard practice, the size of a polyp is measured only after the polyp has been removed from the patient, as described, for example, in Kume, Keiichiro, et al. “Endoscopic measurement of poly size using a novel calibrated hood.” Gastroenterology research and practice 2014 (2014). The importance of polyp size measurement is described, for example, in Summers, Ronald M. “Polyp size measurement at CT colonography: What do we know and what do we need to know?” Radiology 255.3 (2010): 707-720. In contrast, at least some of the systems, methods, apparatuses, and / or code instructions described herein calculate the size of a polyp in vivo while the polyp is attached to the colon wall, before the polyp is removed. Calculating the volume of a polyp before it is removed can provide several advantages. For example, polyps that exceed a threshold volume are targeted for removal and / or polyps that are below the threshold volume are left in the patient's body. The volume of the polyp calculated before removal may be compared to the volume after removal, for example, to determine whether the entire polyp has been removed and / or to compare the volume of the polyp on the surface to the unseen portion of the polyp below the surface as a measure of the risk of cancer and / or to help grade the risk of the polyp and / or cancer.
[0047] At least some implementations of the systems, methods, apparatuses, and / or code instructions described herein address the technical challenges of neural network processing that is slower than the rate of images captured into the video by the camera of a colonoscope. The processes described herein for tracking polyps (and / or associated ROIs) based on the extracted features compensate for the delay in the process for detecting polyps using a neural network that outputs data for the images, and optionally the display and / or location of the detected polyps. The neural network-based detection process is computationally more expensive than the feature extraction and tracking processes (e.g., 25 milliseconds (ms) - 40 ms on a typical personal computer (PC) with an I7 Intel processor and an Nvidia GTK 1080 TI GPU, while the typical time difference between consecutive frames is in the range of 20 ms - 40 ms). Thus, a delay scenario can be created in the sense that the detection result output by the neural network for the frame number indicated by i is ready only when subsequent frames (e.g., the number indicated by i+2, or later frames) have already been presented. Such a delay scenario results in a strange situation when the first frame depicts a polyp but later frames do not (e.g., the camera has been shifted to a position where the polyp is not captured in the image), and the delay in the neural network for detecting polyps is available only when the polyp is no longer depicted, creating a situation where the display of the detected polyp is provided when the presented image does not present the polyp. Note that if frame number i+2 is available for presentation, the delay in the presentation of the frame is unacceptable from a clinical and / or regulatory perspective and can result in losses, for example, in attempts to remove an imaged polyp, so the frame must be presented as soon as it is available (e.g., in real time and / or immediately).Tracking based on computationally efficient features that result in rapid processing, as compared to neural network-based processing (e.g., less than 10 ms on a typical PC with an I7 Intel processor), is used to transform the ROI depicting the polyp detected in frame i to a position within frame i+2. The transformed position is the position presented on the display at frame number i+2. Optionally, the bounding box of the ROI (e.g., only the bounding box of the ROI) is transformed to frame i+2 because the contour transformation can become less accurate due to the influence of different 3D positions that are not necessarily considered in 2D transformation. Note that frame i+2 is an example, and as another example, for instance, i+1, i+3, i+4, i+5 or more may be used and is not necessarily limited to this.
[0048] Before describing at least one embodiment of the present invention in detail, it should be understood that the present invention is not necessarily limited to the details of the construction and arrangement of the components and / or methods described in the following description and / or shown in the drawings and / or examples in its application. The present invention is capable of other embodiments or of being practiced or carried out in various ways.
[0049] The present invention can be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0050] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPRPM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as raised structures within grooves in which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed to be a signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0051] The computer-readable program instructions described herein can be downloaded to respective computing / processing devices from a computer-readable storage medium or through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network, to an external computer or an external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0052] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer, partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions to personalize the electronic circuit and utilize the state information of the computer-readable program instructions to carry out aspects of the present invention.
[0053] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0054] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, and the computer-readable storage medium having the instructions stored therein includes a product comprising instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0055] Alternatively, the computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device, such that the instructions which execute on the computer, other programmable apparatus, or other device create a computer-implemented process for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0056] The flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible embodiments of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special purpose hardware-based systems that perform a specified function or act, or combinations of special purpose hardware and computer instructions.
[0057] Next, referring to FIG. 1, FIG. 1 is a flowchart of a method for processing an image acquired by a camera positioned on an endoscope within the colon of a target patient, tracking a polyp, tracking the movement of the camera, mapping the position of the polyp, calculating the coverage of the inner surface of the colon, tracking the amount of time spent in different portions of the colon, and / or calculating the volume of the polyp, according to some embodiments of the present invention. Referring also to FIG. 2, FIG. 2 is a block diagram of the components of a system 200 for processing an image acquired by a camera positioned on an endoscope within the colon of a target patient, tracking a polyp, tracking the movement of the camera, mapping the position of the polyp, calculating the coverage of the inner surface of the colon, tracking the amount of time spent in different portions of the colon, and / or calculating the volume of the polyp, according to some embodiments of the present invention. System 200 can optionally perform the operations of the method described with reference to FIG. 1 by hardware processor 202 of computing device 204 that executes code instructions stored in memory 206.
[0058] The imaging probe 212, for example, a camera disposed on a colonoscope, captures images within a patient's colon obtained, for example, during a colonoscopy procedure. The colon images are optionally 2D images and optionally color images. The colon images may be acquired as a streamed video and / or a sequence of still images. The captured images may be processed in real-time and / or (e.g., after the procedure is completed) offline.
[0059] The captured images may be stored in the image repository 214 and optionally implemented as an image server, for example, a Picture Archiving and Communication System (PACS) server, and / or an Electronic Health Record (EHR) server. The image repository may communicate with the network 210.
[0060] The computing device 204 receives the captured images, for example, directly in real-time from the imaging probe 212 and / or (e.g., in real-time or offline) from the image repository 214. Real-time images may be received during a colonoscopy procedure to guide an operator as described herein. The captured images can be received by the computing device 204 via one or more imaging interfaces 220, for example, a wire connection (e.g., the output from the imaging probe 212 is plugged into the imaging interface via a connecting wire), a wireless connection (e.g., an antenna), a local bus, a port for connecting a data storage device, a network interface card, other physical interface implementations, and / or a virtual interface (e.g., a software interface, a Virtual Private Network (VPN) connection, an Application Programming Interface (API), a Software Development Kit (SDK)).
[0061] As described herein, computing device 204 analyzes the captured image and generates instructions for dynamically adjusting a graphical user interface presented on a user interface (e.g., a display) 226. For example, as described herein, elements of the GUI are injected as an overlay on the captured image and presented on the display.
[0062] Computing device 204 may be implemented, for example, as a dedicated device, a client terminal, a server, a virtual server, a colonoscopy workstation, a gastrointestinal epidemiology workstation, a virtual machine, a computing cloud, a mobile device, a desktop computer, a thin client, a smartphone, a tablet computer, a laptop computer, a wearable computer, a glasses computer, and a watch computer. Computing 204 may include a gastrointestinal epidemiology and / or colonoscopy workstation, and / or other devices that enable an operator to view a GUI created from the processing of colonoscopy images, such as a real-time presentation that points an arrow towards a polyp not currently visible on the image, and / or a colon map that presents the 2D and / or 3D position of the polyp, and / or an advanced visualization workstation that may include other features described herein.
[0063] Computing device 204 can perform one or more operations represented with reference to FIG. 1 and / or operate as one or more servers (e.g., a network server, a web server, a computing cloud, a virtual server), and include locally stored software that provides services (e.g., one or more of the operations described with reference to FIG. 1) to one or more client terminals 208 (e.g., a client terminal used by a user to view colonoscopy images, e.g., a colonoscopy workstation that includes a display presenting images captured by a colonoscope 212, a remotely located colonoscopy workstation, a PACS server, a remote EHR server, a remotely located display for remote viewing of procedures by a resident, etc.). The services can provide software as a service (SaaS) to the client terminal 208 via the network 210, for example, provide an application for local download as an add-on to a web browser and / or a colonoscopy application to the client terminal 208, and / or provide functionality to the client terminal 208 using a remote access session via, for example, a web browser, an application programming interface (API), and / or a software development kit (SDK), etc., for injection of GUI elements into colonoscopy images and / or presentation of colonoscopy images within the GUI.
[0064] Different architectures of system 200 may be implemented. Examples are shown next. *The computing device 204 is connected between the imaging probe 212 and the display 226 and is, for example, a component of a colonoscopy workstation. Such an implementation may be used for real-time processing of images captured by a colonoscope during a colonoscopy procedure and real-time presentation of the GUIs described herein on the display 226, for example, injecting GUI elements and / or presenting images within the GUI, for example, presenting direction arrows to a currently unseen polyp and / or presenting a colon map indicating the 2D and / or 3D position of the polyp and / or other features described herein. In such embodiments, the computing device 204 may be installed for each colonoscopy workstation (e.g., including the imaging probe 212 and / or the display 226). *The computing device 204 functions as a central server and provides services to a plurality of colonoscopy workstations such as client terminals 208 (e.g., including the imaging probe 212 and / or the display 226) via the network 210. In such embodiments, a single computing device 204 can be installed to provide services to a plurality of colonoscopy workstations. *The computing device 204 is installed as code on an existing device such as a server 218 (e.g., a PACS server, an EHR server) and provides local offline processing for each device, such as offline analysis of colonoscopy videos captured by different operators and stored in the PACS and / or EHR server. The computing device 204 can be installed on an external device that communicates with the server 218 (e.g., a PACS server, an EHR server) via the network 210 to provide local offline processing for a plurality of devices.
[0065] The client terminal 208 may be implemented as a colonoscopy workstation that includes, for example, an imaging probe 212 and a display 226, a desktop computer (e.g., running a viewer application for viewing colonoscopy images), a mobile device (e.g., a laptop, smartphone, glasses, wearable device), and a remote server for remotely viewing colonoscopy images.
[0066] The hardware processor 202 can be implemented, for example, as a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), a digital signal processor (DSP), and an application-specific integrated circuit (ASIC). The processor 202 may include one or more processors (of the same or different types) that can be arranged as a cluster and / or as one or more multi-core processing units for parallel processing.
[0067] The memory 206 (also referred to herein as a program storage device and / or a data storage device) stores code instructions for execution by the hardware processor 202, such as a random access memory, a read-only memory, and / or a storage device, such as a non-volatile memory, a magnetic medium, a semiconductor memory device, a hard drive, a removable storage device, and an optical medium (e.g., DVD, CD-ROM). For example, the memory 206 can store code 206A that implements one or more operations and / or features of the methods described with reference to FIG. 1, and / or GUI code 206B that generates instructions for presentation within the GUI and / or presents the GUI described herein based on the instructions (e.g., injection of GUI elements into colonoscopy images, overlay of GUI elements on images, presentation of colonoscopy images within the GUI, and / or presentation and dynamic update of the colon maps described herein).
[0068] The computing device 204 can include a data storage device 222 for storing data, such as received colonoscopy images, colon maps, and / or processed colonoscopy images presented within the GUI. The data storage device 222 can be implemented, for example, as a memory, a local hard drive, a removable storage device, an optical disk, a storage device, and / or a remote server and / or a computing cloud (e.g., accessed via network 210).
[0069] The computing device 204 can include one or more of a data interface 224 for connecting to network 210, optionally a network interface, such as a network interface card, a wireless interface for connecting to a wireless network, a physical interface for connecting to a cable for network connection, a virtual interface implemented in software, network communication software providing a higher layer of network connection, and / or other implementations. The computing device 204 can use network 210 to access one or more remote servers 218, for example, to download updated imaging processing code, updated GUI code, and / or to obtain images for offline processing.
[0070] The imaging interface 220 and the data interface 224 can be implemented as a single interface (e.g., a network interface, a single software interface), and / or as a software interface (e.g., an API, a network port) and / or a hardware interface (e.g., two network interfaces), and / or as a combination thereof (e.g., a single network interface, and two software interfaces, two virtual interfaces on a common physical interface, a virtual network on a common network port), i.e., as two independent interfaces. It should be noted that the term / component imaging interface 220 may sometimes be exchanged with the term / data interface 224.
[0071] The computing device 204 can communicate with one or more of the server 218, the imaging probe 212, the image repository 214, and / or the client terminal 208 using a network 210 (or another communication channel), for example, via a direct link (e.g., a cable, wireless) and / or an indirect link (e.g., via an intermediate computing device such as a server and / or via a storage device) according to different architecture implementations described herein.
[0072] The imaging probe 212 and / or the computing device 204 and / or the client terminal 208 and / or the server 218 includes or communicates with a user interface 226 that includes a mechanism designed for a user to input data (e.g., marking a polyp for removal) and / or view a GUI including colonoscopy images, direction arrows, and / or a colon map. Exemplary user interfaces 226 include, for example, one or more of a touch screen, a display, a keyboard, a mouse, augmented reality glasses, and voice-activated software using a speaker and a microphone.
[0073] At 100, the endoscope is inserted and / or moved within the patient's colon. For example, the endoscope is advanced forward (i.e., from the rectum to the cecum), retracted (e.g., from the cecum to the rectum), and / or the orientation of at least the camera of the endoscope is adjusted (e.g., up, down, left, right), and / or the endoscope is left in place.
[0074] The endoscope may be adjusted based on the GUI, for example, manually by an operator and / or automatically by a user. For example, it should be noted that the user may adjust the camera of the endoscope according to the presented arrows to recapture a polyp that has moved out of the image.
[0075] At 102, the image is captured by the camera of the endoscope. The image is optionally a color 2D image. The image depicts the inside of the colon and may or may not depict a polyp.
[0076] The image can be captured as a video stream. Individual frames of the video stream can be analyzed.
[0077] The images may be analyzed individually and / or as a set of consecutive images as described herein. Each image in the sequence may be analyzed, or some intermediate images may be ignored. Optionally, a predetermined number of images, e.g., every third image, are analyzed and the two intermediate images are ignored.
[0078] Optionally, one or more polyps depicted in the image are treated. The polyp may be treated via the endoscope. The polyp may be treated by its surgical removal, for example, for sending to a pathology laboratory. The polyp may be treated by ablation.
[0079] Optionally, the treated polyp is marked manually, e.g., by a physician (e.g., by making a selection using a GUI, by pressing a "polyp removal" icon), and / or automatically by code (e.g., detecting the movement of a surgical resection device). As described herein, the marked treated polyp may be tracked and / or presented on a colon map presented on the GUI.
[0080] In 104, the image is supplied to a detection neural network that outputs an indication of whether a polyp is depicted in the image (or not). The detection neural network can include a segmentation process that identifies the location of the detected polyp in the image, e.g., by generating a bounding box and / or other contours depicting the polyp within a 2D frame.
[0081] An exemplary neural network-based process for polyp detection is the Automated Polyp Detection System (APDS) described with reference to International Patent Application Publication (WO2017 / 042812) "A SYSTEM AND METHOD FOR DETECTION OF SUSPICIOUS TISSUE REGIONS IN AN ENDOSCOPIC PROCEDURE" by the same inventor as this application.
[0082] The automated polyp detection process performed by the detection neural network may be executed in parallel with and / or independently of features 106 - 114, e.g., on the same computing device, and / or on a processor, and / or on another real-time connected computing device, and / or on a platform connected to a computing device that executes the features described with reference to 106 - 114.
[0083] The output of the neural network is calculated and provided continuously for each endoscopic image relative to the previous endoscopic image, and if the tracked position of the region depicting the polyp in each endoscopic image (i.e., as described with reference to 106) is at a different position than the region output by the neural network for the previous endoscopic image, an augmented image is generated for each endoscopic image based on the calculated tracked position of the polyp. Such a situation occurs when the frame rate of the images is faster than the processing rate of the detection neural network. The detection neural network completes the processing of the images after one or more consecutive images have been captured. In such a case, when the results of the detection neural network are used, the polyp positions calculated for the older images may not necessarily reflect the polyp positions for the current image.
[0084] Optionally, one or more endoscopic images of a consecutive subset of endoscopic images including each endoscopic image and one or more images positioned continuously prior to each endoscopic image depicting the tracked ROI (as described with reference to 106) (e.g., captured prior to each endoscopic image) are supplied to the detection neural network. The consecutive subset of images can be supplied to the detection neural network in parallel with the tracking process (as described with reference to 106). Alternatively, the subset of images is first processed by the tracking process as described with reference to 106. One or more of the post-processed images may be translated and / or rotated to generate a subset of endoscopic images in which the regions depicting the polyps (e.g., ROIs) detected by the tracking process are at the same approximate position in all images, e.g., at the same pixel position on the display for all images. The consecutive subset of processed images is supplied to the detection neural network to output the current region depicting the polyp. The images may be augmented with the regions detected by the detection neural network.
[0085] Alternatively or additionally, the calculated tracking position of the region depicting the polyp in the image, and / or the output of the tracking process described with reference to 106 (e.g., the 2D transformation matrix between consecutive frames) is supplied to the detection neural network. The tracked position and / or the 2D transformation matrix may be supplied to the neural network when the tracked position of the polyp is within the image, or when the tacked position of the polyp is outside the image. The tracked position may be supplied to the neural network alone or in addition to one or more images (e.g., the current image and / or a previous image). The output of the tracking process (e.g., the tracked position and / or the 2D transformation matrix) may be used by the neural network process, for example, to improve the accuracy of the correlation between detections of the polyp in consecutive frames. The output of the tracking process can increase the reliability of the neural network-based polyp detection process (e.g., when the tracked polyp was detected in a previous frame), and / or can reduce the case of false positive detections.
[0086] Optionally, in the case of a discrepancy between the tracked position of the polyp calculated based on 106 (e.g., the ROI depicting the polyp) and the output of the detection neural network, the position of the neural network is used. The position output by the neural network is used to generate instructions for generating an augmented image augmented with the display of the position of the polyp. The position of the polyp output by the neural network can be considered to be more reliable than the tracking position calculated as described with reference to 106, but the position calculated by tracking is more computationally efficient and / or can be executed in a shorter time than the processing by the neural network.
[0087] Optionally, the 3D reconstruction of the 2D image described with reference to 108 is supplied to the detection neural network alone and / or in combination with the 2D image to output a display of the region depicting at least one polyp.
[0088] Referring back to FIG. 1, at 106, the position of the region (e.g., ROI) depicting one or more polyps is tracked within the current endoscopic image relative to one or more previous endoscopic images.
[0089] It should be noted that the movement of the camera (e.g., orientation, forward, backward) is tracked indirectly by tracking the movement of the ROI between images, because the camera is moving while the polyps remain stationary at their positions within the colon. It should be noted that some movement of the ROI between frames can be due to peristalsis of the colon itself and / or other natural movements, regardless of whether the camera is stationary or moving.
[0090] Optionally, the position is tracked in 2D. The polyps can be tracked by tracking the ROI that depicts the contour of the polyps. The ROI and / or the polyps may be detected in one or more previous images by the detection neural network of feature 104.
[0091] The position of the polyps is tracked even if the polyps are not depicted in the current image. For example, the camera is positioned such that the polyps are no longer present in the image captured by the camera.
[0092] Optionally, the vector is calculated from the position within the current image to the position of the polyp and / or ROI located outside the current image. The vector can be calculated, for example, from the position of the ROI on the last (or previous) image depicting the ROI, from the center of the screen, from the center of the quadrant of the screen closest to the position of the external ROI, and / or from another region of the image closest to the position of the external ROI (e.g., a predetermined distance away from the position at the boundary of the image closest to the position of the external ROI).
[0093] The display of the vector can depict the direction and / or orientation for adjustment of the endoscopic camera to capture another endoscopic image depicting the region of the image.
[0094] Optionally, the tracking algorithm is feature-based. The features can be extracted from the analysis of endoscopic images. The features may be extracted based on the speeded-up robust features (SURF) extraction approach, as described in Bay, Herbert, Tinne Tuytelaars, and Luc Van Gool, "Surf: Speeded up robust features" (European conference on computer vision. Springer, Berlin, Heidelberg, 2006.). The tracking of the extracted features may be performed in two dimensions (2D) between consecutive images (note that one or more intermediate images between the analyzed images may be skipped, i.e., ignored). The features can be matched based on, for example, the K-d tree approach described in Silpa-Anan, Chanop, and Richard Hartley, "Optimised KD-trees for fast image descriptor matching (2008): 1-8". The K-d tree approach may be selected based on the observation that, for example, the main movement in a colonoscopy procedure is the movement of the endoscopic camera in the colon. The features may be matched according to their descriptors between consecutive images.The best homography is estimated by calculating the 2D affine transformation matrix (and its closest 2D geometric transformation matrix) of the camera motion from frame to frame, for example, using the Random Sample Consensus (RANSAC) method described in "Detecting planar homographies in an image pair. ISPA 2001. Proceedings of the 2nd International Symposium on Image and Signal Processing and Analysis. In conjunction with 23rd International Conference on Information Technology Interfaces (IEEE Cat. IEEE, 2001)" by Vincent, Etienne, and Robert Laganiere, and referring to "A survey of planar homography estimation techniques. Centre for Visual Information Technology, Tech. Rep. IIIT / TR / 2005 / 12(2005)" by Agarwal, Anubhav, C.V, Jawahar, and P.J. Narayanan.
[0095] A specific region of interest (ROI) can be tracked, and optionally, a bounding box indicating the position of one or more polyps within it can be tracked. The position of the ROI is tracked while the ROI moves from an image (e.g., video) frame. The ROI may be tracked until the ROI returns (i.e., is redrawn) in the current frame. The ROI may be continuously tracked across the image frame boundary, for example, based on a tracking coordinate system defined externally to the image frame.
[0096] From the time when a polyp is detected (e.g., automatically by code and / or manually by an operator), until the movement of the camera is stopped (e.g., the operator focuses on the detected polyp), during the time interval, the polyp may move out of the frame. An instruction may be generated to present a display of the location where the polyp (or the ROI related to the polyp) is currently located outside the presented image. For example, by augmenting the current image with the presentation of directional arrows. The arrows indicate the location where the operator should move the camera (i.e., the endoscope tip) to return the polyp (or the ROI of the polyp) redrawn in the new frame.
[0097] Optionally, the arrow (or other display) may be presented until the camera is too far from the polyp being tracked (e.g., greater than a defined threshold). The predetermined threshold may be defined, for example, as the distance of the new position of the polyp from the center of the current frame that exceeds three times the length of the diagonal of the frame (in pixel units), or as other values.
[0098] Next, referring to FIG. 3, FIG. 3 is a flowchart of a process for tracking the position of the region depicting one or more polyps in each endoscopic image relative to one or more previous endoscopic images according to some embodiments of the present invention.
[0099] At 302, a new (i.e., current) image, optionally a new frame of video captured by a colon endoscope camera, is received. The image may or may not depict one or more polyps being tracked.
[0100] At 304, irrelevant or misleading portions are removed from the image, e.g., the periphery that is not part of the colon, the lumen in the colon image that is a dark region where light does not return to the camera (e.g., representing the far central portion of the colon) depending on a light source that moves with the camera, and / or inconsistent reflections in consecutive frames.
[0101] At 306, a contrast limited adaptive histogram equalization (CLAHE) process (e.g., as described in Reza, Ali M.'s "Realization of the contrast limited adaptive histogram equalization (CLAHE) for real-time image enhancement. Journal of VLSI signal processing systems for signal, image and video technology 38.1 (2004):35 - 44), and / or other processes are implemented to enhance the image. The speeded up robust features (SURF) technique is optionally implemented to extract features immediately (e.g., without significant delay) after CLAHE.
[0102] At 308, if the current frame is the first frame of the sequence and / or the first frame in which a polyp is detected, 302 is repeated to obtain the next frame. If the current frame is not the first frame, 310 is implemented to process the current frame using the previous frame processed in the previous iteration.
[0103] At 310, a geometric matrix is optionally calculated based on a homography for camera movement between two consecutive frames (e.g., the frame numbers indicated by i and i + n, i.e., a jump of n frames between the two frames processed together).
[0104] At 312, if no transformation is found, the next image (e.g., frame) in the sequence is processed with the previous (i.e., current) frame (i.e., when frames i and i + n do not match and an attempt is made to match frames i and i + n + 1) (by repeating 302), otherwise, 314 is executed.
[0105] At 314, if a transformation between two or more previous frames cannot be found (i.e., the position of the tracked object (i.e., polyp) in the last frame is not calculated and / or cannot be calculated), 318 is performed; otherwise, if the position of the polyp is calculated, 316 is performed.
[0106] At 316, when an object (i.e., polyp) is identified in the image for tracking, the tracked area is fed forward to the current frame according to the transformation found at 310. 322 is implemented to generate instructions for depicting the object (i.e., polyp) in the current frame, for example, by visual markings such as boxes.
[0107] Alternatively, at 318, the transformation of the intermediate frame where the transformation was not found is interpolated according to the last transformation found between two frames before and after the intermediate frame.
[0108] At 320, if the tracked object (i.e., polyp) exists, the tracked area is fed forward to the intermediate frame according to the interpolated transformation found at 318. 322 is implemented to generate instructions for depicting the object (i.e., polyp) in these frames. 316 may also be implemented to depict the object (i.e., polyp) on the current frame.
[0109] At 324, the process is repeated to track when the object (i.e., polyp) is outside the frame. The next frame to be processed may be a jump step, for example, two or more frames are ignored until a new frame is processed or until each frame is processed.
[0110] Referring now to FIG. 4, FIG. 4 is a schematic diagram showing an example of a K-D tree match between two consecutive images 402A-B that are matched together by corresponding features (features shown as plus symbols, one labeled 404, matching features marked using lines, and one shown as 406) for tracking polyps within regions 408A-B, according to some embodiments of the present invention. Image 402A can be denoted as frame number i. Image 402B can be denoted as frame number i+3. Region E08A shows the original bounding box of the region, and region E08B shows the transformation of E08A according to a transformation matrix calculated by the homography of the matches between keypoints, as described herein.
[0111] Referring now to FIG. 5, FIG. 5 is a geometric transformation matrix 502 between frames 402A-B of FIG. 4, according to some embodiments of the present invention.
[0112] Referring now to FIG. 6, FIG. 6 shows a schematic diagram 602 depicting an ROI 604 of a bounding box showing a detected polyp being tracked, according to some embodiments of the present invention, and a subsequent schematic diagram 606 (e.g., 2 frames later) where the tracked bounding box is no longer drawn within the image. Schematic diagram 606 is augmented with the presentation of an arrow 608 indicating the direction to move the camera to redraw the ROI of the polyp within the captured image.
[0113] Referring now to FIG. 7, FIG. 7 shows a schematic diagram 702 showing a detected polyp 704, according to some embodiments of the present invention, and a subsequent frame (i.e., 3 frames later in time) 706 showing a tracked polyp 708 calculated using the transformation matrix described herein.
[0114] Referring back to FIG. 1. At 108, a 3D reconstruction is calculated for the current 2D image. The 3D reconstruction may define 3D coordinate values for the pixels of the 2D image.
[0115] Optionally, the 3D reconstruction neural network is trained to output a 3D image from a 2D image input. The trained 3D reconstruction neural network generates a more accurate 3D image from the 2D image as compared to the 3D reconstruction process alone (e.g., a standard 3D reconstruction process). The 3D reconstruction process is a process different from the 3D reconstruction neural network. The 3D reconstruction process can be based on, for example, a standard 3D reconstruction process that calculates 3D reconstruction using only a single 2D image. The 3D reconstruction neural network is trained using a training dataset of 2D images captured by a colon endoscope camera and corresponding 3D images created from the 2D images by the 3D reconstruction process. The 2D images are specified as inputs and the reconstructed 3D images are specified as ground truths.
[0116] Optionally, the process of calculating the 3D reconstruction of the current 2D image is based on a 3D neural network that outputs the 3D reconstruction. The 3D reconstruction neural network can be trained using a pair of training datasets of 2D endoscope images that define the input image and corresponding 3D coordinate values calculated for the pixels of the 2D endoscope images calculated by the 3D reconstruction process that defines the ground truth. A neural network that can be trained on a large training dataset (e.g., on the order of 10,000 - 100,000 or 100,000 - 1,000,000) can provide increased accuracy as compared to the 3D reconstruction process alone.
[0117] The neural network may be trained using pairs of in-vivo images of the colon captured by an endoscopic camera (shown as the input) and corresponding 3D values estimated by a 3D reconstruction process (shown as the ground truth). The 3D values may be 3D coordinates calculated for the 2D pixels of the image. The 3D reconstruction process may be trained using a large dataset (e.g., at least 100,000 images of the colon from at least 100 different colonoscopy videos, or other smaller or larger values).
[0118] An exemplary process for 3D reconstruction of 2D colon images is described. The 3D geometry of the partial internal surface area of the colon may be reconstructed for each frame using, for example, the "Shape-from-shading (SfS)" process described in Zhang, Ruo, et al., "Shape-from-shading: a survey", IEEE transactions on pattern analysis and machine intelligence 21.8 (1999): 690-706, and / or Prasath, V.B., Surya, et al., "Mucosal region detection and 3D reconstruction in wireless capsule endoscopy videos using active contours", 2012 Annual International Conference of the IEEE Engineering in Medicine and Biology Society, IEEE, 2012. The camera motion parameters can be calculated based on, for example, the Shape from Motion (SfM) process described in Szeliski, Richard, and Sing Bing Kang, "Recovering 3D shape and motion from image streams using nonlinear least squares", Journal of Visual Communication and Image Representation 5.1 (1994): 10-28. Optionally, the 3D positions of one or more feature points can be calculated to integrate the partial surfaces reconstructed by the SfS algorithm. The SfS algorithm processes moving local light and light attenuation. The real situation during endoscopy within a human organ may be mimicked. One exemplary advantage of the 3D reconstruction process described herein compared to other SfS-based processes is that the 3D reconstruction process described herein can calculate a clear reconstructed surface for each frame.The 3D reconstruction process can use intensity thresholds to remove non-Lambertian (e.g., more specular reflection) regions in order to make the SfS process function for other regions. The SfS process implemented by the inventors (described with reference to Prados, E., Faugeras, O., et al., "Shape from shading: a well-posed problem?", IEEE Conference on Computer Vision and Pattern Recognition, 870-877 (2005)) generates a clear surface. The clear surface can be calculated by taking into account the light attenuation of 1 / r. 2 and / or by assuming that the spot light source is attached to the center of the camera's projection, and thus the brightness of the image is
Number
[0119] The described 3D reconstruction process was applied to colonoscopy video frames obtained according to Kaufman A., Wang J. (2008) 3D Surface Reconstruction from Endoscopic Videos. In: Linsen L., Hagen H., Hamann B. (eds) Visualization in Medicine and Life Sciences. Mathematics and Visualization. Springer, Berlin, Heidelberg, and the average reprojection error of the selected feature points is 0.066 pixels.
[0120] The standard 3D process described herein can be used to calculate an estimate of the 3D reconstruction of the colon for each 2D frame. The 2D frames and the corresponding calculated ground truth (i.e., the 3D values of the 3D reconstruction of the 2D frames) can be split to create a training dataset (e.g., 70% of the data) and test the dataset (e.g., 30% of the data). For example, a convolutional neural network (CNN) based on an implementation similar to an encoder-decoder architecture (as described, for example, in Shan, Hongming, et al. 「3-D convolutional encoder-decoder network for low-dose CT via transfer learning from a 2-D trained network」 IEEE transactions on medical imaging 37.6 (2018): 1522-1534) may be trained and tested with the training and validation datasets as it is robust enough to predict its 3D reconstruction for each 2D frame.
[0121] In an exemplary architecture, the encoder and decoder portions of the 3D reconstruction CNN may be constructed from 2D convolutional layers. The last layer may be cubed to predict 3D coordinate values for each pixel of the input image. The 3D reconstruction CNN can be used in real time to predict the 3D values for each pixel within each 2D frame of the colon surface within a colonoscopy video. The CNN is trained on data output by the 3D reconstruction process and optionally uses a number of processed 2D images (e.g., on the order of 100,000, or more or less), but can be more robust and / or accurate than the 3D reconstruction process alone.
[0122] Referring now to FIG. 8, FIG. 8 is a flowchart of a process for 3D reconstruction of 2D images captured by a camera within a patient's colon, according to some embodiments of the present invention. 3D reconstruction can be defined as a ground truth value associated with the 2D image defined as an input value to create a training data set for training a 3D reconstruction neural network, as described herein.
[0123] At 802, the image is optionally provided as a video stream of frames.
[0124] At 804, the current frame, denoted as n, is processed. Optionally, an SfS process for non-specular regions is used to compute a 3D reconstruction of frame n.
[0125] At 806, a plurality of images are processed that include images acquired before and after the current image. The frames may be denoted as n - 2 to n + 2, or other numbers may be used. Using an SfM process, a general 3D reconstruction of frame #n can be computed based on frames #n - 2 to #n + 2, and updated features in progress can be computed based on the estimation of intrinsic camera parameters.
[0126] At 808, the 3D reconstructions calculated at I04 and I06 are integrated for frame n.
[0127] At 810, the 3D reconstruction is provided for each pixel within the 2D frame #n that shows the observed colon region. Each pixel of the 2D frame #n is assigned (x, y, z) values in the 3D coordinate system. The origin of the 3D coordinate system can be correlated with the center point of the frame.
[0128] Next, an exemplary data flow of the 3D reconstruction CNN will be described. In the first phase shown, consecutive 2D frames (i.e., at least 3 odd numbers, the number of consecutive input frames represented by n) are supplied to the 3D CNN and pass through a layer of 3D convolutional kernels (n×3×3). The first layer is designed to process all three color channels of the 2D frame by replicating them three times. The data flow continues through batch normalization and ReLU layers and a max pooling layer that reduces the output size by 2 in each of the frame dimensions. The data passes through an additional four layer groups, using 2D convolutional kernels that represent the encoder part of the 3D CNN. The data flow continues through an additional four layer groups with 2D convolutions, using an upsampling layer instead of the max pooling layer, and then passes through a fifth layer replicated three times that represents the encoder part of the CNN. The output has the same resolution as the input 2D frame. For each 2D pixel, three values that are the 3D coordinate values of the pixel (x, y, z) of the 2D pixel are output.
[0129] Optionally, the calculated 3D values of the pixels of the 2D frame can be supplied to the polyp detection process (e.g., the APDS detection process) described with reference to 104. The calculated 3D values provide additional input (e.g., to a neural network) for detecting polyps in the current frame and / or determining the location of polyps in the current frame. For example, the 3D values may be information indicating flatness and / or protrusion with respect to the surrounding of the colon tissue region suspected of being a polyp.
[0130] Return to FIG. 1 for reference. At 110, the current 3D position of the polyp and / or the endoscope is calculated and / or tracked. The 3D position may be calculated based on the output of the 3D reconstruction output by the 3D reconstruction neural network.
[0131] The 3D position is calculated based on a 3D rigid body transformation matrix calculated between successive 2D endoscopic images, based on the 3D coordinates of matching features extracted from successive 2D endoscopic images and / or from the 3D reconstruction of successive 2D endoscopic images. In other words, the 3D transformation matrix is calculated for the matched features using the 3D coordinates of the pixels corresponding to the features, and uses the 3D coordinates of the features rather than 2D coordinates, for example, to track 2D images based on corresponding features between 2D images, similar to the processes described herein. The 3D body transformation matrix indicates the movement of the endoscope camera between successive 2D endoscopic images. The current 3D position of the endoscope camera is calculated according to the 3D rigid body transformation matrix taking into account each 2D endoscopic image. The trajectory of the movement of the camera may be calculated based on successive 3D transformation matrices. The trajectory of the movement of the camera may be presented on the colon map as described herein.
[0132] Optionally, the detected 3D position of the polyp and / or the endoscope is iteratively tracked according to the 3D positions calculated for a plurality of successive images.
[0133] The 3D tracking process reception may be based on the output of the 2D tracking process described with reference to 106 and / or the 3D reconstruction process described with reference to 108. The feature extraction unit and / or the matching unit can be implemented as described with reference to 2D tracking on the 2D consecutive frames of the colonoscopy video. However, it should be noted that the homography is calculated for the 3D values (x, y, z) of the extracted features (keypoints) rather than for the 2D values. The 3D values are reconstructed by the 3D reconstruction process described herein to find the best-fit 3D affine transformation. The 3D rigid transformation closest to the calculated 3D affine transformation is calculated using, for example, the process described with reference to Yuan, Jie et al., "Application of Feature Point Detection and Matching in 3D Objects Reconstruction", PATTERNS 2011: the Third International Conferences on Pervasive Patterns and Applications, 19 - 24. The 3D rigid transformation matrix describes the 3D movement of the camera in the colon from frame to frame.
[0134] Referring now to FIG. 9, FIG. 9 is a flowchart showing an exemplary 3D tracking process for tracking the 3D movement of a camera according to some embodiments of the present invention.
[0135] At 1102, the matched features between frame n and the next analyzed consecutive frame n + i are provided as the output of the process execution 510 described with reference to FIG. 5, for example.
[0136] At 1104, the calculated 3D coordinate values of the matched features (keypoints) between frames n and n + i are optionally provided from the 3D reconstruction process as described herein (e.g., output by a 3D reconstruction neural network).
[0137] At 1106, the 3D holography is calculated according to the 3D coordinate values (1104) of the matching features (1102).
[0138] At 1108, the closest 3D rigid body transformation matrix is found.
[0139] At 1110, the affine 3D transformation matrix is calculated.
[0140] At 1112, a 3D rigid body transformation matrix that depicts the movement of the camera from frame n to frame n + i is calculated.
[0141] Referring now to FIG. 10, FIG. 10 is an example of a 3D rigid body transformation matrix 1202 for tracking the 3D movement of a colon endoscope camera according to some embodiments of the present invention.
[0142] Referring back to FIG. 1. The 3D tracking algorithm makes it possible to construct a 3D trajectory of the movement of the camera (e.g., during a colonoscopy procedure), and is calculated based on the 3D tracking of the position of the camera. The 3D trajectory can be defined according to a 3D coordinate system. For example, the origin correlates with the camera position when tracking is started (e.g., when the endoscope just enters the colon, e.g., the point [0,0,0] in the 3D coordinate system where the 3D values of the first frame in the colon are reconstructed). The trajectory can be constructed by calculating the 3D position (e.g., its principal point) of the camera for each new frame (and / or multiple new frames, e.g., 1 - 25 new frames) relative to the 3D position of the camera at the last frame where the 3D position was calculated. The calculation of the new 3D position can be derived from the 3D rigid body transformation matrix between two frames (e.g., multiplying the previously calculated 3D position of the camera by the newly calculated transformation matrix), (e.g., immediately, without significant delay, during the time until the current frame is presented on the display).
[0143] Optionally, the 3D movement of the camera is tracked relative to the 3D position(s) of anatomical landmarks and / or the 3D position(s) of previously detected polyps. Anatomical landmarks may be predefined and / or set by an operator, and may be, for example, the position of the anus, the position of the cecum, anatomical abnormalities, and / or parts of the colon (e.g., transverse, ascending). Polyps are automatically and / or manually detected during forward movement of the camera and removed during backward movement of the camera.
[0144] At 112, the portions of the inner surface of the colon depicted in the image are calculated. For example, each portion may be defined as one-third, one-fourth, one-eighth (or other divisions) of the circumference of the inner surface of the colon. The coverage of each portion may be calculated dynamically in real time, for example, regardless of whether the currently captured image depicts (e.g., mostly depicts) each portion. Coverage of the entire colon (or a portion thereof) can be calculated as a set of coverages of individual portions, providing results such as, for example, about 94% of the inner surface of the colon is covered and / or about 6% of the inner surface of the colon is not covered.
[0145] Optionally, the presentation of the total amount of the inner surface depicted in the image relative to the amount of the non-depicted inner surface is calculated by aggregating the portions covered during the spiral scanning movement with respect to the portions not covered during the spiral scanning movement.
[0146] The portions of the inner surface of the colon depicted within the endoscopic image may be calculated based on, for example, analysis of the 3D reconstruction of the image and / or based on analysis of the image itself, for example, according to the identification of the position of the lumen within the image. The lumen may be identified, for example, as a region of pixels having intensity values below a threshold indicating darkness (i.e., the lumen does not reflect the light source to the camera).
[0147] Optionally, multiple consecutive images are aggregated to generate a single mosaic image. Portions of the inner surface of the colon may be calculated for the single mosaic image and / or the individual images. The multiple consecutive images may be taken at different orientations of the camera, for example, when the camera is oriented using clockwise, counterclockwise, x-pattern, or other movements. The camera may remain stationary at the same position along the longitudinal axis of the colon, or may be moved, for example, while being oriented in a spiral pattern, such as when the camera is pushed forward and / or pulled back.
[0148] The cumulative portions of the inner surface of the colon depicted in consecutive endoscopic images may be tracked for each position. Portions of the inner surface of the colon may be calculated for each position within the colon for different orientations of the camera without displacing the camera forward or backward (e.g., the displacement is zero or less than a predetermined threshold). When the camera is moved, the portions of the inner surface of the colon depicted in the captured images may be recalculated for each new position, for example, as the colonoscope is withdrawn from the colon. For example, Q1, Q2, Q3, Q4 are a set for the current location, and another set of Q1, Q2, Q3, Q4 is a set for the new location. Alternatively, the cumulative portions of the inner surface of the colon depicted in consecutive endoscopic images are tracked during the spiral movement of the camera. The portions of the inner surface of the colon depicted in the captured images may be recalculated in a spiral pattern as the camera is moved, for example, the quarters are repeated individually for the spiral movement, for example, Q1, Q2, Q3, Q4, Q1, Q2, Q3, Q4, Q1, Q2, Q3, Q4.
[0149] Optionally, each portion corresponds to a time window having intervals corresponding to the amount of time to cover the entire inner circumference during the helical scan operation, i.e., after a single helical scan, it returns to the same arc position, i.e., completes approximately 360 degrees, or returns to an arc range that defines the same quarter, i.e., after Q1 is completed, Q2, Q3, and Q4 are completed and it returns to Q1. The display of the coverage of each portion may be updated as a correlation to the time window. For example, if all portions are continuously associated with the display of coverage (e.g., if all portions are colored red or another color for proper coverage), the operator has covered all portions appropriately. Each portion may be associated with the display of proper coverage during the time interval corresponding to the time window, i.e., the portion needs to be appropriately covered up to the current circumferential helix end and appropriately depicted in the new circumferential helix. The operator can use the display that all portions are appropriately covered as a guideline during the continuous helical scan in which the image of the colon is appropriately captured. When one of the displays of one of the portions changes to an inappropriate coverage display, the operator can capture each portion within the image.
[0150] The time window can be, for example, based on clinical guidelines and / or based on a doctor's diagnosis, such as about 2.5 seconds per quarter, or about 5 seconds per quarter, or other values per section, or about 10 seconds or about 20 seconds per circumferential spiral scan, or other values. The time window may be dynamically calculated and adjusted based on real-time measurement of the spiral scan movement being performed by the operator. For example, when the operator stops the spiral scan to focus on a polyp, the time window is stopped and resumed when the operator resumes the spiral scan. When the operator slows down the spiral scan speed, such as by having the same operator or a student (e.g., a trainee doctor) perform the scan, the time window increases accordingly. The real-time speed of the spiral scan movement may be measured, for example, by analysis of the captured images (e.g., the tracking distance between the features extracted by matching from the perspective of the frame capture rate), and / or by one or more sensors that sense the movement of the femoral endoscope. An appropriate coverage display is generated when one or more images depicting most of each part are captured during the time window, and / or another display of inappropriate coverage is generated when images depicting most of each part are not captured during the time window.
[0151] An appropriate coverage display may be generated, for example, when at least 50%, or 60%, or 70%, or 80%, or 90%, or other intermediate values or larger values of each part are depicted in each image. The threshold for determining the amount of the required part in each image may be set, for example, based on the lens and the area of the inner surface of the colon depicted in the image.
[0152] Optionally, the 3D values of the pixels within the tracked frames are integrated into one 3D coordinate system. A 3D panoramic (mosaic) image of the colon surface can be generated based on a single 3D coordinate system, for example, based on the process described in Morimoto, Carlos, and Rama Chellappa et al., "Fast 3D stabilization and mosaic construction", cvpr. IEEE, 1997. The mosaic image may be constructed continuously (e.g., when the endoscope is being withdrawn from the colon).
[0153] For each frame included in the mosaic image, the lumen region (e.g., the center of the colon tube) can be detected by identifying dark pixels as, for example, pixels having an intensity value lower than a threshold (e.g., 20 or lower than other values if the intensity range is 0 - 255), in which case the dark pixels have the 3D position that is farthest from the camera position when the frame was captured (or the 3D positions correlated to the pixels within their proximity environment).
[0154] Optionally, the positions surrounding the inner surface of the colon covered by the current frame are calculated. Optionally, the inner surface of the colon is divided, for example, into four quarters. The quarter of the inner surface of the colon shown by the current frame can be calculated, for example, as the first, second, third, or fourth quarter of the colon. The captured images may be aggregated into the mosaic image to incrementally cover the quarters, in some cases until the mosaic image depicts all four quarters, which indicates that the entire circumference of the inner surface of the colon at the current position is depicted in the image.
[0155] The quarter(s) of the inner surface of the colon depicted in the individual images and / or mosaic image(s) may be calculated, for example, based on the detected lumen area. For each new frame (e.g., a frame presented in the mosaic image and / or an individual frame added), the 3D position of the central pixel and / or the direction of the 3D position of the central pixel relative to the detected 3D position of the lumen is calculated. The 3D position of the lumen may be estimated from the 3D positions of the pixels in its near environment. The quarter covered by the current frame can be calculated.
[0156] The estimated quarter of the colon depicted by the currently presented frame and / or mosaic image may be shown to the user. When all four quarters are covered, a display may be output indicating that the camera may be moved to a new position, and / or when one or more quarters are not covered and the camera is moved, another display may be generated indicating the lack of proper imaging of the local colon area. Using the display, the physician performing the colonoscopy procedure can determine whether the four quarters of the current local colon area are sufficiently covered in consecutive frames. The quarters covered by the frame may be shown to the user (e.g., for at least 5 seconds), such that the user can view the combination of the four shown quarters together on the screen as an indication that the inner surface of the colon is sufficiently covered during the scanning of the colon. The covered quarters may be indicated, for example, by their number appearing and / or blinking on the screen.
[0157] Optionally, while the endoscope is advancing and / or retracting (e.g., while being pulled out of the cecum, while exiting the colon and until the endoscope is completely removed from the colon), the covered quarter sequences are dynamically aggregated and / or tagged at their respective 3D positions. The calculated trajectory of the endoscope (e.g., the endoscope is advancing and / or retracting) may be scanned with a window of a predetermined length (e.g., about 2-3 cm) and / or a predetermined stride length (e.g., about 0.5-1 cm). In the number of consecutive scanning windows (denoted by Ncs and having, for example, a value of 3), if a quarter is not covered at all, it is registered as a missed quarter. The percentage of the colon covered by the endoscope camera during advancement and / or while being pulled out is a mathematical relationship of [Number] It can be calculated using, where Tms represents the total number of missed quarters and Nsw represents the total number of scanned windows. For example, if the total number of missed quarters (Tms) is 8, the total number of scanned windows (Nsw) is 100, and the number of consecutive scanned windows (Ncs) is 3, the percentage of the colon covered according to the mathematical relationship is 76%. The calculation of the percentage of the colon covered by the endoscopic camera can be performed dynamically in real time, for example, based on a set of images captured during the procedure, and / or offline after the procedure is completed using the set of images captured during the procedure. Referring now to FIG. 11, FIG. 11 is a schematic diagram of a mosaic and / or panoramic image of the colon 1302 being constructed for each pixel in the combined 2D image, where the 3D position is calculated within a single 3D coordinate system according to some embodiments of the present invention. As shown, three 2D frames 1304 having a step of three frames between each consecutive frame are used to generate the panoramic image 1302. The lumen region 1306 is shown as the dark region at the center of the colon tube. The panoramic image 1302 covers only a portion of the local colon region. Region 1308 remains unfilled by the mosaic image since the image is captured from the mosaic image.
[0158] Referring now to FIG. 12, FIG. 12 is a schematic diagram showing individual images 1402 displayed within each quarter of the inner surface of the colon shown herein according to some embodiments of the present invention. Portion 1404A shows quarter 1, portion 1404B shows quarter 2, portion 1404C shows quarter 3, and portion 1404D shows quarter 4.
[0159] Next, referring to FIG. 13, FIG. 13 is a flowchart of a method for calculating the quarter indicated by the frame according to some embodiments of the present invention. Note that the selected quarter indicates the quarter most covered by the frame since the frame may overlap between two or more quarters.
[0160] At 1502, the current frame (indicated by #n) is provided.
[0161] At 1504, the frame #n is registered in the 3D mosaic image currently being constructed.
[0162] At 1506, a lumen can be detected within the current frame. The 3D position of the lumen can be estimated.
[0163] At 1508, if no lumen (e.g., the dark area at the center of the colon tube) has been detected yet in the last frame of the current mosaic (e.g., in the last 10 seconds of the video), R02 - R08 is repeated.
[0164] At 1510, when a lumen is detected, the 3D position of the central pixel of the image is calculated. If the central pixel is determined to be within the lumen area, the 3D position is calculated according to the pixels within the proximity environment.
[0165] At 1512, the relative direction (e.g., 3D vector direction) between the 3D position of the central pixel of the current frame and the estimated 3D position of the last detected lumen area is calculated.
[0166] At 1514, based on the assumption that the vector whose direction was calculated at 1512 starts from the center of the colon tube (i.e., the center of the lumen), a quarter of the colon that covers most of the area of the current frame is determined according to the direction of the vector. (If two or more quarters are equally covered and the uncovered quarter is not displayed,) the selected quarter is the quarter that is mostly covered by the current frame.
[0167] At 1516, the quarter shown in the current frame is provided. As described herein, instructions for presenting the display of the quarter determined in the GUI may be generated.
[0168] Referring back to FIG. 1, at 114, the dimensions of the detected polyp are calculated. The dimensions may be 2D and / or 3D dimensions, such as the volume of the polyp, the radius of a sphere representing the polyp, the surface area of the polyp (e.g., a flat polyp), and / or the radius of a 2D circle representing the polyp. The dimensions may be calculated by performing a best fit of a 3D sphere and / or 2D circle to the pixels of a 3D image (created from a 2D image as described herein and optionally by a 3D CNN) representing the polyp. The dimensions are obtained from the best fit 3D sphere and / or 2D circle, e.g., the radius of the best fit 3D sphere and / or the radius of the best fit 2D circle.
[0169] The dimensions of the polyp are calculated based on the ROI depicting the calculated polyp and / or the 3D coordinates of the pixels of the 3D image output by the 3D reconstruction neural network as described with reference to 108 and / or 110 (i.e., 3D reconstruction of a 2D image). The ROI depicting the polyp may be manually set by an operator (e.g., using a GUI) and / or automatically output by the detection network described with reference to 104 when one or more 2D images are supplied.
[0170] The process described herein enables the calculation of the 3D volume and / or 2D surface area dimensions of the polyp using (optionally only using) the 3D coordinate values (x, y, z) calculated for the pixels in a 2D image (e.g., output by a 3D CNN) as described herein. In contrast, other processes that use 2D images to calculate 3D and / or 2D dimensions require knowing camera characteristics (e.g., the pose of the camera with respect to the polyp) that can be difficult to obtain.
[0171] The size of the polyp can be calculated based on the radius of a circle that best fits the polyp slice.
[0172] The volume of the polyp can be automatically calculated by taking the calculated 3D values of the 2D pixels inside the contour and / or bounding box that depicts the polyp and finding the best-fit 3D sphere (or circle if the polyp is flat) of the exposed 3D surface created by interpolating between the 3D values of the pixels. The radius of the sphere (or circle) is the polyp size.
[0173] Note that the 3D volume can be calculated from the calculated radius of the sphere correlated with the polyp, as described herein.
[0174] The estimated dimensions (e.g., 3D volume) of the polyp(s) may be calculated, for example, by the following process. Calculate the best-fit 3D surface for the 3D coordinates of the pixels in the region of at least one 2D image. Calculate a plurality of normal vectors proximate to the 3D position correlated with the center of gravity of the region (i.e., ROI) that depicts the polyp. Determine the side of the 3D surface that is concave according to the relative direction of the normal vectors. Calculate the plane containing the normal vectors of the plurality of normal vectors correlated with the 3D position correlated with the center of gravity. Calculate the vector correlated with the center of gravity that touches the best-fit 3D surface at the 3D position. Calculate the first curvature as the radius of the tangent parabola of the contour that intersects between the best-fit 3D surface and the calculated plane. Calculate the second curvature as the radius of the tangent parabola of the contour that intersects between the best-fit 3D surface and the orthogonal plane orthogonal to the calculated plane, and the orthogonal plane contains the normal vector. Then, calculate the 3D radius of the 3D volume of the polyp as the average of the first curvature and the second curvature.
[0175] Referring now to FIG. 14, FIG. 14 includes a schematic diagram showing a process for calculating the volume of a polyp from a 3D reconstructed surface calculated from a 2D image, according to some embodiments of the present invention. Schematic 1602 shows a 2D colonoscopy frame having a detected polyp 1604 obtained from the web address site.google.com / site / suryaiit / research / endoscopy / sfs. Schematic 1606 shows the 3D reconstructed surface of frame 1602. Polyp 1608 is the 3D reconstruction of polyp 1604 of 2D image 1602. Image 1610 shows a process for calculating the volume of the polyp by finding the best fit (tangent) 3D sphere 1612 on the concave side of the surface interpolated from the 3D values of the pixels within the polyp region, as described herein. The volume is the radius (R) of 3D sphere 1612.
[0176] Next, referring to FIG. 15, FIG. 15 is a flowchart of an exemplary process for calculating the volume of a polyp from a 2D image, according to some embodiments of the present invention.
[0177] At 1702, a 2D frame, denoted as #n, having a detected polyp is received. A plurality of frames before and after the current frame are received.
[0178] At 1704, a 3D reconstruction of frame #n is calculated, as described herein.
[0179] At 1706, the pixels within the boundary contour and / or boundary box of the polyp (as available) are identified and the best fit surface of their 3D values is calculated.
[0180] At 1708, a plurality of normal vectors in the vicinity of a 3D point correlated with the centroid of the boundary box of the polyp are calculated (if there is a delineated contour and the centroid is outside the delineated contour, the centroid is replaced with the nearest pixel within the delineated contour).
[0181] At 1710, the concave side of the surface is identified according to the calculated normal vectors and their relative directions.
[0182] At 1712, an infinite plane is calculated that includes the normal vector correlated with the upper centroid (or its replacement) and the vector at this 3D point (correlated with the centroid) that is tangent to the surface.
[0183] At 1714, the curvature (radius of the tangent parabola) of the contour that is the intersection between the polyp surface and the calculated plane is calculated based on, for example, the approach described in "Curvature of curves and surfaces - a parabolic approach" by Har’el, Zvi et al., Department of Mathematics, Technion - Israel Institute of Technology (1995).
[0184] At 1716, (to obtain an additional estimate of the curvature) the last step is repeated for a plane orthogonal to the previous plane and including the normal vectors from 1712 and 1714.
[0185] At 1718, the average of the last two calculated curvatures is calculated, providing the 3D radius of the polyp.
[0186] At 1720, if the 3D radius is greater than a predetermined threshold (e.g., 7 millimeters, 7 centimeters, or other value), at 1722, the polyp size is shown to be flat. The best - fit 2D circle is calculated. The size of the polyp is defined according to the size of the 2D circle. At 1726, if the 3D radius is less than the predetermined threshold, the size of the polyp is defined according to the 3D radius.
[0187] Return to FIG. 1 for reference. At 116, instructions are generated to update the presentation of the GUI. The image may be augmented. The instructions may be for introducing GUI elements into the image, for example, to present an overlay on the image and / or to present the image within a portion of the GUI and other graphical elements within another portion of the GUI, such as a covered quadrant and / or within a colon map.
[0188] The instructions are generated based on the output of one or more of the features described with reference to 104 - 114. * For example, instructions are generated to augment the image at the location of the detected polyp as described with reference to 104 by marking an ROI (e.g., a bounding box) depicting the detected polyp and / or color - coding the polyp and / or the bounding box and / or referring to an arrow indicating the polyp. * As described with reference to feature 106, instructions are generated based on the display of a vector pointing from the current image towards the ROI, and a polyp located outside the current image is depicted. The instructions are for generating an augmented endoscopy image by augmenting each endoscopic image with the display of the vector.
[0189] The vector may be presented as an arrow. The arrow indicates the direction in which the camera should be moved to reacquire the polyp within the image. The arrow and the image may be presented in 2D.
[0190] Optionally, when the polyp is reacquired within the image (e.g., after moving the camera in the direction of the arrow), instructions are generated to augment the image with the ROI depicting the polyp. The ROI may be remarked on the image based only on tracking, without necessarily performing feature 104. Optionally, instructions are generated to present a colon map in the colon based on the output of the features described with reference to 106 and / or 110. The 2D and / or 3D positions of the polyps (e.g., the ROI depicting the polyps) are output, for example, by a detection neural network and / or a 3D reconstruction neural network, and / or manually marked by an operator (e.g., using a GUI by pressing an "polyp" icon). The positions of the polyps are marked on a schematic of the patient's colon, for example, in a display such as an X and / or a circle, and optionally color-coded according to whether the polyps were identified during insertion and / or removal of the colonoscope. The colon map is dynamically updated with newly detected polyps. Optionally, instructions are generated to mark polyps presented on the colon map with an indication of treatment (e.g., surgical removal, excision). For example, a polyp presented as an ellipse in a removed colon map is marked with an X. An ellipse in a map that has not been removed is not marked with an X. The polyps are marked as treated based on an indication of removal of the polyps, which are provided manually by the user (e.g., by pressing a "polyp removal" icon on the GUI), and / or automatically detected by code (e.g., based on detection of movement of a surgical tool within the colonoscope). Optionally, instructions are generated to plot the tracked 3D position of the endoscope on a colon map presented within the GUI, e.g., as a trajectory and / or curve and / or dotted line. The trajectory and / or curve and / or dotted line are dynamically updated as the endoscope moves (e.g., advances and / or retreats) within the colon. Optionally, the forward direction tracked 3D position is marked on the colon map with a marking, e.g., an arrow in the forward direction, and / or color coding, that indicates the forward direction of the endoscope camera as it enters deeper into the colon (e.g., from the rectum to the cecum). The reverse direction tracked 3D position presented on the colon map is marked with a different marking, e.g., an arrow in the reverse direction and / or a different color, that indicates the reverse direction of the endoscope camera as it is removed from the colon (e.g., from the cecum to the rectum).
[0191] Optionally, instructions are generated to present the calculated dimensions of the detected polyp. For example, the calculated dimensions are presented as a numerical value that optionally has a value of the calculated volume in units, e.g., cubic millimeters, proximate to the ROI marked on the image. In another example, the calculated dimensions are presented as a numerical value having units proximate to the display of the polyp on the colon map. Optionally, instructions to present a warning within the GUI indicating a proposal to remove a polyp are generated when the dimensions exceed a threshold. For example, when the radius of the sphere defining the 3D volume of the polyp exceeds 2 millimeters. The warning may be generated, e.g., as a coloring of the boundary of the polyp and / or ROI that depicts the polyp on the image, and / or as a coloring of the polyp representation on the colon map, using a distinct color (e.g., red for removal and green for leaving) that indicates the proposal to remove, and / or as an audio message within the GUI, and / or as a pop-up text message. Optionally, instructions are generated to present an estimate of the remaining portion of the inner surface area that has not yet been depicted in a previously captured endoscopic image. Alternatively or additionally, instructions are generated to present an estimate of the portion of the inner surface area that has already been depicted. For example, a circular GUI element divided into four quadrants (or, for example, another number of divisions such as slices) is presented. The quadrants depicted in the image are indicated by one marking, for example, in green. The quadrants that have not yet been depicted are indicated by a different marking, for example, in red. The operator can view the marked ones (e.g., the color-coded circle) and direct the camera in the direction towards the quadrants that have not yet been imaged. When the GUI element indicates that all quadrants have been imaged, the operator can displace the colonoscope (e.g., forward or backward).
[0192] The quadrant can depict a mostly imaged area where, for example, more than 50%, more than 70%, or more than 80% of the surface of the quadrant is depicted in one or more images. Other divisions can be selected according to the imaging ability of the camera lens. For example, a larger number of divisions can be used for a narrow-angle lens. Optionally, instructions to present a warning are generated according to a predefined system definition (which may sometimes be set by a setup menu) when the current position of the camera is close to an anatomical landmark and / or a polyp. For example, when the position of the camera is close to the position of a landmark, and when the distance is before a predetermined threshold, the definition can be set to generate a warning. The distance between the camera and the landmark and / or the polyp can be calculated, for example, by the L2 metric and / or a neighborhood that is small enough to give an alarm (e.g., about 3 centimeters). The threshold distance value may take into account the aggregated calculation error when constructing the trajectory and / or the fact that the colon itself is not completely stable in the abdomen, i.e., not stable in the coordinate system in which the trajectory is constructed.
[0193] Optionally, the amount of time spent by the endoscope at each defined portion of the colon can be calculated and presented. Each portion may be defined according to, for example, transition points between anatomical landmarks that divide the colon into portions, such as the ascending colon and the transverse colon (e.g., between the transverse colon and the ascending colon). The anatomical landmarks may be detected manually by the user (e.g., the user marks the landmarks using a GUI) and / or automatically by code (e.g., based on analysis of an image), and the 3D position of the endoscope camera with respect to the anatomical landmarks is tracked. The amount of time spent by the endoscope camera at each portion of the colon is calculated. Instructions are generated for presenting the amount of time spent by the endoscope camera at each portion of the colon within the GUI. The time may be presented, for example, as markings on respective portions of a colon map corresponding to portions of the patient's colon.
[0194] At 118, the generated instructions are executed. The GUI is updated according to the instructions.
[0195] Next, referring to FIG. 16, FIG. 16 is a schematic diagram showing a sequence of magnified images for tracking the ROI of a polyp presented within a GUI according to some embodiments of the present invention, in accordance with the generated instructions. This sequence shows the 2D tracking of the polyp and the generation of an arrow indicating the direction of the polyp when the polyp is located outside the image, as described herein. Image 1902 is a first image augmented with ROI 1904 depicting the automatically detected polyp. In the second image of sequence 1906, ROI 1904 is moving towards the left of the image for reorientation of the camera. In the third image of sequence 1908, ROI 1904 is moving further to the left of the image. In the fourth image of sequence 1910, the ROI is located outside the image and thus not shown in image 1910. Arrow 1912 is shown to point in the direction of the ROI located outside. In the fifth image of sequence 1914, ROI 1904 is reproduced after the operator has reoriented the camera in the direction indicated by arrow 1912 in image 1910. In the sixth image of sequence 1916, ROI 1904 is moving further to the right due to movement of the camera by the operator.
[0196] Referring now to FIG. 17, FIG. 17 is a schematic diagram of a colon map 2002 presenting the movement trajectory of an endoscope, the positions of detected polyps, removed polyps, and anatomical landmarks presented within a GUI according to some embodiments of the present invention. The colon map 2002 may be presented, for example, as an overlap presented at the corner of an image or within a designated window of a GUI having another designated window for presenting the image. The trajectory 2004 shows the tracked 3D movement of the endoscope during forward movement into the colon, color-coded for example. The trajectory 2006 shows the tracked 3D movement of the endoscope during reverse movement out of the colon, color-coded with a different color for example. The ellipse 2008 indicates a polyp automatically detected during forward movement of the colonoscope and is colored, for example, with the same color used for the forward trajectory 2004. The ellipse 2010 indicates a polyp automatically detected during reverse movement of the colonoscope and is colored, for example, with the same color used for the reverse trajectory 2006. The circle 2012 indicates polyps manually found by the user and / or automatically detected polyps manually marked by the user. The X marking 2014 indicates polyps removed by the operator. The arrowhead 2016 indicates the current 3D position of the endoscope, optionally the distal end (e.g., tip) of the camera. The box 2018 indicates manually specified landmarks, such as hemorrhoids, manually entered by the user (e.g., via the GUI). The circle 2020 indicates automatically identified landmarks detected, for example, by analysis of 3D reconstructed images. The colon map 2002 is dynamically updated when the colonoscope is moved when a polyp is detected and / or when a polyp is removed.
[0197] Referring to FIG. 18, FIG. 18 is a schematic 2102 showing the quadrants of the inner surface of the colon depicted in one or more images 2104A-C and / or quadrants of the inner surface of the colon not yet depicted in image 2106, according to some embodiments of the present invention. Schematic diagram 2102 is created for each new position of the camera as the camera is displaced forward and / or backward. Alternatively or additionally, schematic diagram 2102 is calculated for a recently defined time window (e.g., 5 seconds, or 10 seconds, or other value). The time window may be short enough to exclude movement of the camera forward and / or backward, but long enough to allow for complete imaging around the inner wall of the colon. The quadrants are updated when a new image is acquired at the same location by changing the orientation of the camera. Quadrants 2104A-C and 2106 may be color-coded. Schematic diagram 2102 is dynamically updated as the operator moves the camera to image the remaining portions of the inner surface of the colon.
[0198] At 120, one or more of the features described with reference to 100-118 are repeated. The repetition can, for example, dynamically track a polyp, dynamically generate an arrow indicating an ROI depicting a polyp located outside the current image, update the colon map at the new 2D and / or 3D position of the polyp, update the camera trajectory of the colon map with the movement of the camera, update the GUI to depict the coverage of the inner surface of the colon (e.g., quadrants), and / or dynamically update the GUI to update the calculated 2D and / or 3D dimensions of the detected polyp.
[0199] The updated GUI may be used by the operator, for example, to operate the camera to capture an image of the polyp, to determine which polyp(s) should be removed, to operate the camera to ensure complete coverage of the inner surface of the colon, and / or to track the position of the camera and / or detected polyp(s) within the colon.
[0200] The descriptions of the various embodiments of the present invention are presented for illustrative purposes, but are not intended to be exhaustive and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terms used herein are selected to best explain the principles of the embodiments, the practical application to technologies found in the market, or the technological improvements, or to enable other practitioners in the art to understand the embodiments disclosed herein.
[0201] During the term of the patent commencing from this application, it is expected that many related anatomical images will be developed, and the scope of the term "anatomical image" is intended to include a priori all such new technologies.
[0202] As used herein, the term "about" refers to ±10%.
[0203] The terms "comprises," "comprising," "includes," "including," "having," and their conjugates mean "including but not limited to." This term encompasses the terms "consisting of" and "consisting essentially of."
[0204] The expression "consisting essentially of" means that a composition or method may include additional ingredients and / or steps, but only if the additional ingredients and / or steps do not substantially alter the basic and novel characteristics of the composition or method described in the claims.
[0205] As used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "compound" or "at least one compound" can include a plurality of compounds including mixtures thereof.
[0206] In this specification, the term "exemplary" is used in the sense of "provided as an example, instance, or illustration". Any embodiment described as "exemplary" need not be construed as preferred or advantageous over other embodiments, and / or need not be construed as excluding the incorporation of features from other embodiments.
[0207] The phrase "optionally" is used in this specification in the sense of "provided in some embodiments and not provided in other embodiments". Any particular embodiment of the invention may include a plurality of "optional" features as long as such features are not inconsistent.
[0208] Throughout this application, various embodiments of the invention may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, a description of a range should be considered to specifically disclose not only the individual numerical values within that range but also all possible sub-ranges. For example, a range description such as from 1 to 6 should be considered to specifically disclose sub-ranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as the individual numerical values within that range, e.g., 1, 2, 3, 4, 5, 6, etc. This applies regardless of the breadth of the range.
[0209] Whenever a numerical range is indicated in this specification, it is meant to include any numerical value (fractional or integral) cited within the indicated range. In this specification, the expressions "range between / over the first reference number and the second reference number" and "range from the first reference number to the second reference number" are used interchangeably and are meant to include the first and second reference numbers, as well as all fractional and integral numbers therebetween.
[0210] For clarity, it is understood that certain features of the invention, which are described in the context of separate embodiments, may be provided in combination in a single embodiment. Conversely, for brevity, the various features of the invention that are described in the context of a single embodiment may be provided separately, or in any suitable sub - combination, or as appropriate in other described embodiments of the invention. Specific features described in the context of various embodiments are not considered essential features of those embodiments, unless the embodiments would not operate without those elements.
[0211] Although the invention has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations will be apparent to those skilled in the art. Accordingly, the invention is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and broad scope of the appended claims.
[0212] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Further, the citation or identification of a reference in this application shall not be construed as an admission that such reference is available as prior art to the present invention. To the extent that section headings are used, they should not necessarily be construed as limiting. Also, the priority documents of this application are hereby incorporated by reference in their entirety.
Claims
**Claim 1** A system for generating instructions for presenting a graphical user interface (GUI) for dynamically tracking at least one polyp in a plurality of endoscopic images of a patient's colon, tracking the position of an area depicting at least one polyp in each endoscopic image relative to at least one previous endoscopic image, wherein each of the endoscopic images represents at least one current endoscopic image, the tracking, in the case where the position of the area is outside the at least one current endoscopic image, and the area depicting the at least one polyp has disappeared from the at least one current endoscopic image and is no longer depicted in the at least one current endoscopic image, calculating a vector from within the at least one current endoscopic image to the position of the area that is no longer present within the at least one current endoscopic image and is located outside the at least one current endoscopic image, creating an augmented endoscopic image by augmenting each of the endoscopic images with the display of the vector, generating instructions for displaying the augmented endoscopic image within the GUI, A system comprising at least one processor that executes code for iterating over the plurality of endoscopic images. **Claim 2** The system of claim 1, wherein the display of the vector depicts a direction and / or orientation for adjustment of an endoscopic camera for capturing another at least one endoscopic image depicting the area of the at least one endoscopic image. **Claim 3** The system of claim 1, wherein when the position of the area depicting the at least one polyp appears in each endoscopic image, the augmented endoscopic image is generated by the at least one processor executing code for augmenting each endoscopic image at the position of the area, and the display of the vector is excluded from the augmented endoscopic image. **Claim 4** calculating the position of an area depicting at least one polyp in a patient's colon, creating a colon map by marking a schematic diagram representing the patient's colon using a display indicating the position of the area depicting the at least one polyp. Generating an instruction to present the colon map within the GUI, the colon map being dynamically updated at the location of the newly detected polyp, the system according to claim 1 further comprising code for generating the instruction.
5. Translating and / or rotating at least one endoscopic image among successive subsets of processed endoscopic images, each endoscopic image being included in the successive subsets of endoscopic images, such that the region depicting at least one polyp is at the same approximate position in all of the images of the successive subsets of endoscopic images, and the translating and / or rotating; Feeding the processed successive subsets of the endoscopic images to a detection neural network; Outputting, by the detection neural network, a current region depicting the at least one polyp for each of the endoscopic images; Creating an augmented image of each of the endoscopic images by augmenting each of the endoscopic images with the current region; The system according to claim 1 further comprising code for generating an instruction to present the augmented image within the GUI.
6. When the output of the detection neural network is provided for a previous endoscopic image that is serially earlier than each of the endoscopic images, and the tracked position of the region depicting at least one polyp within each of the endoscopic images is at a different position from the region output by the detection neural network for the previous endoscopic image, the at least one processor executes code for generating the augmented image for each of the endoscopic images based on the tracked position. The system according to claim 5.
Citation Information
Patent Citations
System and method for combined display in medicine
JP2009022446A
Medical imaging apparatus and method using comparison image
US20160171158A1
System and Method for Guiding and Tracking a Region of Interest Using an Endoscope
US20170258295A1
Methods for polyp detection
US20190080454A1
Method and system for navigating within a colon
US8795157B1