Systems and methods for contextualized image analysis
By using context-aware neural networks for medical imaging, the systems address inefficiencies in existing imaging systems by accurately detecting and classifying objects based on user interactions, improving accuracy and reducing unnecessary information display.
Patent Information
- Application Number
- JP2022529511
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-03
- Filing Date
- 2021-01-29
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2041-01-29
AI Technical Summary
Existing medical imaging systems lack the ability to accurately detect objects and provide classification while considering real-time context and user interactions, leading to inefficiencies and false positives.
Implement computer-implemented systems and methods that utilize contextual information to perform image processing operations, such as object detection and classification, by applying trained neural networks to process image frames from medical imaging devices, and modify visualization based on user interactions.
Enhances the accuracy and efficiency of medical imaging by enabling context-aware object detection and classification, reducing false positives and providing relevant information only when needed, thereby optimizing the user experience.
Smart Images

Figure 0007722993000002 
Figure 0007722993000003 
Figure 0007722993000004
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 62 / 969,643, provisionally filed February 3, 2020, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure relates generally to computer-implemented systems and methods for contextualized image analysis. More particularly, and without limitation, the present disclosure relates to computer-implemented systems and methods for processing real-time video and performing image processing operations based on context information. The systems and methods disclosed herein may be used in a variety of applications and vision systems, such as medical image analysis and systems where accurate image processing capabilities are an advantage. [Background technology]
[0003] In image analysis systems, it is often desirable to detect objects of interest in images. The objects of interest may be people, places, or things. In some applications, such as systems for medical image analysis and diagnosis, the location and classification of detected objects (e.g., anomalies on or in human tissue) are equally important. However, existing computer-implemented systems and methods suffer from many drawbacks, including an inability to accurately detect objects and / or an inability to provide the location or classification of detected objects. In addition, existing systems and methods are inefficient in that they may perform image processing operations unnecessarily and / or indiscriminately without considering the real-time context or use of the imaging device. As used herein, "real-time" means occurring or processing immediately.
[0004] Some existing medical imaging systems are built on a single detector network. When a detection is made, the network simply outputs the detection result, for example, to a physician or other medical professional. However, such detection results may be false positives, such as non-polyps or the like in an endoscopy. Such systems do not provide a separate network to distinguish between false positives and true positives.
[0005] Furthermore, neural network-based object detectors typically feed features identified by the neural network into a detector, which may include a second neural network. However, such networks are often inaccurate because feature detection is performed by a generalized network and only the detector portion is specialized.
[0006] Existing medical imaging systems for real-time applications have other disadvantages, such as the fact that they are often designed to operate without regard for the context of use or the real-time interaction between a physician or other user and the medical imaging device that generates the video frames for processing.
[0007] Furthermore, existing medical imaging systems for real-time applications do not use contextualized information derived from interactions between a physician or other user and the medical imaging device to aggregate objects identified by an object detector along the time dimension.
[0008] Furthermore, existing medical imaging systems for real-time applications do not use contextualized information derived from the interaction between a user and a medical imaging device to enable or disable specific neural networks that can perform specific tasks, such as detecting objects, classifying detected objects, outputting object characteristics, or modifying the way information is visualized on a medical display for the user's benefit.
[0009] In view of the foregoing, the inventors have determined that there is a need for improved systems and methods for image analysis, including those directed to medical image analysis and diagnosis. There is also a need for improved medical imaging systems that can accurately and efficiently detect objects and provide classification information. There is also a need for image analysis systems and methods that can perform real-time image processing operations based on contextual information. Summary of the Invention
[0010] In view of the foregoing, embodiments of the present disclosure provide computer-implemented systems and methods for processing real-time video from an imaging device, e.g., a medical imaging system. The disclosed systems and methods may be configured to perform image processing operations, such as object detection and classification. The disclosed systems and methods may also be configured to identify user interactions with the imaging device using contextual information and perform image processing based on the identified interactions, e.g., by applying one or more neural networks trained to process image frames received from the imaging device or to modify the manner in which information is visualized on a display based on the contextual information. The disclosed systems and methods provide advantages over existing systems and techniques, including by addressing one or more of the above-mentioned and / or other shortcomings of existing systems and techniques.
[0011] In some embodiments, the image frames received from the imaging device may include image frames of a human organ. For example, the human organ may include the digestive tract. The frames may include images from a medical imaging device used during at least one of an endoscopy, a gastroscopy, a colonoscopy, an enteroscopy, a laparoscopy, or a surgical endoscopy. In various embodiments, the object of interest included in the image frames may be a portion of a human organ, a surgical instrument, or an abnormality. The abnormality may include formations on or in human tissue, a change in human tissue from one type of cell to another type of cell, and / or the absence of human tissue where human tissue is expected to be present. The formations on or in human tissue may include lesions, such as polypoid lesions or non-polypoid lesions. As a result, the disclosed embodiments may be utilized in a medical context in a generally applicable manner, rather than being specific to any single disease.
[0012] In some embodiments, the context information may be used to determine which image processing operations to perform. For example, the image processing operations may include enabling or disabling specific neural networks, such as object detectors, image classifiers, or image similarity evaluators. Additionally, the image processing operations may include enabling or disabling specific neural networks adapted to provide information about the detected object, such as the type of object or specific characteristics of the object.
[0013] In some embodiments, the context information may be used to identify user interactions with the imaging device. For example, the context information may indicate that a user is interacting with the imaging device to identify an object of interest within an image frame. Thereafter, the context information may indicate that the user is no longer interacting with the imaging device to identify the object of interest. As a further example, the context information may indicate that a user is interacting with the imaging device to examine one or more detected objects within the image frame. Thereafter, the context information may indicate that the user is no longer interacting with the imaging device to examine one or more detected objects within the image frame. However, it will be understood that the context information may also be used to identify any other user interactions with the imaging device or associated equipment comprising the medical imaging system, such as showing or hiding displayed information, performing video functions (e.g., zooming to an area including an object of interest, image color distribution, or the like), saving captured image frames to storage, powering the imaging device on or off, or the like.
[0014] In some embodiments, contextual information may be used to determine whether to perform aggregation of objects of interest across multiple image frames along the time dimension. For example, it may be desirable to capture every image frame containing an object of interest, such as a polyp, for future examination by a physician. In such a situation, it may be advantageous to group every image frame captured by the imaging device that contains the object of interest. Information, such as a label, timestamp, location, or distance traveled, may be associated with each group of image frames to distinguish them from one another. Other methods of performing aggregation of objects of interest may also be used, such as changing the color distribution of the image frames (e.g., using green to indicate a first object of interest and red to indicate a second object of interest) or adding alphanumeric information or other characters to the image frames (e.g., using "1" to indicate a first object of interest and "2" to indicate a second object of interest).
[0015] The context information may be generated by various means consistent with disclosed embodiments. For example, the context information may be generated by using an intersection over union (IoU) value for the location of detected objects in two or more image frames over time. The IoU value may be compared to a threshold to determine the context of a user's interaction with the imaging device (e.g., a user navigating the imaging device and attempting to identify an object). In some embodiments, the IoU value meeting a threshold over a predetermined number of frames or time may establish the persistence required to determine user interaction with the imaging device.
[0016] In some embodiments, the context information may be generated by using image similarity values or other particular image features of detected objects in two or more image frames over time. The image similarity values or other particular image features of detected objects may be compared to a threshold to determine the context of a user's interaction with the imaging device (e.g., a user navigating the imaging device and attempting to identify an object). In some embodiments, the image similarity values or other particular image features of detected objects may meet a threshold over a predetermined number of frames or time to establish the persistence required to determine user interaction with the imaging device.
[0017] The disclosed embodiments may also be implemented to derive context information based on the presence or analysis of multiple objects simultaneously present within the same image frame. The disclosed embodiments may also be implemented to derive context information based on analysis of the entire image (i.e., not just identified objects). In some embodiments, the context information is derived based on classification information. Additionally or alternatively, the context information may be generated based on user input received by the imaging device indicating user interaction (e.g., input indicating that the user is examining an identified object by focusing or zooming the imaging device). In such embodiments, persistence of the user input over a predetermined number of frames or time may be required to determine user interaction with the imaging device.
[0018] Embodiments of the present disclosure include computer-implemented systems and methods that perform image processing based on context information. For example, in some embodiments, object detection may be invoked when context information indicates that a user is interacting with an imaging device to identify an object. As a result, object detection is unlikely to be performed, for example, when an object of interest is not present or the user is otherwise not ready to initiate a detection process or one or more classification processes. As a further example, in some embodiments, classification may be invoked when context information indicates that a user is interacting with an imaging device to examine a detected object. Thus, the risk of premature classification being performed, for example, before the object of interest is properly framed or before the user intends to know the classification information for the object of interest, is minimized.
[0019] Additionally, embodiments of the present disclosure include performing image processing operations by applying trained neural networks to process frames received from an imaging device, such as a medical imaging system. In this manner, the disclosed embodiments may be adapted for a variety of applications, such as real-time processing of medical images, in a non-disease specific manner.
[0020] Embodiments of the present disclosure also include systems and methods configured to display real-time video (such as endoscopy video or other medical images) along with object detection and classification information obtained from image processing. Embodiments of the present disclosure further include systems and methods configured to display real-time video (such as endoscopy video or other medical images) along with image modifications introduced to draw a physician's attention to features of interest within the image and / or provide information about the features or objects of interest (e.g., overlays including a boundary indicating the location of the object of interest within the image frame, classification information for the object of interest, a zoomed image of the object of interest or a particular region of interest within the image frame, and / or a modified image color distribution). Such information may be presented together on a single display device for viewing by a user (e.g., a physician or other medical professional). Furthermore, in some embodiments, such information may be displayed in response to corresponding image processing operations being invoked based on contextual information. Thus, as described herein, embodiments of the present disclosure provide such detection and classification information efficiently and when needed, thereby avoiding overcrowding the display with unnecessary information.
[0021] In one embodiment, a computer-implemented system for real-time video processing may include at least one memory configured to store instructions and at least one processing device configured to execute instructions. The at least one processing device may execute the instructions to receive real-time video generated by a medical imaging system, the real-time video including a plurality of image frames. While receiving the real-time video generated by the medical imaging system, the at least one processing device may be further configured to obtain context information to direct a user's interaction with the medical imaging system. The at least one processing device may be further configured to perform object detection to detect at least one object within the plurality of image frames. The at least one processing device may be further configured to perform classification to generate classification information for the at least one detected object within the plurality of image frames. The at least one processing device may be further configured to perform image modification to modify the received real-time video based on at least one of the object detection and the classification, and generate a representation of the real-time video with the image modification on a video display device. The at least one processing device may be further configured to invoke at least one of the object detection and the classification based on the context information.
[0022] In some embodiments, at least one of object detection and classification may be performed by applying at least one trained neural network to process frames received from the medical imaging system. In some embodiments, the at least one processing unit may be further configured to invoke object detection when the context information indicates that a user may be interacting with the medical imaging system and attempting to identify the object. In some embodiments, the at least one processing unit may be further configured to disable object detection when the context information indicates that a user may no longer be interacting with the medical imaging system and attempting to identify the object. In some embodiments, the at least one processing unit may be configured to invoke classification when the context information indicates that a user may be interacting with the medical imaging system and examining at least one object in the plurality of image frames. In some embodiments, the at least one processing unit may be further configured to disable classification when the context information indicates that a user may no longer be interacting with the medical imaging system and examining at least one object in the plurality of image frames. In some embodiments, the at least one processing unit may be further configured to enable object detection when the context information indicates a region in the plurality of image frames that includes the at least one object is likely to be of user interest, and to invoke classification when the context information indicates a region in the plurality of image frames that includes the at least one object is likely to be of user interest. In some embodiments, the at least one processing unit may be further configured to perform aggregation of two or more frames that include the at least one object, where the at least one processing unit may be further configured to invoke aggregation based on the context information.In some embodiments, the image modification includes at least one overlay including at least one boundary indicating the location of the at least one detected object, classification information for the at least one detected object, a zoomed image of the at least one detected object, or a modified image color distribution.
[0023] In some embodiments, the at least one processing device may be configured to generate the context information based on an intersection-over-union (IoU) value of the location of at least one detected object in two or more image frames over time. In some embodiments, the at least one processing device may be configured to generate the context information based on image similarity values in two or more image frames. In some embodiments, the at least one processing device may be configured to generate the context information based on detection or classification of one or more objects in the plurality of image frames. In some embodiments, the at least one processing device may be configured to generate the context information based on input received by the medical imaging system from a user. In some embodiments, the at least one processing device may be further configured to generate the context information based on the classification information. In some embodiments, the plurality of image frames may include image frames of the digestive tract. In some embodiments, the frames may include images from a medical imaging device used during at least one of endoscopy, gastroscopy, colonoscopy, enteroscopy, laparoscopy, or surgical endoscopy. In some embodiments, the at least one detected object may be an anomaly. The abnormality may be a formation on or in human tissue, a change in human tissue from one type of cell to another type of cell, the absence of human tissue where human tissue is expected to be present, or a lesion.
[0024] In a further embodiment, a method for real-time video processing is provided. The method includes receiving real-time video generated by a medical imaging system, the real-time video including a plurality of image frames. The method further includes providing at least one neural network, the at least one neural network being trained to process the image frames from the medical imaging system to obtain contextual information that directs a user's interaction with the medical imaging system. The method further includes identifying an interaction based on the contextual information and performing real-time processing on the plurality of image frames based on the identified interaction by applying the at least one trained neural network.
[0025] In some embodiments, performing real-time processing includes performing at least one of object detection to detect at least one object in the plurality of image frames, classification to generate classification information for the at least one detected object, and image correction to correct the received real-time video.
[0026] In some embodiments, object detection is invoked when the identified interaction is a user interacting with and navigating the medical imaging system and attempting to identify an object, hi some embodiments, object detection is disabled when the context information indicates that the user is no longer interacting with or navigating the medical imaging system and attempting to identify an object.
[0027] In some embodiments, classification is invoked when the identified interaction is a user interacting with the medical imaging system to examine at least one detected object in the plurality of image frames. In some embodiments, classification is disabled when the context information indicates that the user is no longer interacting with the medical imaging system to examine at least one detected object in the plurality of image frames.
[0028] In some embodiments, object detection is invoked when the context information indicates that the user is interested in an area in the plurality of image frames that includes at least one object, and classification is invoked when the context information indicates that the user is interested in at least one object.
[0029] In some embodiments, object detection and / or classification is performed by applying at least one neural network trained to process frames received from a medical imaging system.
[0030] In some embodiments, the method further includes performing aggregation of two or more frames including the at least one object based on the context information. In some embodiments, the image modification includes at least one overlay including at least one boundary indicating a location of the at least one detected object, classification information for the at least one detected object, a zoomed image of the at least one detected object, or a modified image color distribution.
[0031] The plurality of image frames may include image frames of a human organ, for example, the digestive tract. By way of example, the frames may include images from a medical imaging device used during at least one of an endoscopy, a gastroscopy, a colonoscopy, an enteroscopy, a laparoscopy, or a surgical endoscopy.
[0032] According to embodiments of the present disclosure, at least one detected object is an abnormality, which may be a formation on or in human tissue, a change in human tissue from one type of cell to another, an absence of human tissue where human tissue is expected to be present, or a lesion.
[0033] Additional objects and advantages of the disclosure will be set forth in part in the detailed description which follows, and in part will be obvious from the description or may be learned by the practice of the disclosure. The objects and advantages of the disclosure will be realized and attained by means of the elements and combinations particularly pointed out in the appended claims.
[0034] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments.
[0035] The accompanying drawings, which form a part of this specification, illustrate several embodiments and, together with the description, serve to explain the principles and features of the disclosed embodiments. In these drawings: [Brief explanation of the drawings]
[0036] [Figure 1] FIG. 1 is a schematic diagram of an exemplary computer-implemented system for real-time video processing and overlaying information onto a video feed, according to an embodiment of the present disclosure.
[0037] [Figure 2A] FIG. 2A is a schematic diagram of an exemplary computer-implemented system for real-time image processing using context information, according to an embodiment of the present disclosure. [Figure 2B] FIG. 2B is a schematic diagram of an exemplary computer-implemented system for real-time image processing using context information, according to an embodiment of the present disclosure.
[0038] [Figure 3] FIG. 3 is a flowchart of an exemplary method for processing real-time video received from an imaging device, according to an embodiment of the present disclosure.
[0039] [Figure 4]FIG. 4 is a flowchart of an exemplary method for invoking image processing operations based on context information directing a user's interaction with an imaging device, according to an embodiment of the present disclosure.
[0040] [Figure 5] FIG. 5 is a flowchart of an exemplary method for generating overlay information on a real-time video feed from an imaging device, according to an embodiment of the present disclosure.
[0041] [Figure 6] FIG. 6 is an example display with an overlay of object detection and associated classification information within a video, according to an embodiment of the present disclosure.
[0042] [Figure 7A] FIG. 7A is an example visual representation of determining an intersection-over-union (IoU) value for detected objects in two image frames, according to an embodiment of the present disclosure.
[0043] [Figure 7B] FIG. 7B is another example visual representation of determining an intersection-over-union (IoU) value for detected objects in two image frames according to an embodiment of the present disclosure.
[0044] [Figure 8] FIG. 8 is a flowchart of another exemplary method for performing real-time image processing consistent with embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0045] Disclosed embodiments of the present disclosure generally relate to computer-implemented systems and methods for processing real-time video from an imaging device, e.g., a medical imaging system. In some embodiments, the systems and methods of the present disclosure may be configured to perform image processing operations, such as object detection and classification. As disclosed herein, the systems and methods may also be configured to identify user interactions with the imaging device using contextual information and to perform image processing based on the identified interactions. Furthermore, embodiments of the present disclosure may be implemented with artificial intelligence, e.g., one or more neural networks trained to process image frames received from the imaging device. These and other features of the present invention are further disclosed herein.
[0046] As will be understood from the present disclosure, the disclosed embodiments are provided for illustrative purposes and may be implemented and used in a variety of applications and vision systems. For example, embodiments of the present disclosure may be implemented for medical image analysis systems and other types of systems that perform image processing, including real-time image processing operations. While embodiments of the present disclosure are described herein with general reference to medical image analysis and endoscopy, it will be understood that the embodiments may also be applied to other medical imaging procedures, such as endoscopy, gastroscopy, colonoscopy, enteroscopy, laparoscopy, or surgical endoscopy. Furthermore, embodiments of the present disclosure may be implemented for other environments and vision systems, such as for or including LIDAR systems, surveillance, autopilots, and other imaging systems.
[0047] According to one aspect of the present disclosure, a computer-implemented system is provided that uses context information to identify user interactions and performs image processing based on the identified interactions. The system may include at least one memory (e.g., ROM, RAM, local memory, network memory, etc.) configured to store instructions and at least one processing device configured to execute the instructions (see, e.g., FIGS. 1 and 2). The at least one processing device may receive real-time video generated by an imaging device, the real-time video representing a plurality of image frames. For example, the at least one processing device may receive real-time video from a medical imaging system, such as one used during an endoscopy, gastroscopy, colonoscopy, or enteroscopy procedure. Additionally or alternatively, the image frames may include medical images, such as images of the digestive tract or other organs or regions of human tissue.
[0048] As used herein, the term "image" refers to any digital representation of a scene or field of view. The digital representation may be encoded in any suitable format, such as the Joint Photographic Experts Group (JPEG) format, the Graphics Interchange Format (GIF) format, a bitmap format, a Scalable Vector Graphics (SVG) format, or an Encapsulated PostScript (EPS) format. Similarly, the term "video" refers to any digital representation of a scene or region of interest that is composed of a sequence of multiple images. The digital representation may be encoded in any suitable format, such as the Moving Picture Experts Group (MPEG) format, Flash Video format, or Audio Video Interleave (AVI) format. In some embodiments, the sequence of images may be paired with audio.
[0049] The image frame may include a representation of a feature of interest (i.e., an anomaly or object of interest). For example, the feature of interest may include an anomaly on or in human tissue. In some embodiments, the feature of interest may include an object, such as a vehicle, a person, or other entity.
[0050] In accordance with the present disclosure, "abnormality" may include formations on or in human tissue, changes in human tissue from one type of cell to another, and / or the absence of human tissue where human tissue is expected to be present. For example, a tumor or other tissue growth may include an abnormality because there are more cells than expected. Similarly, a bruise or other change in cell type may include an abnormality because there are blood cells in a location other than where expected (i.e., outside of capillaries). Similarly, a depression in human tissue may include an abnormality because there are no cells in an expected location, resulting in a depression.
[0051] In some embodiments, the abnormality may comprise a lesion. The lesion may comprise a lesion of the gastrointestinal mucosa. The lesion may be classified histologically (e.g., according to the NICE (Narrow-Band Imaging International Colorectal Endoscopy) or Vienna classification), or morphologically (e.g., according to the Paris classification), and / or architecturally (e.g., as serrated or non-serrated). The Paris classification includes polypoid and non-polypoid lesions. Polypoid lesions may include protruding lesions, pedunculated and protruding lesions, or sessile lesions. Non-polypoid lesions may include raised, flat, depressed, or indented lesions.
[0052] Regarding the detected abnormalities, serrated lesions may include sessile serrated adenomas (SSAs); conventional serrated adenomas (TSAs); hyperplastic polyps (HPs); fibrous polyps (FPs); or mixed polyps (MPs). According to the NICE classification system, abnormalities are classified into three types: (Type 1) sessile serrated polyps or hyperplastic polyps, (Type 2) conventional adenomas, and (Type 3) carcinomas with deep submucosal invasion. The Vienna Classification divides abnormalities into five categories: (Class 1) neoplasia / dysplasia negative; (Class 2) neoplasia / dysplasia indeterminate; (Class 3) noninvasive low-grade neoplasia (low-grade adenoma / dysplasia); (Class 4) mucosal high-grade neoplasia, e.g., high-grade adenoma / dysplasia, noninvasive carcinoma (carcinoma in situ), or suspicious for invasive carcinoma; and (Class 5) invasive neoplasia, intramucosal carcinoma, submucosal carcinoma, or the like.
[0053] The processing unit of the system may include one or more image processing units. The image processing unit may be implemented as one or more neural networks trained to process real-time video and perform image operations, such as object detection and classification. In some embodiments, the processing unit includes one or more CPUs or servers. According to one aspect of the present disclosure, the processing unit may obtain contextual information to direct a user's interaction with the imaging device. In some embodiments, the contextual information may be generated by the processing unit analyzing two or more image frames in the real-time video over time. For example, the contextual information may be generated from intersection-over-union (IoU) values for the locations of detected objects in two or more image frames over time. In some embodiments, the IoU value may be compared to a threshold to determine the context of the user's interaction with the imaging device (e.g., a user navigating the imaging device and attempting to identify an object). Furthermore, in some embodiments, persistence of IoU values that meet a threshold over a predetermined number of frames or time may be required to determine user interaction with the imaging device. The processing unit may be implemented to obtain the contextual information based on an analysis of the entire image (i.e., not just identified objects). In some embodiments, the context information is obtained based on the classification information.
[0054] Additionally or alternatively, the context information may be generated based on user input received by the imaging device indicating user interaction (e.g., input indicating that the user is examining an identified object by focusing or zooming the imaging device). In such embodiments, the imaging device may provide a signal to the processing device indicating the user input received by the imaging device (e.g., by pressing a focus or zoom button). In some embodiments, persistence of the user input over a predetermined number of frames or time may be required to determine user interaction with the imaging device.
[0055] The system's processor may identify user interaction based on the context information. For example, in embodiments employing an IoU method, an IoU value of 0.5 or greater (e.g., about 0.6 or 0.7 or greater, e.g., 0.8 or 0.9) between two consecutive image frames may be used to identify that the user is examining an object of interest. In contrast, an IoU value of 0.5 or less (e.g., about 0.4 or less) between those same image frames may be used to identify that the user is navigating the imaging device or moving away from the object of interest. In either case, persistence of the IoU value (above or below a threshold) over a predetermined number of frames or time may be required to determine user interaction with the imaging device.
[0056] Additionally or alternatively, context information may be obtained based on user input to the imaging device. For example, a user pressing one or more buttons on the imaging device may provide context information indicating that the user desires to know classification information, e.g., type information, about an object of interest. Examples of user input indicating that the user desires to know more information about an object of interest include focus operations, zoom operations, image stabilization operations, dimming operations, and the like. As a further example, other user input may indicate that the user desires to navigate and identify an object. As a further example, in the case of a medical imaging device, a user may control and navigate the device to move the field of view and identify an object of interest. In the above embodiments, persistence of user input over a predetermined number of frames or time may be required to determine user interaction with the imaging device.
[0057] In some embodiments, a processing unit of the system may perform image processing on the plurality of image frames based on the obtained context information and the determined user interaction with the imaging device. In some embodiments, the image processing may be performed by applying at least one neural network (e.g., an adversarial network) trained to process frames received from the imaging device. For example, the neural network may include one or more layers configured to receive image frames as input and to output indicators of location and / or classification information of objects of interest. In some embodiments, the image processing may be performed by applying a convolutional neural network.
[0058] Consistent with embodiments of the present disclosure, a neural network may be trained by adjusting the weights of one or more nodes of the network and / or by adjusting the activation (or transfer) functions of one or more nodes of the network. For example, the weights of a neural network may be adjusted to minimize a loss function associated with the network. In some embodiments, the loss function may include a squared loss function, a hinge loss function, a logistic loss function, a cross-entropy loss function, or any other suitable loss function or combination of loss functions. In some embodiments, the activation (or transfer) function of a neural network may be modified to improve the fit between one or more models of the node and the inputs to the node. For example, the processing unit may increase or decrease the degree of a polynomial function associated with the node, change the associated function from one type to another type (e.g., from polynomial to exponential, from logarithmic to polynomial, or the like), or perform any other adjustment to the model of the node.
[0059] In some embodiments, processing the plurality of image frames may include performing object detection to detect at least one object in the plurality of image frames. For example, if an object in the image frames includes non-human tissue, the at least one processing device may identify the object (e.g., based on characteristics such as texture, color, contrast, etc.).
[0060] In some embodiments, processing the plurality of image frames may include performing classification to generate classification information for at least one detected object in the plurality of image frames. For example, if the detected object includes a lesion, the at least one processing unit may classify the lesion into one or more types (e.g., cancerous, or non-cancerous, or the like). However, the disclosed embodiments are not limited to performing classification on objects identified by an object detector. For example, classification may be performed on an image without first detecting an object in the image. Additionally, classification may be performed on sections or regions of the image that are likely to contain the object of interest (e.g., identified by a region proposal algorithm, such as a region proposal network (RPN), a fast region-based convolutional neural network (FRCN), etc.).
[0061] In some embodiments, processing the plurality of image frames may include determining an image similarity value or other specific image feature between two or more image frames or portions thereof. For example, the image similarity value may be generated based on the motion of one or more objects in the plurality of image frames, the physical similarity between one or more objects in the plurality of image frames, the similarity between two or more image frames as a whole or portions thereof, or any other feature, characteristic, or information between two or more image frames. In some embodiments, the image similarity value may be determined based on historical data of object detection, classification, and / or any other information received, captured, or calculated by the system. For example, the image similarity value may be generated from an intersection-over-union (IoU) value of the locations of detected objects in two or more image frames over time. Additionally, the image similarity value may be generated based on whether a detected object is similar to a previously detected object. Additionally, the image similarity value may be generated based on whether at least one object is part of a classification in which a user previously expressed interest. Additionally, image similarity values may be generated based on whether the user has previously performed an action (e.g., stabilized a frame, focused on an object, or any other interaction with the imaging device). In this way, the system may learn to recognize the user's preferences, thereby resulting in a more personalized and enjoyable user experience. As can be appreciated from the foregoing, the disclosed embodiments are not limited to any particular type of similarity value or process for generating it, but rather may be used with any suitable process for determining similarity values between two or more image frames or portions thereof, such as processes involving aggregating information over time, integrating information over time, averaging information over time, and / or any other method of processing or manipulating data (e.g., image data).
[0062] In some embodiments, object detection, classification, and / or similarity value generation for at least one object in the plurality of image frames may be controlled based on information received, captured, or generated by the system. For example, object detection, classification, and / or similarity value generation may be invoked or disabled based on context information (e.g., object detection may be invoked when context information indicates that a user is interacting with the imaging device to identify an object, and / or classification may be invoked when context information indicates that a user is interacting with the imaging device to examine a detected object). As an example, object detection may be invoked to detect any objects within the region of interest when context information indicates that a user is interested in an area in one or more image frames or portions thereof. Subsequently, classification may be invoked to generate classification information for the object of interest when context information indicates that the user is interested in one or more particular objects within the region of interest. In this manner, the system may continuously provide information of user interest in real time or near real time. Furthermore, in some embodiments, at least one of object detection, classification, and / or similarity value generation may be continuously enabled. For example, object detection may be performed continuously to detect one or more objects of interest in multiple frames, and the resulting output may be used in other processes of the system (e.g., classification and / or similarity value generation to generate context information, or any other function of the system). Continuous activation may be controlled automatically by the system (e.g., upon power-up), as a result of input from a user (e.g., pressing a button), or by a combination thereof.
[0063] As disclosed herein, the processor of the system may generate an overlay for display with the plurality of image frames on the video display device. Optionally, if no object is detected in the plurality of image frames, the overlay may include a null indicator or other indicator that no object was detected.
[0064] The overlay may include a boundary indicating the location of at least one detected object in the multiple image frames. For example, in embodiments where the location of at least one detected object includes a point, the overlay may include a circle, a star, or any other shape placed over the point. Additionally, in embodiments where the location includes an area, the overlay may include a boundary around the area. In some embodiments, the shape or boundary may be animated. Thus, the shape or boundary may be generated for multiple frames not only to trace the location of the detected object across frames, but also to appear animated when the frames are displayed in sequence.
[0065] In some embodiments, an overlay may be displayed with classification information, such as classification information for at least one detected object in the video feed. For example, in embodiments using the NICE classification system, the overlay may include an indicator that may be either "Type 1," "Type 2," "Type 3," "No Polyps," or "Unknown." The overlay may also include information such as a confidence score (e.g., "90%)" or the like. In some embodiments, the color, shape, pattern, or other aspect of the overlay may depend on the classification. Additionally, in embodiments providing an audio and / or vibration indicator, the duration, frequency, and / or amplitude of the sound and / or vibration may depend on whether an object was detected or on the classification.
[0066] Consistent with this disclosure, the system processor may receive real-time video from an imaging device and output video, including overlays, in real-time to a display device. Exemplary disclosures of embodiments suitable for receiving video from an imaging device and outputting video, including overlays, to a display device are provided in U.S. Application Nos. 16 / 008,006 and 16 / 008,015, both filed June 13, 2018, and both of which are expressly incorporated herein.
[0067] In some embodiments, an artificial intelligence (AI) system including one or more neural networks may be provided to determine the behavior of a physician or other medical professional during interaction with an imaging device. Several possible methods can be used to train the AI system. In one embodiment, video frames can be grouped, for example, according to a specific task-organ-disease combination. For example, a series of video frames can be collected for colonic adenoma detection or esophageal characterization of Barrett's syndrome. In these video frames, the behavior of different physicians performing the same task may share some common characteristics in the multidimensional domain analyzed by the system. When presented with similar video frames, an AI system, if properly trained, may be able to identify with a given accuracy that the physician is performing a given task in these video frames. The system may then activate an appropriate artificial intelligence sub-algorithm trained to analyze the video frames with high performance and assist the physician with on-screen information.
[0068] In other embodiments, similar results can be achieved using computer vision analysis of basic image features in the time-space domain, analyzing image features such as changes in color, velocity, contrast, motion speed, optical flow, entropy, binary patterns, texture, and the like.
[0069] In this disclosure, embodiments are described in the context of polyp detection and characterization during colonoscopy. During a traditional colonoscopy, a flexible tube containing a video camera is passed through the anus. The primary objective is to examine the entire length of the colon to identify and potentially remove small lesions (polyps) that may be precursors to colon cancer. A physician or other user may navigate through the colon by moving the flexible tube, while simultaneously continuously inspecting the colon wall for the presence of potential lesions (detection). Whenever a particular area of the image that may be a polyp catches the physician's attention, the physician may alter the navigation method by slowing down or attempting to zoom in on the suspicious area. Once the nature of the suspicious lesion is determined (characterized), appropriate action may then be taken. The physician may perform an on-site resection if the lesion is deemed a precursor to cancer, or may resume navigation for detection if not.
[0070]
[0003] Artificial intelligence systems and algorithms trained to detect polyps may be useful during the detection phase but may be intrusive at other times, such as during surgery. Similarly, artificial intelligence algorithms trained to characterize potential lesions as adenomas or non-adenomas may be useful during the characterization phase but are not necessary during the detection phase. Thus, the inventors have discovered that it is desirable to enable an artificial intelligence system or algorithm for detection only during the detection phase and an artificial intelligence system or algorithm for characterization only during the characterization phase.
[0071] Referring now to FIG. 1 , a schematic diagram of an exemplary computer-implemented system 100 for real-time video processing and overlaying information onto a video feed is provided, according to an embodiment of the present disclosure. As shown in FIG. 1 , system 100 includes an operator 101 controlling an imaging device 103. In embodiments in which the video feed includes medical video, operator 101 may include a physician or other medical professional. Imaging device 103 may include a medical imaging device, such as an X-ray device, a computed tomography (CT) device, a magnetic resonance imaging (MRI) device, an endoscopy device, or other medical imaging device that produces a video or one or more images of a human body or portion thereof. Operator 101 may control imaging device 103 by controlling the capture rate of imaging device 103 and / or the movement of imaging device 103, for example, through or relative to the human body. In some embodiments, instead of an external imaging device, such as an X-ray device, the imaging device 103 may include a Pill-Cam™ device or other form of capsule endoscopy device, or an imaging device inserted through a cavity in the body, such as an endoscopy device.
[0072] As further illustrated in FIG. 1 , the imaging device 103 may transmit captured video as multiple image frames to the overlay device 105. The overlay device 105 may include one or more processing devices that process the video as described herein. Additionally or alternatively, one or more processing devices may be implemented as separate components (not shown) that are not part of the overlay device 105. In such embodiments, the processing device may receive multiple image frames from the imaging device 103 and communicate with the overlay device 105 to transfer control or information signals for purposes of creating one or more overlays. Also, in some embodiments, the operator 101 may control the overlay device 105 in addition to the imaging device 103, for example, by controlling the sensitivity of an object detector (not shown) of the overlay device 105.
[0073] As illustrated in FIG. 1 , the overlay device 105 may augment video received from the image device 103 and then transmit the augmented video to the display device 107. In some embodiments, this augmentation may include providing one or more overlays to the video, as described herein. As further illustrated in FIG. 1 , the overlay device 105 may be configured to directly relay video from the image device 103 to the display device 107. For example, the overlay device 105 may perform a direct relay under predetermined conditions, e.g., if there are no augmentations or overlays to be generated. Additionally or alternatively, the overlay device 105 may do so if the operator 101 inputs a command to the overlay device 105 to perform a direct relay. The command may be received via one or more buttons included on the overlay device 105 and / or via an input device such as a keyboard or the like. In the event of video modifications or one or more overlays, the overlay device 105 may generate and transmit a modified video stream to the display device. The modified image may include original image frames with overlays and / or classification information to be displayed to the operator via display device 107. Display device 107 may include any suitable display or similar hardware for displaying the image or the modified image. Other types of image modifications (e.g., zoomed image of at least one object, modified image color distribution, etc.) are described herein.
[0074] 2A and 2B are schematic diagrams of exemplary computer-implemented systems 200a and 200b, respectively, for real-time image processing using context information, in accordance with embodiments of the present disclosure. Figures 2A and 2B illustrate exemplary components of exemplary computer-implemented systems 200a and 200b, respectively, consistent with disclosed embodiments. It will be understood that other configurations may be implemented, and that components may be added, removed, or rearranged in light of the present disclosure and various embodiments herein.
[0075] 2A and 2B, one or more image processors 230a and 230b may be provided. The image processors 230a and 230b may process image frames acquired by the image devices 210a and 210b, respectively. The image processors 230a and 230b may include object detectors 240a and 240b, respectively, for detecting at least one object of interest in the image frames and classifiers 250a and 250b, respectively, for generating classification information for the at least one object of interest. In some embodiments, the object detectors 240a and 240b and the classifiers 250a and 250b may be implemented using one or more neural networks trained to process the image frames. Image processors 230a and 230b may perform other image processing functions, including image modifications such as generating an overlay including at least one boundary indicating the location of at least one detected object, generating classification information for at least one object, zooming on at least one object, modifying image color distribution, or making any other adjustments or modifications to one or more image frames. Image processors 210a and 210b (similar to image processor 103 of FIG. 1) may be image processors of a medical imaging system or other types of image processors. Display devices 260a and 260b may be the same as or similar to display device 107 of FIG. 1 and may operate in the same or similar manner as described above.
[0076] The context analyzers 220a and 220b may be implemented separately from the image processing devices 230a and 230b (as shown in FIGS. 2A and 2B) or may be implemented as an integrated component (not shown) of the image processing devices 230a and 230b. The context analyzers 220a and 230b may determine operator or user interactions with the imaging devices 210a and 210b, respectively, and generate one or more outputs based on the determined user interactions. Context information may be obtained or generated by the context analyzers 220a and 220b to determine user interactions with the imaging devices 210a and 210b, respectively. For example, in some embodiments, the context analyzers 220a and 220b may calculate an intersection-over-union (IoU) value related to the location of an object in two or more image frames over time. The context analyzers 220a and 220b may compare the IoU value to a threshold to determine user interaction with the imaging devices. Additionally or alternatively, context information may be generated by context analyzers 220a and 220b by using image similarity values or other specific image features of detected objects in two or more image frames over time. The image similarity values or other specific image features of detected objects may be compared to a threshold to determine the context of a user's interaction with the imaging device (e.g., a user navigating the imaging device to identify an object). If the image similarity values or other specific image features of detected objects meet the threshold for a predetermined number of frames or time, the persistence required to determine user interaction with the imaging device may be established. Additionally or alternatively, context information may be generated manually by a user, for example, by the user pressing a focus or zoom button or providing other input to imaging devices 210a and 210b, as described herein.In these embodiments, (i) the IoU or image similarity value relative to a threshold or (ii) the identified user input may be required to persist over a predetermined number of frames or time to determine user interaction with the imaging device.
[0077] In some embodiments, similarity value generation may be performed using one or more neural networks trained to determine image similarity values or other specific image features between two or more image frames or portions thereof. In such embodiments, the neural network may determine the similarity value based on any features, characteristics, and / or information between the two or more image frames, including IoU values, whether the detected object is similar to previously detected objects, whether at least one object is part of a class in which the user has previously expressed interest, and / or whether the user is performing an action previously performed. In some embodiments, similarity value generation may be invoked or disabled based on information received, captured, and / or generated by the system, including contextual information, as described herein.
[0078] According to the exemplary configuration of FIG. 2A , the context analyzer 220a may determine operator or user interaction with the imaging device 210a and may generate instructions for the image processing device 230a based on the determined user interaction with the imaging device 210a. Context information may be obtained or generated by the context analyzer 220a to determine user interaction with the imaging device 210a. For example, in some embodiments, the context analyzer 220a may calculate an intersection-over-union (IoU) value associated with the location of an object in two or more image frames over time. The context analyzer 220a may compare the IoU value to a threshold to determine user interaction with the imaging device. Additionally or alternatively, context information may be generated manually by a user, such as by a user pressing a focus or zoom button or providing other input to the imaging device 210a, as described above. In these embodiments, (i) the IoU value relative to a threshold or (ii) the identified user input may be required to persist for a predetermined number of frames or time to determine user interaction with the imaging device.
[0079] The image processing device 230a may process the image frames based on input received by the context analyzer 220a regarding context analysis. The image processing device 230a may perform one or more image processing operations, for example, by invoking an object detector 240a, a classifier 250a, and / or other image processing components (not shown). In some embodiments, the image processing may be performed by applying one or more neural networks trained to process the image frames received from the image device 210a. For example, the context analyzer 220a may instruct the image processing device 230a to invoke the object detector 240a when the context information indicates that a user is using the image device 210a to navigate. As a further example, the context analyzer 220a may instruct the image processing device 230a to invoke the classifier 250a when the context information indicates that a user is inspecting an object of interest. As will be appreciated by those skilled in the art, image processing is not limited to object detection or classification. For example, image processing may include applying a region proposal algorithm (e.g., a Region Proposal Network (RPN), a Fast Region-Based Convolutional Neural Network (FRCN), etc.), applying an interest point detection algorithm (e.g., Features from Accelerated Segment Test (FAST), Harris's method, Maximally Stable Extremal Regions (MSER), or the like), image modification (e.g., overlaying boundary or classification information as described herein), or performing any other adjustments or modifications to one or more image frames.
[0080] As further shown in Figure 2A, image processor 230a may generate output to display device 260a, which may be the same as or similar to display device 107 of Figure 1 and may operate in the same or similar manner as described above. The output may include, for example, the original image frame with one or more overlays, such as a boundary indicating the location of detected objects within the image frame and / or classification information for objects of interest within the frame.
[0081] In the example configuration of FIG. 2B , the image processor 230b may process image frames using information provided by the context analyzer 220b or may directly process images captured by the image device 210b. The context analyzer 220b may run consistently throughout the process to determine contextual information that directs the user's interaction with the image device 210a, if possible, and provide instructions to the image processor 230b accordingly. The context analyzer 220b may also be implemented to analyze historical data, including IoU values, similarity determinations, and / or other information over time. The image processor 230b may provide a video output to the display device 260b and / or may provide one or more outputs of its image processing functions to the context analyzer 220b. The video output to the display device 260b may include the original video with or without modifications (e.g., one or more overlays, classification information, etc.) as described herein.
[0082] The context analyzer 220b may determine an operator or user interaction with the imaging device 210b and may generate instructions for the image processing device 230b based on the determined user interaction with the imaging device 210b. The context analyzer 220b may determine the user interaction using one or more image frames captured by the imaging device 210b (e.g., by calculating an IoU value between two or more frames) as disclosed herein. The context analyzer 220b may receive historical data generated by the image processing device 230b, such as object detections generated by the object detector 240b or classifications generated by the classifier 250b. The context analyzer 220b may use this information to determine the user interaction with the imaging device 210b, as described herein. Additionally, the context analyzer 220b may determine the operator interaction or user interaction based on context information previously obtained by the context analyzer 220b itself (e.g., previously calculated IoU values, similarity values, user interactions, and / or other information generated by the context analyzer 220b), as described herein.
[0083] In some embodiments, the context analyzer 220b may process multiple image frames from the image device 210b and determine that a particular region within the image frame is of user interest. The context analyzer 220b may then provide instructions to the image processor 230b to cause the object detector 240b to perform object detection and detect any objects within the identified region of interest. Thereafter, if the context information indicates that the user is interested in an object within the region of interest, the context analyzer 220b may provide instructions to the image processor 230b to cause the classifier 250b to generate classification information for the object of interest. In this manner, the system may continuously provide information of interest to the user in real time or near real time, while avoiding displaying information for objects of no interest. Advantageously, using the context information in this manner also avoids excessive processing by the object detector 240b and the classifier 250b, because processing is performed only with respect to the region of interest and the object of interest within the region derived from the context information.
[0084] The image processor 230b may process image frames based on input received by the context analyzer 220b for context analysis. Additionally, the image processor 230b may directly process image frames captured by the image device 210b without first receiving instructions from the context analyzer 220b. The image processor 230b may perform one or more image processing operations, for example, by invoking an object detector 240b, a classifier 250b, and / or other image processing components (not shown). In some embodiments, image processing may be performed by applying one or more neural networks trained to process image frames received from the image device 210b. For example, the context analyzer 220b may instruct the image processor 230b to invoke the object detector 240b when the context information indicates that a user is using the image device 210b to navigate. As a further example, the context analyzer 220b may instruct the image processor 230b to invoke the classifier 250b when the context information indicates that a user is looking up an object or feature of interest. As will be appreciated by those skilled in the art, image processing is not limited to object detection and classification. For example, image processing may include applying a region proposal algorithm (e.g., a Region Proposal Network (RPN), a Fast Region-Based Convolutional Neural Network (FRCN), etc.), applying an interest point detection algorithm (e.g., features derived from an Accelerated Segment Test (FAST), the Harris method, Maximally Stable Extremal Regions (MSER), or the like), image modification (e.g., overlaying boundary or classification information as described herein), or performing any other adjustments or modifications to one or more image frames.
[0085] 2B , the image processor 230b may generate output to the display device 260b. The output may include the original image frame with one or more image modifications (e.g., an overlay of a boundary indicating the location of the detected object within the image frame, classification information for the object of interest within the frame, a zoomed image of the object, a modified image color distribution, etc.). Additionally, the image processor 230b may provide image processing information to the context analyzer 220b. For example, the image processor 230b may provide information related to the objects detected by the object detector 240b and / or classification information generated by the classifier 250b. As a result, the context analyzer 220b may utilize this information to determine operator or user interactions, as described herein.
[0086] FIG. 3 is a flowchart of an exemplary method for processing real-time video received from an imaging device, according to an embodiment of the present disclosure. The embodiment of FIG. 3 may be implemented by one or more processing devices and other components (such as those shown in the exemplary systems of FIGS. 1 or 2). In FIG. 3, video is processed based on context information. In step 301, video is received from an imaging device, e.g., a medical imaging system. The video may include multiple image frames, which may include one or more objects of interest. In step 303, one or more neural networks trained to process the image frames may be provided. For example, an adversarial neural network may be provided to identify the presence of an object of interest (e.g., a polyp). As a further example, a convolutional neural network may be provided to classify images into one or more types (e.g., cancerous or non-cancerous) based on texture, color, or the like. In this manner, image frames may be processed in an efficient and accurate manner while being tailored to a desired application.
[0087] In step 305, context information may be obtained. The context information may indicate a user's interaction with the imaging device, as described herein. In step 307, the context information may be used to identify the user's interaction. For example, an IoU or image similarity value may be used to identify whether the user is navigating to identify an object of interest, examining the object of interest, or moving away from the object of interest. Additionally or alternatively, user input to the imaging device may provide context information that may be used to determine user interaction with the imaging device. As part of step 307, the IoU or similarity value relative to a threshold and / or the presence of user input may be required to persist for a predetermined number of frames or time before the processing device identifies that a particular user interaction with the imaging device exists. In step 309, image processing may be performed based on the identified interaction (context information) using one or more trained neural networks, as described above. For example, if the identified interaction is navigating, the image processing device may perform object detection. As another example, if the identified interaction is examining, the image processing device may perform classification. In step 311, image modifications to the received video may be performed based on the image processing. For example, as part of step 311, one or more overlays and / or classification information may be generated based on the image processing performed in step 309. As disclosed herein, the overlays may be displayed to a user or operator via a display device. For example, the displayed video output may include a boundary (e.g., a box or a star) indicating the detected object in the image frame and / or classification information of the object of interest in the image frame (e.g., a text indicator such as "Type 1," "Type 2," or "Type 3").
[0088] FIG. 4 is a flowchart of an exemplary method for invoking image processing operations based on context information directing user interaction with an imaging device, according to an embodiment of the present disclosure. The embodiment of FIG. 4 may be implemented by one or more processing devices and other components (such as those shown in the exemplary systems of FIGS. 1 or 2). In FIG. 4, object detection and classification operations are invoked based on identified user interaction with the imaging device. In step 401, the processing device may determine whether the user is navigating using the imaging device (e.g., navigating through a body part during a colonoscopy to identify an object of interest). If it is determined that the user is navigating, an object detector may be invoked in step 403. For example, a neural network trained to detect adenomas in the colon may be invoked. In step 405, the processing device may determine whether the user is inspecting an object of interest (e.g., holding the imaging device steady to analyze the object of interest in the frame). If it is determined that the user is inspecting, a classifier may be invoked in step 407. For example, a neural network trained to characterize signs of Barrett's syndrome in the esophagus may be invoked. In step 409, it may be detected whether the user is leaving the object of interest. If it is determined that the user is leaving, in step 411, the classifier may be stopped.
[0089] FIG. 5 is a flowchart of an exemplary method for generating overlay information on a real-time video feed from an image device, according to an embodiment of the present disclosure. The embodiment of FIG. 5 may be implemented by one or more processing devices and other components (such as those shown in the exemplary systems of FIGS. 1 or 2). In FIG. 5, the overlay is generated based on an analysis of context information, where the overlay representation provides, for example, location and classification information of an object within the image frames. In step 501, the processing device may detect an object within multiple image frames in the real-time video feed. This may be done by applying an object detection algorithm or a trained neural network, as described above. In step 503, a first overlay representation may be generated that includes a boundary indicating the location of the detected object within the image frames. For example, the first overlay representation may include a circle, a star, or other shape that specifies a point location of the detected object. As a further example, if the object location includes an area, the first overlay representation may include a box, rectangle, circle, or another shape positioned over the area. In step 505, the processing device may obtain context information to direct user interaction. As discussed above, the context information may be obtained by analyzing the video (i.e., IoU or image similarity methods) and / or user input (i.e., focus or zoom operations). In step 506, classification information for the object of interest in the image frame may be generated by invoking a classifier or classification algorithm as described herein. In step 504, a second overlay representation may be generated that includes the classification information. For example, the second overlay representation may include an overlay with a boundary indicating the location of the object of interest and a text indicator (e.g., "polyp" or "non-polyp") that provides the classification information.Additionally or alternatively, in some embodiments, the color, shape, pattern, or other aspect of the first and / or second overlay may depend on the detection and / or classification of the object.
[0090] FIG. 6 is an example of a display with an overlay in video based on object detection and classification, according to an embodiment of the present disclosure. In the example of FIG. 6 (as well as FIGS. 7A and 7B), the example video samples 600a, 600b, and 600c are obtained from a colonoscopy procedure. It will be understood from this disclosure that video from other procedures and imaging devices may be utilized when implementing embodiments of the present disclosure. Thus, the video samples 600a, 600b, and 600c (as well as FIGS. 7A and 7B) are non-limiting examples of the present disclosure. Additionally, by way of example, the video display of FIG. 6 (as well as FIGS. 7A and 7B) may also be presented on a display device, such as display device 107 of FIG. 1.
[0091] First overlay 601 represents an example of a graphical boundary used as an indicator for detected objects (e.g., anomalies) in the video. In the example of FIG. 6, first overlay 601 includes an indicator in the form of a solid rectangular boundary. In other embodiments, first overlay 601 may be a different shape (regular or irregular). Additionally, first overlay 601 may be displayed in a predetermined color or by transitioning from one color to another. First overlay 601 may appear in video frames 600b and 600c, which may follow video frame 600a in sequence.
[0092] The second overlay 602 presents an example classification (e.g., anomaly) of an object of interest in the video. In the example of FIG. 6, the second overlay 602 includes a text indicator identifying the type of anomaly (e.g., "Type 1" according to a classification system, such as the NICE classification system). As can be seen from the video sample 600c, the second overlay 602 may include other information besides the classification indicator. For example, a confidence indicator (e.g., "95%") associated with the classification may be included in the second overlay 602.
[0093] FIG. 7A is an example visual representation of determining an intersection-over-union (IoU) value for an object in two image frames, according to an embodiment of the present disclosure. As shown in FIG. 7A, images 700a and 700b include frames of video that include an object of interest. FIG. 7A illustrates image 700a and a subsequent image 700b. In the example of FIG. 7A, areas 701a and 701b represent the location and size of the object of interest detected in images 700a and 700b, respectively. Additionally, area 702 represents the combination of areas 701a and 701b, and represents a visual representation of determining an IoU value for the object detected in images 700a and 700b. In some embodiments, the IoU value is calculated using the following formula:
[0094]
number
[0095] In the above IoU formula, the Area of Overlap is the area where the detected object is present in both images, and the Area of Union is the total area where the detected object is present in the two images. In the example of FIG. 7A , the IoU value may be estimated using the ratio of the overlap area between areas 701a and 701b (i.e., the center of area 702) to the union area between areas 701a and 701b (i.e., the entire area 702). In the example of FIG. 7A , the IoU value may be considered low given that the center of area 702 is relatively smaller than the entire area 702. In some embodiments, this may indicate that the user is moving away from the object of interest.
[0096] FIG. 7B is another example of a visual representation of determining an intersection-over-union (IoU) value for an object in two image frames, according to an embodiment of the present disclosure. As shown in FIG. 7B, images 710a and 710b comprise frames of video containing an object of interest. FIG. 7B illustrates image 710a and a subsequent image 710b (similar to images 700a and 700b). In the example of FIG. 7B, areas 711a and 711b represent the location and size of the object of interest detected in images 710a and 710b, respectively. Additionally, area 712 represents the combination of areas 711a and 711b, providing a visual representation of determining the IoU value for the object detected in images 710a and 710b. The same IoU formula as described above for FIG. 7A may be used to determine the IoU value. In the example of Figure 7B, the IoU value may be estimated using the ratio between the overlapping area of areas 711a and 711b (i.e., the center of area 712) and the combined area of areas 711a and 711b (i.e., the entire area 712). In the example of Figure 7B, the IoU value may be considered high given that the center of area 712 is relatively equal to the entire area 712. In some embodiments, this may indicate that the user is examining an object of interest.
[0097] 8 is a flowchart of an exemplary method for invoking an object detector and classifier in which context information is determined based on image similarity values between multiple frames, according to an embodiment of the present disclosure. However, it will be understood that this method may be used in combination with other methods for determining context information, such as IoU values, detection or classification of one or more objects in an image frame, or methods based on input received by a medical imaging system from a user. The embodiment of FIG. 8 may be implemented by one or more processing units and other components (such as those shown in the exemplary systems of FIGS. 1 or 2).
[0098] In step 801, an object detector (e.g., object detectors 240a and 240b of FIGS. 2A and 2B) is invoked to detect an object of interest in a first image frame. For example, one or more neural networks trained to detect a particular disease or abnormality (e.g., a colon adenoma) may be invoked to determine whether the particular disease or abnormality is present in the first image frame. The object detector may be invoked for the same or similar reasons as discussed above in connection with other embodiments. In step 803, the object detector processes a second image frame obtained subsequent to the first image frame to determine the presence or absence of an object of interest in the second image frame. For example, the one or more neural networks may detect the presence of a polyp in the second image frame that is consistent with a colon adenoma.
[0099] In step 805, a determination is made as to whether a similarity value between the first image frame and the second image frame is greater than or equal to a predetermined threshold for determining context information. The determination may be made using an image similarity evaluator (not shown). The similarity evaluator may be implemented with a processing device and may include one or more algorithms that process the image frames as input and output a similarity value between the two or more image frames along with image features such as image overlap, edges, interest points, regions of interest, color distribution, or the like. In some embodiments, the similarity evaluator may be configured to output a number between 0 and 1 (e.g., 0.587), where a similarity value of 1 means the two or more image frames are identical and a similarity value of 0 means the two or more image frames have no similarity. In some embodiments, the image similarity evaluator may be part of a context analyzer (e.g., context analyzers 220a and 220b of FIGS. 2A and 2B) or an image processor (e.g., image processors 230a and 230b of FIGS. 2A and 2B), such as an object detector (e.g., object detectors 240a and 240b of FIGS. 2A and 2B) or a classifier (e.g., classifiers 250a and 250b of FIGS. 2A and 2B).
[0100] The calculation of the similarity value may be performed using one or more features of the first and second image frames. For example, a determination may be made as to whether a sufficient portion of the first image frame is included in the second image frame to identify that the user is looking at the object of interest. As a non-limiting example, if at least 0.5 (e.g., about 0.6 or 0.7 or more, e.g., 0.8 or 0.9) of the first image frame is included in the second image frame, this may be used to identify that the user is looking at the object of interest. In contrast, if less than 0.5 (e.g., about 0.4 or less) of the first image frame is included in the second image frame, this may be used to identify that the user is navigating the imaging device or moving away from the object of interest. However, it will be understood that this determination may also be made using other image features, such as edges, interest points, regions of interest, and color distribution.
[0101] In step 807, if the context information indicates that the user is not examining the object of interest, such as by determining that the image similarity value is below a predetermined threshold, the object detector remains invoked to obtain its output, process the next image frame, and start over at step 803 of the example method of FIG. 8 . In some embodiments, the object detector may be disabled in step 807. For example, the object detector may be disabled if the context information indicates that the user no longer wishes to detect objects. This may be determined, for example, when the user interacts with an input device (e.g., a button, mouse, keyboard, etc.) to disable the object detector. In this way, detection is performed efficiently and only when needed, thereby, for example, preventing the display from being overcrowded with unnecessary information.
[0102] In step 809, image modifications are performed based on the output of the object detector to modify the received image frame. For example, embodiments of the present disclosure may generate overlay information on a real-time video feed from an imaging device. The overlay information may include, for example, a location of an object of interest detected by the object detector, e.g., a circle, star, or other shape that designates the location of the detected object. As a further example, if the object location includes an area, the overlay information may include a box, rectangle, circle, or another shape positioned over the area. However, it will be understood that other image modifications may be used to draw a user's attention to the detected object, for example, to zoom in on the area of the detected object, change the image color distribution, or the like.
[0103] In step 811, a classifier (e.g., classifiers 250a and 250b in FIGS. 2A and 2B) is invoked to generate classification information for at least one detected object, consistent with disclosed embodiments. For example, if the detected object includes a lesion, the classifier may classify the lesion into one or more types (e.g., cancerous or non-cancerous, or the like). In some embodiments, one or more neural networks (e.g., adversarial neural networks) trained to classify objects may be invoked to classify the detected object, consistent with disclosed embodiments. In step 813, both the object detector and the classifier process the next frame (e.g., a third image frame obtained following the second image frame) to determine whether an object of interest is present or absent in that image frame and generate classification information if an object of interest is detected. For example, one or more neural networks may detect the presence of a polyp in an image frame that is consistent with a colon adenoma, and may then generate a label such as "Adenoma" if it determines that the polyp is in fact an adenoma, or "Non-Adenoma" if it determines that the polyp is not an adenoma, along with a confidence score (e.g., "90%).
[0104] In step 815, a determination is made as to whether a similarity value between the image frames (e.g., the second and third image frames) is greater than or equal to a predetermined threshold for determining context information. This may be performed in the same or similar manner as described above in connection with step 805. In step 817, if the context information indicates that the user is no longer examining the object of interest, such as by determining that the image similarity value is less than a predetermined threshold, the classifier is disabled, and the object detector remains invoked to process the next image frame and start over at step 803. In this manner, classification is performed efficiently and only when needed, thereby, for example, preventing the display from becoming overcrowded with unnecessary information. In contrast, in step 819, if the context information indicates that the user continues to examine the object of interest, the classifier processes N (i.e., two or more) image frames to generate classification information for at least one detected object. An algorithm may be applied to the classifier outputs for all N image frames to generate a single output. For example, a moving average calculation may be applied to combine the classifier outputs for each image frame across the time dimension. Because the classification of a particular polyp into type (e.g., adenoma or non-adenoma) may be affected by different characteristics (e.g., texture, color, size, shape, etc.), the output of the classifier may be affected by noise in some of the N frames in which a polyp is present. To reduce this phenomenon, a form of moving average can be implemented that combines the output of the classifier for the last N frames. As a non-limiting example, an arithmetic mean may be calculated, although other mathematical and statistical formulations can be used to achieve the same result.
[0105] In step 821, image modification is performed based on the output of the classifier to modify the received image frame. For example, overlay information may be generated for the detected object on a real-time video feed from the imaging device in the same or similar manner as described above in connection with step 809. In addition, the overlay information may be displayed with classification information generated by the classifier for the detected object. The classification information may include information the same as or similar to that described above in connection with step 813. In steps 823a, 823b, and 823c, for example, different classification information may be generated for the detected object depending on the classification. In step 823a, if the detected object is a polyp classified as an adenoma by the classifier, an "Adenoma" label may be generated with a red square around the detected object. In step 823b, if the detected object is a polyp classified as a non-adenoma by the classifier, a "Non-Adenoma" label may be generated with a white square around the detected object. In step 823c, if the detected object cannot be classified by the classifier, for example as a result of being out of focus, corrupted image data, etc., an "Unclassified" label may be generated with a grey box around the detected object.
[0106] In step 825, both the object detector and the classifier process the next available image frame to determine whether or not an object of interest is present, and if an object of interest is detected, generate classification information and start over at step 815 of the method of FIG. 8.
[0107] The present disclosure has been presented for illustrative purposes. It is not exhaustive and is not limited to the precise form or embodiment disclosed. Modifications and adaptations of the embodiments will become apparent from consideration of the specification and practice of the disclosed embodiments. For example, while the described implementations include hardware, systems and methods consistent with this disclosure can be implemented using both hardware and software. In addition, while certain components have been described as being coupled to each other, such components may be integrated with each other or distributed in any suitable manner.
[0108] Furthermore, while exemplary embodiments have been described herein, the scope includes any and all embodiments having equivalent elements, modifications, omissions, combinations (e.g., combinations of aspects across various embodiments), adaptations, and / or alterations based on the present disclosure. Claim elements are to be interpreted broadly based on the language employed in the claims and are not limited to the examples described herein or during the prosecution of this application, which examples are to be construed as non-exclusive. Furthermore, the steps of the disclosed methods can be modified in any way, including by changing the order of steps and / or inserting or deleting steps.
[0109] The features and advantages of the present disclosure will be apparent from the detailed specification, and accordingly, the appended claims are intended to cover all systems and methods that fall within the true spirit and scope of the present disclosure. As used herein, the indefinite articles "a" and "an" mean "one or more." Similarly, the use of plural terms does not necessarily indicate a plurality unless the given context clearly indicates otherwise. Also, words such as "and" or "or" mean "and / or" unless specifically indicated otherwise. Moreover, since numerous modifications and variations will readily occur from review of the present disclosure, it is not desired to limit the disclosure to the exact structure and operation as illustrated and described, and therefore, resort may be made to all suitable modifications and equivalents that fall within the scope of the present disclosure.
[0110] Other embodiments will be apparent from consideration of the specification and practice of the embodiments disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the disclosed embodiments being indicated by the appended claims.
[0111] According to some embodiments, the operations, techniques, and / or components described herein can be implemented by a device or system that can include one or more special-purpose computing devices. The special-purpose computing device can include digital electronic devices, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), that can be hardwired to perform the operations, techniques, and / or components described herein or that are permanently programmed to perform the operations, techniques, and / or components described herein, or can include one or more hardware processing devices that are programmed to perform such features of the present disclosure according to program instructions in firmware, memory, other storage, or a combination. Such special-purpose computing devices can also combine custom hardwired logic, ASICs, or FPGAs with custom programming to achieve the techniques and other features of the present disclosure. The special-purpose computing device can be a desktop computing system, a portable computing system, a handheld device, a network device, or any other device that can incorporate hardwired and / or program logic to implement the techniques and other features of the present disclosure.
[0112] One or more special-purpose computing devices may generally be controlled and regulated by operating system software, such as iOS, Android, Blackberry, Chrome OS, Windows XP, Windows Vista, Windows 7, Windows 8, Windows Server, Windows CE, Unix, Linux, SunOS, Solaris, VxWorks, or other compatible operating systems. In other embodiments, the computing device may be controlled by its own operating system. An operating system may, among other things, control and schedule computer processes for execution, provide memory management, file system, networking, I / O services, and provide user interface functionality, such as a graphical user interface ("GUI").
[0113] Additionally, while aspects of the disclosed embodiments are described as relating to data stored in memory and other tangible computer-readable storage media, those skilled in the art will appreciate that these aspects can also be stored on and executed from many types of tangible computer-readable media, for example, secondary storage such as a hard disk, floppy disk, or CD-ROM, or other forms of RAM or ROM. Accordingly, the disclosed embodiments are not limited to the examples set forth above, but rather are defined by the appended claims in light of their fullest range of equivalents.
[0114] Furthermore, although exemplary embodiments have been described herein, the scope includes any and all embodiments having equivalent elements, modifications, omissions, combinations (e.g., combinations of aspects across various embodiments), adaptations, or variations based on this disclosure. The elements in the claims are to be interpreted broadly based on the language employed in the claims and are not limited to the examples described herein or during the prosecution of this application, which examples are to be construed as non-exclusive. Furthermore, the steps of the disclosed methods can be modified in any way, including changing the order of steps or inserting or deleting steps.
[0115] Accordingly, it is intended that the specification and examples be considered as exemplary only, with a true scope and spirit being indicated by the following claims and their full scope equivalents.
Claims
1. at least one memory configured to store instructions; receiving real-time video generated by a medical imaging system and including a plurality of image frames; and While receiving real-time video generated by the medical imaging system: processing the plurality of image frames of the real-time video to obtain contextual information that directs a user's interaction with the medical imaging system; performing object detection to detect at least one object within the plurality of image frames; performing classification to generate classification information for the at least one detected object in the plurality of image frames; and performing image modification to modify the received real-time video based on at least one of the object detection and the classification, and generating a representation of the real-time video with the image modification on a video display device; at least one processing unit configured to execute instructions to perform operations including:
1. A computer-implemented system for real-time video processing, comprising: The system, wherein the at least one processing unit is further configured to invoke at least one of the object detection and the classification based on the user's interaction with the medical imaging system as directed by the context information.
2. 10. The system of claim 1, wherein at least one of the object detection and classification is performed by applying at least one neural network trained to process frames received from the medical imaging system.
3. 3. The system of claim 1, wherein the at least one processing unit is configured to invoke the object detection when the context information indicates the user's interaction with the medical imaging system to identify the object.
4. 4. The system of claim 3, wherein the at least one processing unit is configured to disable the object detection when the context information indicates that the user is no longer interacting with the medical imaging system to identify the object.
5. 5. The system of claim 1, wherein the at least one processing unit is configured to invoke the classification when the context information indicates the user's interaction with the medical imaging system to examine at least one object in the plurality of image frames.
6. 6. The system of claim 5, wherein the at least one processing unit is further configured to invalidate the classification when the context information indicates that the user is no longer interacting with the medical imaging system to examine at least one object in the plurality of image frames.
7. 7. The system of claim 1, wherein the at least one processing unit is further configured to invoke object detection when the context information indicates the user's interaction with the medical imaging system in which the user is interested in an area within the plurality of image frames that includes at least one object, and to invoke classification when the context information indicates the user's interaction with the medical imaging system in which the user is interested in the at least one object.
8. 8. The system of claim 1, wherein the at least one processing unit is further configured to perform aggregation of two or more frames containing the at least one object, and wherein the at least one processing unit is further configured to invoke the aggregation based on the context information.
9. 9. The system of claim 1, wherein the image modifications include at least one overlay including at least one boundary indicating the location of the at least one detected object, classification information for the at least one object, a zoomed image of the at least one object, or a modified image color distribution.
10. 10. The system of claim 1, wherein the at least one processing unit is configured to generate the context information based on an Intersection over Union (IoU) value of regions of the at least one detected object location in two or more image frames over time.
11. The system of claim 1 , wherein the at least one processing unit is configured to generate the context information based on image similarity values in two or more image frames.
12. The system of claim 1 , wherein the at least one processing unit is configured to generate the context information based on input received from the user by the medical imaging system.
13. The system of claim 1 , wherein the plurality of image frames includes image frames of the gastrointestinal tract.
14. 14. The system of claim 1, wherein the frames include images from a medical imaging device used during at least one of endoscopy, gastroscopy, colonoscopy, enteroscopy, laparoscopy, or surgical endoscopy.
15. The system of claim 1 , wherein the at least one detected object is an anomaly.
16. 16. The system of claim 15, wherein the abnormality comprises at least one of a formation on or in human tissue, a change in human tissue from one type of cell to another type of cell, an absence of human tissue where human tissue is expected to be present, or a lesion.
17. receiving real-time video generated by a medical imaging system and including a plurality of image frames; providing at least one neural network trained to process image frames from the medical imaging system; processing the plurality of image frames of the real-time video to obtain contextual information that directs a user's interaction with the medical imaging system; identifying a type of the interaction based on the context information; applying the at least one trained neural network to perform real-time processing on the plurality of image frames based on the identified interaction type; A method for real-time video processing comprising:
18. 20. The method of claim 17, wherein performing real-time processing includes performing at least one of object detection to detect at least one object in the plurality of image frames, classification to generate classification information for the at least one detected object, and image modification to modify the received real-time video.
19. The method of claim 18 , wherein the object detection is invoked when the identified interaction is a user interaction with the medical imaging system to navigate and identify an object.
20. 20. The method of claim 19, wherein the object detection is disabled when the context information indicates that there is no longer user interaction with the medical imaging system to navigate and identify objects.
21. 21. The method of claim 18, wherein the classification is invoked when the identified interaction is a user interaction with the medical imaging system to examine at least one detected object in the plurality of image frames.
22. 22. The method of any one of claims 18 to 21, wherein the classification is disabled when the context information indicates that there is no longer any user interaction with the medical imaging system to examine at least one detected object in the plurality of image frames.
23. 23. The method of any one of claims 18 to 22, wherein object detection is invoked when context information indicates the user's interaction with the medical imaging system in which the user is interested in an area within the plurality of image frames that includes at least one object, and wherein classification is invoked when context information indicates the user's interaction with the medical imaging system in which the user is interested in the at least one object.
24. 24. The method of any one of claims 18 to 23, wherein at least one of the object detection and classification is performed by applying at least one neural network trained to process frames received from the medical imaging system.
25. 25. The method of any one of claims 18 to 24, wherein the image modification includes at least one overlay including at least one boundary indicating the location of the at least one detected object, classification information for the at least one detected object, a zoomed image of the at least one detected object, or a modified image color distribution.
26. 26. The method of any one of claims 18 to 25, wherein the at least one detected object is an anomaly.
27. 27. The method of claim 26, wherein the abnormality comprises at least one of a formation on or in human tissue, a change in human tissue from one type of cell to another type of cell, an absence of human tissue where human tissue is expected to be present, or a lesion.
28. 28. The method of any one of claims 17 to 27, further comprising performing an aggregation of two or more frames containing at least one object based on the context information.
29. 29. The method of any one of claims 17 to 28, wherein the plurality of image frames comprises image frames of the gastrointestinal tract.
30. 30. The method of any one of claims 17 to 29, wherein the frames comprise images from a medical imaging device used during at least one of endoscopy, gastroscopy, colonoscopy, enteroscopy, laparoscopy, or surgical endoscopy.
Citation Information
Patent Citations
Advanced medical image processing wizard
JP2018161488A
Endoscopy Video Feature Enhancement Platform
US20190297276A1
System and methods for automatic polyp detection using convolutional neural networks
WO2016161115A1
Endoscope system, processor device, and method for operating endoscope system
WO2018179986A1
Endoscope system and method for operating same
WO2018179991A1