A computer-implemented system and method for intelligent image analysis using spatio-temporal information
A computer-implemented system using spatio-temporal processing and neural networks addresses the inefficiencies in medical image analysis by accurately detecting features of interest in endoscopic images, improving diagnosis and reducing examination time through intelligent image processing and real-time feedback.
Patent Information
- Application Number
- JP2025500223
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-08
- Filing Date
- 2023-07-07
- Publication Date
- 2025-07-10
AI Technical Summary
Existing video processing and image analysis systems in medical procedures, such as endoscopy and capsule endoscopy, struggle with accurately detecting features of interest, like abnormalities in gastrointestinal tracts, due to inefficiencies and the inability to intelligently analyze image sequences, leading to potential misdiagnosis and increased examination time.
A computer-implemented system utilizing spatio-temporal processing modules and neural networks to analyze temporally ordered images from endoscopic or capsule procedures, employing local and global spatio-temporal processing to detect and score features of interest, generating reports on the presence and likelihood of abnormalities, and providing real-time feedback.
Enhances the accuracy and efficiency of medical image analysis by intelligently processing large image collections, reducing the risk of misdiagnosis and minimizing examination time, while providing physicians with real-time feedback and recommended actions.
Smart Images

Figure 2025521918000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to the field of video processing and image analysis. More specifically, without limitation, the present disclosure relates to systems, methods, and computer-readable media for performing intelligent image analysis that processes imaging video content from an imaging device to determine the presence of features of one or more objects of interest or actions performed during a medical procedure. The systems and methods disclosed herein can be used in various applications including medical image analysis and diagnosis.
Background Art
[0002] In video processing and image analysis systems, it is often desirable to detect objects or features of interest. The features of interest may be humans, places, or things. In some applications such as systems and methods for medical image analysis, the location and classification of detected features of interest (e.g., abnormalities such as on human tissue or its formation) are important for patient diagnosis. However, existing computer-implemented systems and methods have a number of drawbacks including the inability to accurately detect features of interest and / or the inability to recognize characteristics regarding features of interest. In addition, existing systems and methods are inefficient and do not provide a way to intelligently analyze images, including regarding the presence of image sequences or events.
[0003] Modern medical procedures require accurate and precise examinations of patients' bodies and organs. Endoscopy is a medical procedure aimed at providing an internal video image of a patient's body and organs to a physician for diagnosis. In the gastrointestinal tract of the human body, the procedure can be performed by introducing a probe with a video camera through the patient's mouth or anus. During the endoscopy procedure, the physician manually navigates the probe through the gastrointestinal tract while viewing the video on a display device in real time. The video can also be imaged, stored, and examined after the endoscopy procedure. Alternatively, capsule endoscopy is a procedure in which the patient swallows a capsule containing a small camera to examine the patient's gastrointestinal tract. The sequence of images captured by the capsule during movement is wirelessly transmitted to a receiver and stored for examination by a physician after the procedure is completed. The frame rate of the capsule device can be varied (e.g., 2 - 6 frames per second), and a large number of images can be captured during the examination process.
[0004] From a computer vision perspective, the imaging content from either real-time video endoscopy or capsule procedures is a temporally ordered sequence of consecutive images that contains information about the patient, e.g., information about the internal mucosa of the gastrointestinal tract. Accurate and precise analysis of the imaging data is important for properly examining the patient and identifying features of lesions, polyps, or other areas of interest. Also, typically, there are a large number of images collected for each patient. One of the most important medical tasks that need to be performed by a physician is to examine this large collection of images in order to perform an appropriate diagnosis, including determining the presence or absence of features of interest such as pathological sites in the imaged mucosa. However, manually examining these images is time-consuming and inefficient. As a result, the review process can lead to the physician making errors and / or misdiagnosing.
[0005] To improve diagnosis, reduce the time required for medical imaging examinations, and reduce the potential for errors, the inventors have determined that it would be desirable to have a computer-implemented system and method that can intelligently process images and identify conditions or other features of interest in all images from video endoscopy or capsule procedures, or other medical procedures. For example, features of interest can also include actions being performed on or in the image, anatomical or other locations of interest in the image, clinical metric levels of the image, etc. Trained neural networks, spatio-temporal image analysis, and other features and techniques are disclosed herein for this purpose. As will be appreciated from this disclosure, the invention and embodiments are applicable to a variety of imaging and analysis applications and are not limited to the examples presented herein.
Summary of the Invention
[0006] Embodiments of the present disclosure include systems, methods, and computer-readable media for performing intelligent image analysis, such as processing an image captured by an imaging device and determining the presence of features of one or more objects of interest. Systems and methods consistent with the present disclosure can provide advantages over existing systems and technologies, including addressing one or more of the disadvantages mentioned above and / or the limitations of existing systems and technologies. Consistent with some of the disclosed embodiments, systems, methods, and computer-readable media are provided for processing images from video endoscopy or capsule procedures or other medical procedures, where the images are temporally ordered. Example embodiments include systems and methods for intelligently processing captured images using spatio-temporal information to accurately evaluate the likelihood of the presence of abnormalities, medical conditions, or other features of interest in the image. As a further example, the features of interest may be parameters or statistics related to an endoscopy or capsule procedure or other medical procedure. For example, features of interest in an endoscopy procedure may be the clean withdrawal time (withdrawal time excluding biopsy time) or the time for a probe or capsule to pass through an organ. Features of interest in an image can also be determined based on the presence or absence of characteristics related to that feature of interest. These and other embodiments, features, and implementations are described more fully herein. A feature of interest may be any feature in or related to one or more images, particularly any feature in or related to a scene or field of view represented in one or more images that is distinguishable or detectable by analyzing the image or each image. A feature of interest may be, for example, an object, a location, an action, or a state (e.g., clinical indicator level).
[0007] In some embodiments, the images captured by an imaging device such as an endoscopic video camera or a capsule camera include images of the gastrointestinal tract or organs. The images can be collected from a medical imaging device used, for example, during a gastroscopy, a colonoscopy, or a small bowel endoscopy. The features of interest in the images can be, for example, abnormalities or other medical conditions. The abnormalities or medical conditions can include on the human tissue, or the formation of human tissue, the change from one cell type of human tissue to another cell type, the disappearance of human tissue from its predicted position, or on the human tissue, or the formation of human tissue. The formation can include lesions, polypoid lesions, or non-polypoid lesions. Other examples of features of interest include anatomical or other locations, actions, clinical metrics (e.g., cleanliness), and the like. Thus, as recognized from the present disclosure, the embodiments as examples are not specific to any single disease in a medical situation but can be utilized in a generally applicable manner.
[0008] According to one general aspect of the present disclosure, a computer-implemented system is provided for processing an image captured by an imaging device. The computer-implemented system can include at least one processor configured to detect features of at least one object of interest in the image captured by the imaging device. The at least one processor receives an ordered set of temporally ordered images from the captured image, determines the presence of characteristics regarding features of at least one object of interest in each image of each subset of the images, and uses a local spatio-temporal processing module configured to annotate the subset images with feature vectors based on the determined characteristics in each image of each subset of the images to individually analyze one or more subsets of the ordered set of images, and uses a global spatio-temporal processing module configured to improve the determined characteristics associated with each subset of the images to process a set of feature vectors of the ordered set of images, where each feature vector includes information about each determined characteristic of features of at least one object of interest, and can be configured to calculate, for each image, a numerical value representing the presence of features of at least one object of interest and calculated using the improved characteristics associated with each subset of the images and spatio-temporal information. Further, the at least one processor can be configured to generate a report regarding features of at least one object of interest using the numerical values associated with each image of each set of the ordered set of images. The report can be generated after the completion of an endoscopic examination or other medical procedure. The report can include information regarding all features of interest identified in the processed images.
[0009] At least one processor of a system implemented by a computer can further be configured to determine the likelihood of characteristics regarding at least one feature of interest in each image of a subset of images. Additionally, the at least one processor can be configured to encode each image of the subset of images and determine the likelihood of characteristics in each image of the subset of images by aggregating spatio-temporal information of the determined characteristics using a regression neural network or a temporal convolutional network.
[0010] To improve the determined characteristics, a non-causal temporal convolutional network can be utilized. For example, at least one processor of the system can be configured to apply a non-causal temporal convolutional network to improve the likelihood of characteristics in each image of the subset of images. The at least one processor can further be configured to improve the likelihood of characteristics by applying one or more signal processing techniques including, for example, low-pass filtering and / or Gaussian smoothing.
[0011] According to a further aspect, at least one processor of the system can be configured to analyze an ordered set of images using a local spatio-temporal processing module to determine the presence of a characteristic by determining a vector of quality scores, where each quality score in the vector of quality scores corresponds to each image of a subset of the images. Additionally, at least one processor can be configured to process the ordered set of images using a global spatio-temporal processing module by applying signal processing techniques to improve the quality scores of each image of one or more subsets of the ordered set of images. At least one processor can further be configured to analyze one or more subsets of the ordered set of images using a local spatio-temporal processing module to determine the presence of a characteristic by generating a binary mask for pixels for each image of a subset of the images using a deep convolutional neural network. At least one processor can further be configured to process one or more subsets of the ordered set of images using a global spatio-temporal processing module by applying morphological operations that utilize prior information about the shape and distribution of the determined characteristic to improve the binary mask for image segmentation.
[0012] As disclosed herein, the implementation forms can include one or more of the following features. The determined probability of the characteristics in each image of the subset of images can include a floating-point value between 0 and 1. The quality score can be an ordinal number between 0 and R, where score 0 represents the lowest quality and score R represents the highest quality. The numerical value can be associated with each image and is interpretable to determine the probability of identifying the characteristics of at least one object of interest in the image. When at least one feature of the object of interest is not detected, the output can be a first numerical value for the image. When at least one feature of the object of interest is detected, the output can be a second numerical value for the image. The size or volume of the subset of images can be configurable by the user of the system. The size or volume of the subset of images can be determined dynamically based on the required features of the object of interest. The size or volume of the subset of images can be determined dynamically based on the determined characteristics. One or more subsets of the images can include shared images.
[0013] Other general aspects of the present disclosure relate to a computer-implemented system for spatio-temporal analysis of images captured by an imaging device. The computer-implemented system can comprise at least one processor configured to receive a video captured by an imaging device, the video including a plurality of image frames. The at least one processor is further configured to access a temporally ordered set of images from the captured images and use an event detector module to detect the occurrence of an event in the temporally ordered set of images, wherein a start time and an end time are identified by a start image frame and an end image frame in the temporally ordered set of images, use a frame selector module to select an image from a group of images whose range is determined by a start image frame and an end image frame in the temporally ordered set of images, based on an associated score and a quality score of the image, wherein the associated score of the selected image indicates the presence of at least one feature of interest, use an object descriptor module to fuse a subset of images identified based on spatial and temporal coherence, using spatio-temporal information from the selected images, based on the matching presence of at least one feature of interest, and configure the temporally ordered set of images to be divided into time intervals that satisfy the temporal coherence of the selected task.
[0014] According to the disclosed system, at least one processor can further be configured to use a local spatio-temporal processing module to determine spatio-temporal information about characteristics of at least one feature of interest for a subset of images of video content, and use a global spatio-temporal processing module to determine spatio-temporal information for all images of the video content. Additionally, at least one processor can be configured to divide a temporally ordered set of images into temporal intervals by identifying a subset of the temporally ordered set of images in which at least one feature of interest is present. At least one processor can also be configured to identify a subset of the temporally ordered set of images in which at least one feature of interest is present by adding bookmarks to images in the temporally ordered set of images, where the images to which bookmarks are added are part of the subset of the temporally ordered set of images. Additionally, or alternatively, at least one processor can be configured to identify a subset of the temporally ordered set of images in which at least one feature of interest is present by extracting a set of images from a subset of the temporally ordered set of images.
[0015] Embodiments can include one or more of the following features. The extracted set of images can include characteristics regarding at least one feature of interest. Color can be varied according to the level of relevance of the images of the subset of the temporally ordered set of images to at least one feature of interest. Color can be varied according to the level of relevance of the images of the subset of the temporally ordered set of images to characteristics regarding at least one feature of interest.
[0016] Another general aspect involves a computer-implemented system for performing multiple tasks on a set of images. The computer-implemented system can comprise at least one processor configured to receive video captured by an imaging device, the video including a set of image frames. The at least one processor can further receive a plurality of tasks, where at least one task is associated with a request to identify features of at least one object of interest in the set of images, and can use a local spatio-temporal processing module to analyze a subset of the images of the set of images to identify the presence of characteristics associated with the at least one object of interest, and can be configured to repeatedly execute a time-series analysis module for each task of the plurality of tasks to associate a numerical score for each task with each image of the set of images.
[0017] In accordance with the present disclosure, a system of one or more computers has software, firmware, hardware, or a combination thereof installed on the system and, when in operation, can be configured to cause the system to perform operations or actions. One or more computer programs, when executed by a data processing apparatus (such as one or more processors), include instructions to cause the apparatus to perform such operations or actions and can be configured to perform such operations or actions.
[0018] The systems and methods consistent with the present disclosure can be implemented using any suitable combination of software, firmware, and hardware. Embodiments of the present disclosure can include programs and machine-implemented instructions for specifically performing functions associated with the disclosed operations or actions. Additionally, a non-transitory computer-readable storage medium can be used, which stores program instructions executable by at least one processor for performing the steps and / or methods described herein.
[0019] It will be understood that the foregoing general description and the following detailed description are by way of example and explanation only and are not restrictive of the disclosed embodiments.
Brief Description of the Drawings
[0020] The following drawings, which form a part of this specification, illustrate several embodiments of the present disclosure and, together with the description, serve to explain the principles and features of the disclosed embodiments.
[0021]
Figure 1A
Figure 1B
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 6C
Figure 6D
Figure 6E
Figure 7
Figure 8
Figure 9
Figure 10
[0022] Exemplary embodiments are described below with reference to the accompanying drawings. The drawings are not necessarily drawn to scale. Examples and features of the disclosed principles are described herein, but modifications, adaptations, and other implementations are possible without departing from the spirit and scope of the disclosed embodiments. Also, the terms "comprising," "having," "containing," "including," and other similar forms are intended to be equivalent in meaning, and it is not intended that the items following any one of these terms be an exhaustive list of such items, nor is it intended that the items be limited to those listed. It should also be noted that, as used herein and in the scope of the appended claims, the singular forms "a" and "the" include plural referents unless the context clearly indicates otherwise.
[0023] In the following description, various examples are provided for illustrative purposes. However, it will be recognized that the present disclosure can be practiced without one or more of these details.
[0024] Throughout this disclosure, reference is made to "the disclosed embodiments," which refers to the inventive ideas, concepts, and / or disclosures described herein. A number of related or unrelated embodiments are described throughout this disclosure. The fact that some "disclosed embodiments" are described as presenting a certain feature or characteristic does not mean that other disclosed embodiments necessarily share that feature or characteristic.
[0025] The embodiments described herein include a non - transitory computer - readable medium containing instructions that, when executed by at least one processor, cause the at least one processor to execute a set of methods or operations. The non - transitory computer - readable medium can be any medium that can store data in any memory such that the data can be read by any computing device having a processor for executing any other instructions stored in some method or memory. The non - transitory computer - readable medium can be implemented as software, firmware, hardware, or any combination thereof. The software can preferably be realized as an application program tangibly embodied on a computer - readable medium that is composed of a program storage unit or component, or a certain device and / or combination of devices. The application program can be uploaded to and executed by a machine having any suitable architecture. Preferably, the machine can be realized on a computer platform having hardware such as one or more central processing units (CPUs), memory, and an input / output interface. The computer platform can also include an operating system and microinstruction code. The various processes and functions described in this disclosure can be part of the microinstruction code, part of the application program, or any combination thereof, and they can be executed by the CPU regardless of whether such a computer or processor is explicitly shown. In addition, various other peripheral units can be connected to the computer platform such as additional data storage units and printing units. Further, the non - transitory computer - readable medium can be any computer - readable medium except a transitory propagation signal.
[0026] The memory can include any mechanism for storing electronic data or instructions, including random access memory (RAM), read-only memory (ROM), hard disks, optical disks, magnetic media, flash memory, and other permanent, fixed, volatile, or non-volatile memory. The memory can be arranged in a specific location or distributed and can include one or more separate storage devices capable of storing data structures, instructions, or any other data. The memory can further include a portion of the memory that contains instructions for execution by the processor. The memory can also be used as a working memory device for the processor or as temporary storage.
[0027] Some embodiments can include at least one processor. The processor can be any physical device or group of devices having an electronic circuit that performs logical operations on inputs. For example, the at least one processor can include all or a portion of an application-specific integrated circuit (ASCI), microchip, microcontroller, microprocessor, central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), field programmable gate array (FPGA), server, virtual server, or one or more integrated circuits (ICs) including other circuits suitable for executing instructions or performing logical operations. The instructions executed by the at least one processor can be pre-loaded, for example, into memory integrated with or included in the controller, or stored in a separate memory.
[0028] In some embodiments, at least one processor can include two or more processors. Each processor can have a similar configuration, or the processors can have different configurations that are electrically connected or disconnected from each other. For example, the processors can be discrete circuits or integrated into a single circuit. When two or more processors are used, the processors can be configured to operate independently or in cooperation. The processors can be coupled by electrical, magnetic, optical, acoustic, mechanical, or other means that allow the processors to interact.
[0029] Embodiments consistent with the present disclosure can include a network. The network can comprise any type of physical or wireless computer network arrangement used to exchange data. For example, the network can be a public network, a Wi-Fi network, a LAN or WAN network, and / or the Internet, a private data network, a virtual private network using other suitable connections that can enable information exchange between various components of the system. In some embodiments, the network can include one or more physical links used to exchange data, such as Ethernet, coaxial cable, twisted pair cable, fiber optic, or any other suitable physical medium for exchanging data. The network can also include one or more networks such as a private network, a public switched telephone network ("PSTN"), the Internet, and / or a wireless cellular network. The network can be a secure network or an unsecured network. In other embodiments, one or more components of the system can communicate directly through a dedicated communication network. Direct communication can be, for example, BLUETOOTH TM (Bluetooth) (registered trademark), Bluetooth Low Energy TMAny suitable technology can be used, including, for example, (BLE), Wi-Fi, Near Field Communication (NFC), or any other suitable communication method that provides a medium for the exchange of data and / or information between separate entities.
[0030] In some embodiments, for example, in the cases described below, a machine learning network or algorithm can be trained using training examples. Non-limiting examples of such machine learning algorithms include classification algorithms, video classification algorithms, data regression algorithms, image segmentation algorithms, temporal video segmentation algorithms, visual detection algorithms (such as object detectors, face detectors, human detectors, motion detectors, edge detectors, etc.), visual recognition algorithms (such as face recognition, human recognition, object recognition, etc.), speech recognition algorithms, action recognition algorithms, mathematical embedding algorithms, natural language processing algorithms, support vector machines, random forests, nearest neighbor algorithms, deep learning algorithms, artificial neural network algorithms, convolutional neural network algorithms, recurrent neural network algorithms, linear machine learning models, non-linear machine learning models, ensemble algorithms, and the like. For example, a trained machine learning network or algorithm can comprise an interface such as a prediction model, a classification model, a regression model, a clustering model, a segmentation model, an artificial neural network (such as a deep neural network, a convolutional neural network, a recurrent neural network, etc.), a random forest, a support vector machine, and the like. In some examples, the training examples can include an example input together with the desired output corresponding to this example input. In still further examples, training a machine learning algorithm using the training examples can generate a trained machine learning algorithm, and the trained machine learning algorithm can be used to estimate the output for inputs not included in the training examples. The training may or may not be with a supervisor (teacher, supervisor, etc.), and may be a combination thereof. In some examples, engineers, scientists, processes, and machines that train machine learning algorithms can further use validation examples and / or test examples.For example, a validation example and / or a test example can include an example input together with a desired output corresponding to this example input, and a trained machine learning algorithm and / or a moderately trained machine learning algorithm can be used to estimate an output for the example input of the validation example and / or the test example, the estimated output can be compared with the corresponding desired output, and the trained machine learning algorithm and / or the moderately trained machine learning algorithm can be evaluated based on the result of the comparison. In some examples, the machine learning algorithm can have parameters and hyperparameters, the hyperparameters can be set manually by a human or automatically by a process external to the machine learning algorithm (such as, for example, a hyperparameter search algorithm), and the parameters of the machine learning algorithm are set by the machine learning algorithm according to training examples. In some embodiments, the hyperparameters are set according to training examples and validation examples, and the parameters are set according to training examples and the selected hyperparameters. The machine learning network or algorithm can be further retrained based on any output.
[0031] Certain embodiments disclosed herein can include a computer-implemented system for performing a method comprising operations or a series of steps. The computer-implemented system and method can be implemented by one or more computing devices, which can include one or more processors as described herein configured to process real-time video. The computing device can be one or more computers or any other device capable of processing data. Such a computing device can include a display such as an LCD display, an augmented reality (AR) or virtual reality (VR) display. However, the computing device can also be implemented in a computing system that includes backend components (such as, for example, a data server), or middleware components (such as, for example, an application server), or frontend components (such as, for example, a graphical user interface or a web browser through which a user can interact with an implementation of the system, and a user device having the technology described herein), or any combination of such backend, middleware, or frontend components. The components of the system and / or the computing device can be interconnected by any form or medium of digital data communication (such as, for example, a communication network). Examples of communication networks include local area networks (“LAN”), wide area networks (“WAN”), and the Internet. The computing device can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship between a client and a server arises from computer programs operating on respective computers and having a client-server relationship with each other.
[0032] FIG. 1A is a block diagram of an exemplary intelligent detector system 100 that is consistent with embodiments of the present disclosure. As further disclosed herein, the intelligent detector system 100 may be a computer-implemented system and may include one or more convolutional neural networks (CNNs) that process images from medical procedures to identify features of a requested object of interest in the image. The features of the object of interest may be a medical condition or a list of medical conditions that a physician is looking for in the image (e.g., to diagnose a patient). As a further example, the features of the object of interest may also include actions taken on or in the image, anatomical locations or other locations of interest in the image, clinical metric levels of the image, and the like. These, and other examples, are within the scope of the present disclosure. As an example, the actions may include actions taken by a physician during a medical procedure or as part of a subsequent procedure, and may include actions or procedures identified by the system 100 as a result of a spatio-temporal examination of images from the medical procedure. For example, the actions may include recommended actions or procedures that follow medical guidelines, such as performing a biopsy, removing a lesion, or exploring / analyzing the surface / mucosa of an organ. The actions or procedures can be identified based on images captured and processed by the intelligent detector system 100.
[0033] The intelligent detector system 100 can receive, as input, a set of temporally ordered images of a medical procedure, such as an endoscopy or a colonoscopy. The intelligent detector system 100 can output a report or information that includes one or more numerical values (e.g., scores) for each image. The numerical values can relate to a medical category, such as a particular medical condition, and can provide information regarding the probability of the presence of the medical category within the image frame. The images processed by the intelligent detector system 100 can be images captured during a medical procedure and can be images stored in a database or memory device for subsequent retrieval and processing by the intelligent detector system 100. In some embodiments, the output provided by the intelligent detector system 100 and resulting from processing the images can include, for example, a report having numerical values assigned to the images and recommended next steps according to medical guidelines. The report can be generated after the completion of an endoscopy or other medical procedure. The report can include information regarding all features of interest identified in the processed images. Further, in some embodiments, the output provided by the intelligent detector system 100 can include recommended actions (e.g., performing a biopsy, removing a lesion, exploring and / or analyzing the surface / mucosa of an organ, etc.) to be performed by a physician in view of the features of interest identified in the images from the medical procedure. During the medical procedure, the intelligent detector system 100 can directly receive video or image frames from a medical imaging device, process the video or image frames, and provide feedback to the operator regarding the actions performed by the operator, along with a final report that includes details of multiple measured variations, clinical metrics, and observations, and / or at any anatomical location, and / or what actions the operator performed during the medical procedure (e.g., within a short time interval of 0 to several minutes). The actions performed can include recommended actions or procedures according to medical guidelines, such as performing a biopsy, removing a lesion, or exploring / analyzing the surface / mucosa of an organ.In some embodiments, the recommended actions may be part of a set of recommended actions based on medical guidelines. A detailed description of an example computer system implementing the intelligent detector system 100 for real-time processing is presented in the description of FIG. 1B below.
[0034] As described herein, the intelligent detector system 100 can generate a report after the completion of a medical procedure, and this report includes information based on the processing of the imaging video by the local spatio-temporal processing module 110, the global spatio-temporal processing module 120, and the time-series analysis module 130. The report can include information regarding the characteristics of the objects of interest identified during the medical procedure, along with other information such as numerical values or scores for each image. As explained, the numerical values can relate to medical categories such as specific medical conditions and can provide information regarding the probability of the presence of a medical category within the image frame. Further details regarding the intelligent detector system 100, and the operations of the local spatio-temporal processing module 110, the global spatio-temporal processing module 120, and the time-series analysis module 130 are provided below with reference to the accompanying drawings.
[0035] In some embodiments, the report generated by the system 100 can include additional recommended actions based on the processing of the stored images by the medical procedure, or the real-time processing of the images by the medical procedure. The additional recommended actions can include actions or procedures that could have been performed during the medical procedure, and actions or procedures to be performed after the medical procedure. The additional recommended actions may be part of a set of recommended actions based on medical guidelines. Further, as described above, the system 100 can process the video in real-time to provide the operator with simultaneous feedback regarding what is happening in the video and what has been identified during the medical procedure.
[0036] The output generated by the intelligent detector system 100 can include a dashboard display or a similar report (see, e.g., FIG. 10). The dashboard can provide a summary report of a medical procedure, such as an endoscopy or a colonoscopy. The dashboard can provide a quality score and / or other information regarding the procedure. The score and / or other information can summarize the inspection actions of healthcare professionals and provide information regarding features of identified objects of interest, such as the number of identified polyps. In some embodiments, the information generated by system 100 is provided and displayed as an overlay on the video from the medical procedure, such that the operator can view the information as part of an extended video feed during or immediately after the medical procedure. This information can be provided with some delay or without delay.
[0037] The intelligent detector system 100 can also generate reports in the form of electronic files, data sets, or data transmissions. As an example, the output generated by the system 100 can conform to a standard format and / or can be integrated into medical records such as an electronic health record (EHR). The output of the system 100 can also comply with regulations such as HIPAA (a U.S. law regarding the interoperability and accountability of medical insurance) regarding interoperability and privacy. In some embodiments, the output can be integrated into other medical records. For example, the output of the intelligent detector system 100 can be integrated into an electronic medical record or electronic health record for a patient. The intelligent detector system 100 can include an API to facilitate such integration and / or can provide output in the form of a standardized data set or template. The standardized template can include a predefined format or table that can be filled with data values generated by the intelligent detector system 100 by processing input videos or image frames from medical procedures. In some embodiments, the report can be generated by the system 100 in a machine-readable format such as an XML file to support its transmission or storage along with integration with other systems and applications. In some embodiments, the report can be provided in other formats such as Word, Excel, HTML, or PDF file formats. In some embodiments, the intelligent detector system 100 can upload data or reports to a server or database via a network (e.g., refer to network 160 in FIG. 1A). The intelligent detector system 100 can also transfer output data or a formatted report to a server or database by making an API call and sending it, for example, as a JSON document.
[0038] FIG. 10 shows, by way of example, a dashboard 1000 having an output summary 1090 generated using an intelligent detector system (such as intelligent detector system 100) that is consistent with an embodiment of the present disclosure. The modules of intelligent detector system 100 can be used to generate an output summary 1090 for a procedure. The output summary 1090 can provide a quality score and / or other information such as procedure time, withdrawal time, and clean withdrawal time. Further examples of information that can be part of summary 1090 include the time spent performing a particular action (such as the recommended actions discussed above) and the time spent at different anatomical locations. Information regarding features of interest such as polyps can also be provided. For example, the time series analysis module 130 can generate a summary of the number of polyps identified based on characteristics observed in the image frames. The time series analysis module 130 can aggregate information generated by processing the images of the input video using the local and global spatio-temporal processing modules 110 and 120. The summary dashboard 1000 can also include a visual depiction of the features of interest identified by the intelligent detector system 100. For example, selected frames of a procedure video can be augmented with marks such as green bounding boxes for the respective locations of the identified features of interest, as shown in frames 1092, 1094, and 1096. The frames can relate to different examined portions of the colon that can themselves be features of interest requested by the user of the intelligent detector system 100, such as the ileocecal valve, small pores, and tri-radiate folds. The example of FIG. 10 illustrates information for a procedure displayed as part of a single dashboard, but multiple dashboards can be generated having output summaries for each part of the colon or other organ examined as part of a medical procedure. In some embodiments, a combined score or value is generated based on inputs received as a plurality of vectors (e.g., image score vectors) generated by the local and global spatio-temporal processing modules 110 and 120.
[0039] As disclosed above, the features of interest can relate to medical categories or medical conditions. The intelligent detector system 100 can be implemented to handle requests to detect one or more medical categories (i.e., one or more features of interest) in an image. In the case of multiple features of interest, one instance of a component of the intelligent detector system 100 can be implemented for each medical category or feature of interest. As will be appreciated from the present disclosure, instances of the intelligent detector system 100 can be implemented in any combination of hardware, firmware, and software, depending on the requirements for speed or throughput, the amount of image data to be processed, and other requirements of the system.
[0040] In some embodiments, a single instance of the intelligent detector system 100 can output a plurality of numerical values for each image with respect to each medical category. In one exemplary embodiment, the medical condition detected by the intelligent detector system 100 can include detecting polyps in the colonic mucosa. As a further example, the intelligent detector system 100 can output a numerical value (e.g., 0) for all images of the input image when no polyps are detected by the intelligent detector system 100, and can output another numerical value (e.g., 1) for all images of the input image when the intelligent detector system detects at least one polyp. In some embodiments, the numerical values can be arranged with respect to a range or scale, and / or can indicate the probability of the presence of a polyp or other feature of interest.
[0041] The source of the input image can vary according to the needs of the imaging device, memory device, and / or application. For example, in accordance with the embodiments disclosed herein, the intelligent detector system 100 can be configured to process a direct video feed from a video endoscope inspection device and receive temporally ordered input images that are subsequently processed by the system. As a further example, the intelligent detector system 100 can be configured to receive input images from a database or memory device, where the stored images are temporally ordered and were previously captured using an imaging device such as a video camera of an endoscope inspection device or a camera of a capsule device. The images received by the intelligent detector system 100 can be processed and analyzed to identify features of one or more objects of interest, such as one or more types of polyps or lesions.
[0042] The system as an example of FIG. 1A can be implemented in various environments and for various applications. For example, the captured input image can be stored in a local database or memory device, or can be accessed and received by the intelligent detector system 100 via a network from a remote storage location such as cloud storage. The intelligent detector system 100 can also be configured to process a streaming video feed from a current medical procedure and process the input images collected from the feed (e.g., via preprocessing and buffering). Further, the operation of the intelligent detector system 100 can be programmed or induced to start based on one or more conditions. For example, the intelligent detector system 100 can be configured to directly analyze the input image when it receives the input image (e.g., via a video feed or a set of stored input images from memory), or when it receives a command from a user. The output of the intelligent detector system 100 can also be configured as desired. For example, as discussed above, the intelligent detector system 100 can analyze the input image for features of one or more objects of interest and generate a report indicating the presence of the features of one or more objects of interest in the processed image. The report can take the form of an electronic file, a graphical display, and / or an electronic transmission of data. As will be appreciated, other outputs and report formats are within the scope of the present disclosure. In some embodiments, reports in different formats can be preconfigured and used as templates for generating reports by filling in the templates with values generated by the intelligent detector system 100. In some embodiments, the report is formatted to be integrated into other reporting systems such as an electronic medical record (EMR). The report format can be a machine-readable format such as XML or Excel for integration with other reporting systems.
[0043] As an example, the intelligent detector system 100 can process recorded video or images and provide a fully automated report and / or other output indicating details of features of the object of interest observed in the processed image. The intelligent detector system 100 can use artificial intelligence or machine learning components to efficiently and accurately process the input image and determine the presence of features of the object of interest based on image analysis and / or spatio-temporal information. Further, for each feature of the object of interest being requested or investigated, the intelligent detector system 100 can estimate its presence in the image and provide a report or other output with information indicating the likelihood of the presence of that feature and other details such as the relative time from the start of treatment or the sequence of images when the feature of the object of interest appeared, the estimated anatomical location, duration, most significant images, the location within these images, and / or the number of occurrences.
[0044] In one embodiment, the intelligent detector system 100 can be configured to automatically determine the presence of gastrointestinal conditions without the assistance of a physician. As discussed above, the input images can be captured and received using different methods and different types of imaging devices. For example, a video endoscope device or a capsule device or other medical device or other imaging device can record and provide the input images. The input images can be part of a live video feed or can be part of a stored set of images received from a local or remote storage location (e.g., a local database or cloud storage). The intelligent detector system 100 can be operated as part of a procedure or medical service in a clinic or hospital or can be provided as an online or cloud medical service to an end user to enable self-diagnosis or remote testing.
[0045] As an example, to initiate an examination procedure, a user can ingest a capsule device or a pill camera. The capsule device can include an imaging device and can wirelessly transmit images of the user's gastrointestinal tract during the procedure to a smartphone, tablet, laptop, computer, or other device (e.g., user device 170). The captured images can then be uploaded via a network connection to a database, cloud storage, or other storage device (e.g., image source 150). The intelligent detector system 100 can receive input images from the image source and analyze the images for one or more requested features of interest (e.g., polyps or lesions). The final report can then be electronically provided as an output to the user and / or physician. The report can include scoring or probability metrics for each observed feature of interest, and / or other relevant information or medical recommendations. Additionally, the intelligent detector system 100 can detect pathophysiological characteristics that are related to the features of interest and are indicators of the features of interest, and can score those characteristics that are determined to be present. Examples of such characteristics include bleeding, inflammation, ulcers, neoplastic tissue, and the like. Further, in response to the detected features of interest, the report can include information or recommendations based on medical guidelines, such as recommendations to consult a physician and / or undergo additional diagnostic tests. One or more actions can also be recommended to the physician in real-time during the medical procedure or after the medical procedure is completed, based on the analysis of the images by the intelligent detector system 100 (e.g., performing a biopsy, removing a lesion, exploring / analyzing the surface / mucosa of an organ, etc.).
[0046] As another example, the intelligent detector system 100 can assist a physician or an expert in analyzing video content recorded during a medical procedure or an examination. The captured images can be, for example, part of the video content recorded during a gastroscopy, a colonoscopy, or a small intestine endoscopy procedure. Based on the analysis performed by the intelligent detector system 100, the entire video recording can be presented to the physician or the expert together with a colored timeline bar where different colors correspond to different features of interest and / or scores for the identified features of interest.
[0047] As a further example, a physician, an expert, or another individual can use the intelligent detector system 100 to create an overview of a video recording or a set of images by focusing on images having the desired features of interest and discarding irrelevant image frames. The intelligent detector system 100 can be configured to allow a physician or a user to adjust or select the features of interest for detection and the duration of each overview based on total duration and / or other parameters such as a preset lead and lag time before and after a sequence of frames having the selected features of interest. The intelligent detector system 100 can also be configured to combine all or the most relevant frames according to the requested features of interest.
[0048] As illustrated in FIG. 1A, the intelligent detector system 100 can include a local spatio-temporal processing module 110, a global spatio-temporal processing module 120, a time series analysis module 130, and a task manager 140. These components can be implemented by any suitable combination of hardware, software, and / or firmware. Furthermore, the number and arrangement of these components can be modified, and it will be recognized that the embodiment as an example in FIG. 1A is provided for illustrative purposes and does not limit the scope of the invention and its embodiments. Further, example features and details regarding these components are provided below and include reference to FIGS. 1B and 3A - 3C.
[0049] Referring again to the embodiment as an example of FIG. 1A, the local spatio-temporal processing module 110 can be configured to provide a local perspective by processing a subset of the images of the input video or a set of input images. The local spatio-temporal processing module 110 can select a subset of the images and process the images to generate a score based on the determined presence of characteristics regarding one or more features of interest. For example, assume that the endoscopic input video V includes a set of T image frames. The characteristics can define the features of interest requested by the user of the intelligent detector system 100. For example, the characteristics can include physical and / or biological aspects such as the size, orientation, color, shape, etc. of the features of interest. The characteristics can also include metadata such as data that identifies a portion of the video or a period in the video. For example, the characteristics of a colonoscopy procedure video can identify portions of the colon such as ascending, transverse, descending, etc. In other examples, the characteristics can relate to one or more portions of the endoscopic procedure video such as the amount of movement in the image, the presence of an instrument, or the duration of a portion with reduced movement. The characteristics that define the content can also indicate the actions of a physician, clinician, or other individual performing a medical procedure. For example, it can be a portion of the video with the longest pause without movement, or a portion of the video with the maximum time for exploring the surface of an organ. In some embodiments, the characteristics can be the features of interest. For example, the features of interest and characteristics of a colonoscopy procedure video can be portions of the colon such as ascending, transverse, descending, etc.
[0050] The local spatio-temporal processing module 110 can be configured to process the entire input video by repeating it for successive batches or a subset of N image frames. The local spatio-temporal processing module 110 can also be configured to provide an output that includes a vector or quality score representing the determined characteristics of the features of interest in each image frame. In some embodiments, the local spatio-temporal processing module 110 can be configured to output a quality value and a segmentation map associated with each image frame. Further details as an example regarding the local spatio-temporal processing module 110 are provided below with reference to the embodiment of FIG. 3A.
[0051] The subset of images processed by the local spatio-temporal processing module 110 can include shared or overlapping images. Further, the size or arrangement of the subset of images can be defined or controlled based on one or more factors. For example, the size or volume of the subset of images can be configurable by a physician or other user of the system. As a further example, the local spatio-temporal processing module 110 can be configured such that the size or amount of the subset of images is dynamically determined based on the required features of interest. Additionally, or alternatively, the size of the subset of images can be dynamically determined based on the determined characteristics regarding the required features of interest.
[0052] The global spatio-temporal processing module 120 can be configured to provide a comprehensive perspective by processing all of a subset of the images analyzed by the local spatio-temporal processing module 110. For example, the global spatio-temporal processing module 120 can process the entire set of input videos or input images by processing all of the outputs of the local spatio-temporal processing module 110 at once or together. Further, the global spatio-temporal processing module 120 can be configured to provide an output that includes a numerical score for each image frame by processing a vector of determined characteristics regarding the features of interest. In some embodiments, the global spatio-temporal processing module 120 can process images and vectors and output a quality score and a segmentation map with improvements added to each image. Further details as an example regarding the global spatio-temporal processing module 120 are provided below with reference to the embodiment of FIG. 3B.
[0053] The time series analysis module 130 uses information about the images determined by the local spatio-temporal processing module 110 and improved by the global spatio-temporal processing module 120 to output a numerical score indicating the presence of one or more features of interest requested by the user of the intelligent detector system 100. For example, the time series analysis module 130 can be configured to use spatio-temporal information about the characteristics regarding the features of interest determined by the local spatio-temporal processing module 110 to perform a time series analysis on the input video or image. Further details as an example regarding the time series analysis module 130 are provided below with reference to the embodiment of FIG. 3C.
[0054] The task manager 140 can assist in managing various tasks requested by the user of the intelligent detector system 100. The tasks can relate to the features of the object of interest requested or solicited and / or the characteristics of the features of the object of interest. One or more of the characteristics and the features of the object of interest can be part of each task for processing by the intelligent detector system 100. The task manager 140 can assist in managing tasks for the detection of a plurality of features of interest in a set of input images. The task manager 140 can determine the number of instances of the components of the intelligent detector system 100 (e.g., the local spatio-temporal processing module 110, the global spatio-temporal processing module 120, and the time series analysis module 130). Further details as an example of a method for handling multiple task requests for detecting features of interest are provided below with reference to the descriptions of FIGS. 6 and 7 below.
[0055] The intelligent detector system 100 can receive a set of input videos or images from the image source 150 via the network 160 for processing. In some embodiments, the intelligent detector system 100 can directly receive the input video from other systems such as, for example, medical devices or systems used to image videos when performing medical procedures, such as colonoscopy. After processing the images, a report of the detected features of the object of interest can be shared via the network 160. As disclosed herein, the report can be transmitted electronically and can take different forms such as an electronic file, a display, or data. In some embodiments, the report is sent to the user device 170 as a file and / or is displayed on the user device 170. The network 160 can take various forms depending on the needs and environment of the system. For example, the network 160 can include the Internet, a wired wide area network (WAN), a wired local area network (LAN), a wireless WAN (e.g., WiMAX), a wireless LAN (e.g., IEEE802.11, etc.), a mesh network, a mobile / cellular network, a corporate or private data network, a storage area network, a virtual private network using a public network, and / or any combination of other types of network communications, or can be utilized. In some embodiments, the network 160 can include an in-facility (e.g., LAN) network, while in other embodiments, the network 160 can be a hybrid in-facility and virtualized, remote and / or cloud network (e.g., AWS TM , Azure TM , IBM Cloud TM , etc.) that can include. Further, the network 160 can, in some embodiments, be a hybrid in-facility and virtualized, remote and / or cloud network that includes one or more types of components of the network architecture.
[0056] The user device 170 can send requests to the intelligent detector system 100 regarding features of objects of interest in input videos or images, and can receive outputs (e.g., reports or data) from the intelligent detector system 100. The user device 170 can control the input video or image for processing, including by means of instructions, commands, video or image set file downloads, and / or storage links to storage locations (e.g., image source 150), and can directly provide it to the intelligent detector system 100. The user device 170 can comprise a smartphone, laptop, tablet, computer, and / or other computing devices. The user device 170 can also include an imaging device (e.g., a video or digital camera) for capturing videos or images for processing. In the case of a capsule inspection procedure, for example, the user device 170 includes a pill camera or the like that is ingested by the user, an input video or image is captured, directly streamed to the intelligent detector system 100, or stored in the image source 150, and then downloaded and received by the system 100 via the network 160. And the results of the image processing are provided as outputs from the intelligent detector system 100 to the user device 170 via the network 160.
[0057] The medical device 180 can also send requests regarding features of interest in the input video or image to the intelligent detector system 100 and receive outputs (e.g., reports or data) from the intelligent detector system 100. Similar to the user device 170, the medical device 180 can control the input video or image for processing, or provide it directly to the intelligent detector system 100, including by commands, commands, video or image set file downloads, and / or storage links to storage locations (e.g., the image source 150). The medical device 180 can comprise a smartphone, laptop, tablet, computer, and / or other computing devices. The medical device 180 can also include an imaging device (e.g., a video or digital camera) for imaging a video or image for processing. In the case of video endoscopy, for example, the medical device 180 can include a colonoscopy probe or the like having an imaging device that images the patient during the examination. The imaged video can be streamed as input video to the intelligent detector system 100 or stored in the image source 150 and then downloaded and received by the system 100 via the network 160. In some embodiments, the medical device 180 can receive notifications for further scrutiny of image frames having characteristics of interest. And the results of the image processing are provided as outputs (e.g., data in the form of an electronic report or file or digital display) from the intelligent detector system 100 to the user device 170 via the network 160.
[0058] The image source 150 can include a storage location or other source for input video or images to the intelligent detector system 100. The image source 150 can comprise any suitable combination of hardware, software, and firmware. For example, the image source 150 can include any combination of computing devices, servers, databases, memory devices, network communication hardware, and / or other devices. As an example, the image source 150 can include a database, memory, or storage (e.g., the storage 220 of FIG. 2) to store a set of input videos or images received from the user device 170 and the physician device 180. The image source 150 storage can include file storage and / or a database accessed using a CPU (e.g., the processor 230 of FIG. 2). As a further example, the image source 150 can also be AMAZON TM S3, Azure TM Storage, GOOGLE TM It can include cloud storage accessible via the network 160 such as Cloud Storage and the like.
[0059] In the system as an example of FIG. 1A, the image source 150, the user device 170, and the physician device 180 may be local or remote from each other and can communicate with each other via wired or wireless communication, including via network communication. The devices may also be local or remote to the intelligent detector system 100 depending on the needs of the application and system implementation form. Further, although the image source 150, the user device 170, and the physician device 180 are shown as separate from the intelligent detector system 100 in FIG. 1A, one or more of these devices may be local or provided as part of the system 100. Also, a portion or part of the network 160 may be local to the system 100 or part of the system 100. Further, it will be recognized that the number and arrangement of components and devices in FIG. 1A are provided for illustrative purposes and are not intended to limit the invention or its disclosed embodiments.
[0060] Embodiments of the present disclosure are described herein with reference to medical image analysis and endoscopy generally, but it will be recognized that the embodiments are applicable to other medical imaging procedures such as gastroscopy, colonoscopy, and small bowel endoscopy. Further, embodiments of the present disclosure can also be realized for other image capture and analysis environments and systems such as those for or including LIDAR, surveillance, autonomous driving, or other imaging systems.
[0061] According to aspects of the present disclosure, a computer-implemented system is provided for intelligently processing a set of input videos or images to determine the presence of features of interest and characteristics thereof. As further disclosed herein, the system (e.g., intelligent detector system 100) can include at least one memory (e.g., ROM, RAM, local memory, network memory, etc.) configured to store instructions, and at least one processor (e.g., processor 230) configured to execute the instructions (e.g., see FIGS. 1 and 2). Using the at least one processor, the system can process a set of input videos or images captured by a medical imaging system such as used during an endoscopic examination, gastroscopy, colonoscopy, or small bowel endoscopy procedure. Additionally, or alternatively, the image frames can comprise medical images such as images of gastrointestinal organs or other organs or regions of human tissue.
[0062] As used herein, the term "image frame" or "image" refers to any digital representation of a scene or field of view captured by an imaging device. The digital representation can be encoded in any suitable format such as the Joint Photographic Experts Group (JPEG) format, Graphics Interchange Format (GIF), bitmap format, Scalable Vector Graphics (SVG) format, Encapsulated PostScript (EPS) format, and the like. Similarly, the term "video" refers to any digital representation of a scene or region of interest composed of a plurality of consecutive images. The digital representation of video can be encoded in any suitable format such as the Moving Picture Experts Group (MPEG) format, flash video format, Audio Video Interleave (AVI) format, and the like. In some embodiments, the sequence of images for an input video can be paired with audio. As will be appreciated from the present disclosure, embodiments of the invention are not limited to processing an input video having image frames arranged in sequence or temporally ordered, but can also process a set of images captured continuously or temporally ordered, either streamed or stored. Accordingly, the terms "input video" and "set of images" should be considered interchangeable and should not limit the scope of the present disclosure.
[0063] As disclosed herein, an image frame or image can include a representation of a feature of interest (i.e., an anomaly or other object of interest). For example, the feature of interest can include or be an anomaly on or of human tissue. In other embodiments for non-medical applications, the feature of interest can include an object such as a vehicle, a person, or other entity.
[0064] According to the present disclosure, an “abnormality” can include on a human tissue, or a change in a human tissue in its formation, from one type of cell to another type of cell in the human tissue, and / or the absence of a human tissue in a location where the human tissue is expected to be present. For example, the growth of a tumor or other tissue may be abnormal because there are more cells than expected. Similarly, a bruise or other change in cell type may be abnormal because blood cells are present outside of where they are expected (i.e., outside of capillaries). Similarly, a depression in a human tissue may be abnormal because cells are not present where they are expected, resulting in the depression.
[0065] In some embodiments, the abnormality can include a lesion. The lesion can include a lesion of the gastrointestinal mucosa. The lesion can be histologically classified (e.g., according to Narrow-Band Imaging International Colorectal Endoscopic (NICE), or the Vienna classification), morphologically classified (e.g., the Paris classification), and / or structurally classified (e.g., serrated or non-serrated, etc.). The Paris classification includes polypoid and non-polypoid lesions. Polypoid lesions can include elevated, pedunculated and elevated, or sessile lesions. Non-polypoid lesions can include superficially elevated, flat, superficially depressed, or excavated lesions. With regard to detecting an abnormality as a feature of interest, serrated lesions can include sessile serrated adenomas (SSA), traditional serrated adenomas (TSA), hyperplastic polyps (HP), fibroblastic polyps (FP), or mixed polyps (MP). According to the NICE classification system, the abnormality is divided into three types as follows: (type 1) sessile serrated polyp or hyperplastic polyp, (type 2) conventional adenoma, and (type 3) cancer with deep submucosal invasion. According to the Vienna classification, the abnormality is classified into five categories as follows: (category 1) negative for neoplasia / dysplasia, (category 2) occult for neoplasia / dysplasia, (category 3) non-invasive low-grade neoplasia (low-grade adenoma / dysplasia), (category 4) high-grade adenoma / dysplasia, mucosal high-grade neoplasia such as non-invasive carcinoma (carcinoma in situ), or suspected invasive carcinoma, and (category 5) invasive neoplasia, intramucosal carcinoma, submucosal carcinoma, etc. These examples and other types of abnormalities are within the present disclosure. It will be recognized that the intelligent detector system 100 can be configured to detect other types of features of interest, including those for medical or non-medical procedures.
[0066] Figure 1B is a schematic representation of an example computer-implemented system that implements the intelligent detector system 100 of Figure 1A for processing real-time video, consistent with an embodiment of the present disclosure. As shown in Figure 1B, system 190 includes an imaging device 192 and an operator 191 that operates and controls the imaging device 192 with control signals sent from the operator 191 to the imaging device 192. As an example, in embodiments where the video feed includes medical video, the operator 191 may be a physician or other healthcare professional. The imaging device 192 can comprise a medical imaging device such as an endoscopy imaging device or other medical imaging device that generates video or one or more images of a human body / tissue / organ, or a portion thereof. The imaging device 192 may be part of the physician device 180 (as shown in Figure 1A) and generates video stored in the image source 150. The operator 191 can control the imaging device 192, for example, by passing through or relative to the body of a patient or individual, and in particular, by controlling the imaging rate or frame rate of the imaging device 192, and / or the movement or navigation of the imaging device 192. In some embodiments, the imaging device 192 can comprise an ingestible capsule device or other form of capsule endoscopy device, as opposed to an endoscopy imaging device inserted through a body cavity of a human body.
[0067] In the example of FIG. 1B, the imaging device 192 can directly transmit the captured video as a plurality of image frames to the computing device 193. The computing device 193 can include a memory (including one or more buffers) and one or more processors for processing video or images, as described above and herein (e.g., see FIG. 2). In some embodiments, one or more of the processors can be implemented not as part of the computing device 193 but as separate components (not shown) that communicate with it over a network (e.g., network 160 of FIG. 1A). In some embodiments, one or more processors of the computing device 193 can implement one or more networks such as a trained neural network. Examples of neural networks include an object detection network, a classification detection network, a position detection network, a size detection network, or a frame quality detection network, as further described herein. The computing device 193 can directly receive and process a plurality of image frames from the imaging device 192. In some embodiments, the computing device 193 can use preprocessing and / or parallel processing and buffering to process video or images in real time, and the levels of such processing and buffering depend on the frame rate of the received video or images and the processing speed of one or more processors or modules of the computing device 193. As recognized, well-matched processing and buffering functions with respect to the frame rate enable real-time processing and output. Further, in some embodiments, control or information signals can be exchanged between the computing device 193 and the operator 191 for the purpose of controlling or instructing the creation of one or more extended videos as output, and the extended videos include the original video with an overlay (graphics, symbols, text, etc.) added that provides information about the identified features of the object of interest and other feedback generated by the computing device 193 to assist the physician or operator performing the medical procedure. With respect to the exchanged control or information signals, they can be communicated as data directly from the imaging device 192 or from the operator 191 to the computing device 193.Examples of control and information signals include signals for controlling components of the arithmetic unit 193, such as an object detection network, a classification detection network, a position detection network, a size detection network, or a frame quality detection network as described herein.
[0068] In the example of FIG. 1B, the arithmetic unit 193 can use one or more modules (such as modules 110-140 of the intelligent detector system 100) to process and expand the video received from the imaging device 192 and transmit the expanded video to the display device 194. The expanded video can provide, for example, real-time feedback and reports of the identified polyps and the actions performed by the operator 191 during or at the end of a medical procedure such as an endoscopy or colonoscopy. Video expansion or modification can comprise providing one or more overlays, alphanumeric characters, text, descriptions, shapes, drawings, images, moving images, and / or other suitable graphical representations in or having the video frame. The video expansion can provide information regarding features of interest such as classification, size, actions performed, and / or position information. Additionally, or alternatively, the video expansion can provide information regarding one or more recommended actions identified by the arithmetic unit 193 in accordance with medical guidelines. It will be recognized that the scope and type of information, reports, and data generated by the arithmetic unit 193 may be similar to those described above for the intelligent detector system 100 to assist the physician or operator and reduce errors. Accordingly, reference is made to the examples provided above for the system 100.
[0069] As further shown in FIG. 1B, the computing device 193 can also be configured to directly relay and send the original unextended video from the imaging device 192 to the display device 194. For example, the computing device 193 can perform a direct relay under certain conditions such as when there is no overlay or other extension being generated, or when the imaging device 192 is off. In some embodiments, the computing device 193 can perform a direct relay if the operator 191 sends a command to the computing device 193 to perform a direct relay as part of a control signal. Commands from the operator 191 can be generated by operations of buttons and / or keys included on the operator device and / or an input device (not shown), such as a mouse click, cursor hover, mouse over, button press, keyboard input, voice command, interactions performed in virtual or augmented reality, or any other input.
[0070] To extend the video, the computing device 193 can process the video from the imaging device 192 and create a modified video stream for sending to the display device 194. The modified video can comprise the original image frame with extension information that is displayed to the operator via the display device 194. The display device 194 can comprise any suitable display or similar hardware for displaying video or modified video, such as an LCD, LED, or OLED display, an augmented reality display, or a virtual reality display.
[0071] FIG. 2 shows an example computing device 200 that can be employed in connection with implementing the system as an example of FIG. 1A and other embodiments of the present disclosure. The computing device 200 can be used in connection with an implementation of one or more components of the system as an example of FIG. 1A (e.g., system 100 and devices 150, 170, and 180). In some embodiments, the computing device 200 can include multiple subsystems such as a cloud computing system, a server, and / or any other suitable components for receiving and processing input video and images.
[0072] As shown in FIG. 2, the computing device 200 can include one or more processors 230 including one or more integrated circuits (ICs) that can include, for example, as described above, all or part of an application specific integrated circuit (ASIC), a microchip, a microcontroller, a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a server, a virtual server, or other circuitry suitable for executing instructions or performing logical operations. In some embodiments, the processor 230 can include a larger processing unit implemented by one or more processors or can be a component of that processing unit. The one or more processors 230 can be implemented by any combination of a general-purpose microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic device (PLD), a controller, a state machine, gate logic, discrete hardware components, a dedicated hardware finite state machine, or any other suitable entity capable of performing calculations or other operations on data or information.
[0073] As further shown in FIG. 2, the processor 230 can be connected to communicate with the memory 240 via a bus or network 250. The bus or network 250 can be adapted to communicate other forms of data and information. The memory 240 can include a memory portion 245 that includes instructions that, when executed by the processor 230, perform the operations and methods described in detail herein. The memory 240 can also serve as a working memory, temporary storage, and in some cases other memory or storage for the processor 230. By way of example, the memory 240 can be, without limitation, a volatile memory such as random access memory (RAM) or, without limitation, a non-volatile memory (NVM) such as flash memory.
[0074] The processor 230 can also be connected to communicate with one or more I / O devices 210 via a bus or network 250. The I / O device 210 can include any type of input and / or output device or peripheral device, including a keyboard, mouse, display device, and the like. The I / O device 210 can include one or more network interface cards, APIs, data ports, and / or other components for supporting connectivity with the processor 230 via the network 250.
[0075] As further shown in FIG. 2, the processor 230 and other components (210, 240) of the arithmetic unit 200 can be connected to communicate with a database or storage 220. The storage 220 can electronically store data (e.g., a set of input videos or images, along with reports and other output data) in an organized format, structure, or set of files. The storage 220 can include a database management system to facilitate data storage and retrieval. Although illustrated as a single device in FIG. 2, it should be understood that the storage 220 can include a plurality of databases or storage devices arranged side by side or distributed at specific locations. In some embodiments, the storage 220 can be realized in whole or in part as part of a remote network, such as cloud storage.
[0076] Processor 230 and / or memory 240 can also include a machine-readable medium for storing a set of software or instructions. As used herein, "software" broadly refers to any type of instructions, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. The instructions can include code (e.g., in source code format, binary code format, executable code format, or any other suitable format of code). The instructions, when executed by one or more processors 230, can cause the processor to perform various operations and functions described in more detail herein.
[0077] The implementation forms of the arithmetic unit 200 are not limited to the example embodiments shown in FIG. 2. The number and arrangement of the components (210, 220, 230, 240) can be modified and rearranged. Further, although not shown in FIG. 2, the arithmetic unit 200 can be in electronic communication with other networks including the Internet, local area network, wide area network, metro area network, and other networks capable of enabling communication between elements of the computing architecture. Also, the arithmetic unit 200 can retrieve the data or other information described herein from any source including the storage 220, together with a network or other database. Further, the arithmetic unit 200 can include one or more machine learning models used to implement the neural networks and other modules described herein, and can retrieve the weights or parameters, training information or training feedback, and / or any other data and information of the machine learning models described herein.
[0078] Figure 3A is a block diagram of an example local spatio-temporal processing module consistent with an embodiment of the present disclosure. The embodiment of Figure 3A can be used to implement the local spatio-temporal processing module 110 of the intelligent detector system 100 as an example in Figure 1A, or other computer-implemented systems such as the arithmetic unit 193 in Figure 1B. As illustrated in Figure 3A, the local spatio-temporal processing module 110 includes a number of components including a sampler 311, an encoder 312, a recurrent neural network (RNN) 313, a temporal convolutional network (TCN) 314, a quality network 315, and a segmentation network 316. These components can be implemented in any suitable combination of hardware, software, and firmware and can be used to select and process a subset of images. The local spatio-temporal processing module 110 can repeatedly apply various networks to the entire set of input videos or images (having T image frames) using a batch of N image frames (where N >= 1) as input. Such image frames may be consecutive or can be sampled at a fixed rate from the set of input videos or images.
[0079] Sampler 311 can select an image frame for processing by other components of the local spatio-temporal processing module 110. In some embodiments, Sampler 311 can perform buffering on a set of input videos or images for a set period of time, and extract the image frames in the buffer as a subset of the images for processing by module 110. In some embodiments, Sampler 311 can enable the configuration of the number of frames or size of the image subset selected for processing by the local spatio-temporal processing module 110. For example, Sampler 311 can be configured to receive user input to set or adjust the number of frames or size of the image subset for processing. Additionally, or alternatively, Sampler 311 can be configured to automatically select the number of image frames or the size of the image subset based on other factors such as the characteristics of the features of interest requested for processing and / or the characteristics related to the features of interest. In some embodiments, the amount or size of the sampled images can be based on the frame rate of the video (24, 30, 60, and 120 FPS). For example, Sampler 311 can perform buffering periodically on the real-time video stream received by the intelligent detector system 100 for a set period of time to extract images from the buffered video. As a further example, buffering can be performed on a stream of images from a pill camera or other imaging device, and the images for processing by system 100 can be extracted from the buffer. In some embodiments, Sampler 311 can selectively perform image sampling based on other factors such as video quality, video length, characteristics related to the features of interest requested, and / or the features of interest. In some embodiments, Sampler 311 can perform sampling on the image frames based on the components of the local spatio-temporal processing module 110 involved in the execution of the tasks requested by the user of the intelligent detector system 100. For example, an encoder 312 using a 3D encoder network can request multiple images to create the three-dimensional structure of the content to be encoded.
[0080] The encoder 312 can determine the presence of characteristics regarding the features of each object of interest as part of a task requested by a user of the intelligent detector system 100. The encoder 312 can be implemented for image analysis using a trained convolutional neural network. The intelligent detector system 100 can include a 2D encoder and a 3D encoder including a non-local layer as the encoder 312. The encoder 312 can be composed of a plurality of convolutional residuals and fully connected layers. Depending on the characteristics to be detected and the features of the object of interest, the encoder 312 can select a 2D or 3D convolutional encoder. The encoder 312 can be trained to detect characteristics in a requested image to detect the features of the object of interest in an image frame. As disclosed herein, the intelligent detector system 100 can process an image and use the encoder 312 to detect desirable characteristics regarding the features of the object of interest. The intelligent detector system 100 can determine desirable characteristics based on the trained network of the encoder 312 and past determinations of the features of the object of interest.
[0081] As shown in FIG. 3A, the local spacetime processing module 110 can also be equipped with a recurrent neural network (RNN) 313. The RNN 313 can function with an encoder 312 to determine the presence of desirable characteristics regarding the features of interest. In some embodiments, a temporal convolutional network (TCN) 314 can also be used to assist in detecting such characteristics in each image frame of the input video. The TCN is a specialized convolutional neural network that can handle a sequence of images, such as a temporally ordered set of frames of the input video from a source (e.g., the image source 150 of FIG. 1A or the imaging device 192 in FIG. 1B). The TCN can function on all the image frames of the input video buffered for processing by the local processing module 110, or on a subset of all the images of the input video processed by the global processing module 120. The TCN can process sequence data using causal convolutions in a fully connected convolutional network.
[0082] The RNN 313 is an artificial neural network having internal feedback connections and an internal memory status used to determine spatio-temporal information in the image frames. In some embodiments, the RNN 313 can include local layers to improve its ability to aggregate spatio-temporal information that is spatially and / or temporally distant in the image frames selected by the sampler 311 on which buffering has been performed. The RNN 313 can be configured, for example, to associate a score between 0 and 1 with each desirable characteristic regarding the requested features of interest. The score indicates the likelihood of the presence of the desirable characteristic in the image, with 0 being the lowest likelihood and 1 being the highest likelihood.
[0083] As shown in FIG. 3A, the local spacetime processing module 110 can include additional components such as a quality network 315 and a segmentation network 316 to further assist in identifying the characteristics required to detect features of interest in the processed image. For example, the quality network 315 can be implemented as a neural network composed of several convolutional layers and several fully connected layers that improve the final score assigned to an image frame. For example, the quality network 315 can remove image frames with low characteristic scores. The quality network 315 can also generate a feature vector based on the determined characteristics for each image frame. Each feature vector can provide a quality score represented as an ordinal number [0, R] indicating the frame image quality where 0 is the lowest quality and R is the highest quality. The quality network 315 can automatically set a quality score of 0 for those characteristics not detected in the image.
[0084] The segmentation network 316 can process an image to compute a segmentation mask for each input image to extract a portion of the image that contains characteristics related to the features of interest. The segmentation mask can be a binary mask for pixels at the same or lower resolution than the input image. The segmentation network 316 can be implemented as a deep convolutional neural network that includes a plurality of convolutional residual layers and a plurality of skip connections. The number and type of layers included in the segmentation network can be based on the characteristics or the features of interest required. In some embodiments, a single model with multiple layers can handle the tasks of the encoder 312 and the segmentation network 316. For example, the single model can be a U-Net model with a ResNet encoder.
[0085] As an example, the segmentation network 316 can take an image of dimension W×H as input and return a segmentation mask represented by a matrix of dimension W’×H’, where W’ is less than or equal to W and H’ is less than or equal to H. Each value in the output matrix represents the probability that a certain image frame coordinate contains a characteristic associated with the feature of interest being requested. In some embodiments, the intelligent detector system 100 can generate multiple output matrices for multiple features of interest. The multiple matrices may have different dimensions.
[0086] FIG. 3B is a block diagram of an exemplary global spatio-temporal processing module consistent with embodiments of the present disclosure. The embodiment of FIG. 3B can be used to implement the global spatio-temporal processing module 120 of the intelligent detector system 100 as an example in FIG. 1A or a system implemented by another computer such as the arithmetic unit 193 in FIG. 1B. The global spatio-temporal processing module 120 can modify or improve the output obtained from the local spatio-temporal processing module 110. Therefore, the global spatio-temporal processing module 120 can be configured to process the complete set of images together after the local spatio-temporal processing module 110 has processed all subsets of the input video image obtained, for example, from the image source 150 via the network 160. As an example, the global spatio-temporal processing module 120 processes the output of the complete set of images processed by the local spatio-temporal processing module 110 to modify its output by improving or removing outliers in the images having the determined characteristics or features of interest. In some embodiments, the global spatio-temporal processing module 120 can improve the quality scores and segmentation masks generated by the quality network 315 and the segmentation network 316.
[0087] The global spatio-temporal processing module 120 can improve the scores of the determined characteristics using one or more non-causal temporal convolutional networks (TCNs) 321. The global spatio-temporal processing module 120 can process the outputs of all images processed by the local spatio-temporal processing module 110 using an extended convolutional network included as the TCN 321. Such an extended convolutional network helps increase the receptive field without increasing the network depth (number of layers) or kernel size and can be used together for multiple images.
[0088] As further described herein, the TCN 321 can examine the entire time series of features K×T’ extracted using the local spatio-temporal processing module 110. The TCN 321 can take as input a matrix of features of dimension K×T’ and can return one or more time series of scalar values of length T”.
[0089] The global spatio-temporal processing module 120 can improve the quality scores generated by the quality network 315 using one or more signal processing techniques such as low-pass filters and Gaussian smoothing. In some embodiments, the global spatio-temporal processing module 120 can improve the segmentation masks generated by the segmentation network 316 using a cascade of morphological operations. The global spatio-temporal processing module 120 can improve the binary mask for segmentation by using prior information about the shape and distribution of the determined characteristics across the entire input image identified by the encoder 312 in a combination of the RNN 313 and the causal TCN 314.
[0090] The TCN321 can function on a complete set of videos or input images, so the local spatio-temporal processing module 110 does not need to wait for the processing of individual image frames to be completed. To meet this requirement, the TCN321 can be trained following the training of the network in the local spatio-temporal processing module 110. The number of layers and the architecture of the TCN321 depend on the task required by the user of the intelligent detector system 100 to detect certain features of interest. The TCN321 can be trained based on the required features of interest imposed on the system. The training algorithm for the TCN321 can adjust the parameters of the TCN321 for each such task or feature of interest.
[0091] As an example, the intelligent detector system 100 can first use the local spatio-temporal processing module 110 to calculate the K-dimensional time series for each video in the training set 415 and apply optimization based on gradient descent to estimate the TCN321 parameters to minimize the loss function L(s, s'), where s is the estimated output time series of the scores and s' is the ground truth time series. The intelligent detector system 100 can calculate the distance between s and s' using, for example, mean squared error (MSE), cross-entropy, and / or Huber loss.
[0092] Similar to the training process for other neural networks, data augmentation, learning rate adjustment, label smoothing, and / or batch training can be used when training the TCN321 to improve the capabilities and results of the TCN321.
[0093] The intelligent detector system 100 can be adapted to specific requirements by adjusting the hyperparameters of components 110-130. In some embodiments, the intelligent detector system 100 can modify the standard architecture of the pipeline to process a set of input videos or images by adding or removing components 110-130, or some part of components 110-130. For example, if the intelligent detector system 100 needs to handle a very local task, the use of TCN321 in the global spatio-temporal processing module 120 can be stopped to avoid any global spatio-temporal processing of the output generated by the local spatio-temporal processing module 110. As another example, if the user of the intelligent detector system 100 requests a diffusion task, RNN313 and / or TCN314 can be removed from the local spatio-temporal processing module 110. The intelligent detector system 100 can remove some of the RNN313 and TCN314 in the global spatio-temporal processing module 120 or the local spatio-temporal processing module 110 by stopping the operation of the associated network from the pipeline used to process the input video or image to detect the required medical condition.
[0094] Other arrangements or realizations of the system are also possible. For example, the segmentation network 316 can be made unnecessary when the required task for detecting features of interest does not address the focal object in the image frame of the image and can be removed from the pipeline. As another example, the quality network 315 can be made unnecessary when all image frames are recognized as useful and its operation can be stopped. For example, when the frame rate of the input video is low or there are too many erroneous image frames, the intelligent detector system 100 can avoid further removal of image frames using the quality network 315. As recognized from the present disclosure, the intelligent detector system 100 can perform preprocessing and / or sampling on the input video or image to determine the components that need to be activated and trained as part of the local spatio-temporal processing module 110 and the global spatio-temporal processing module 120.
[0095] Figure 3C is a block diagram of an example time series analysis module consistent with an embodiment of the present disclosure. The embodiment of Figure 3C can be used to implement the time series analysis module 130 of a system implemented by another computer such as the intelligent detector system 100 as an example in Figure 1A or the arithmetic unit 193 in Figure 1B. The time series analysis module 130 can use the scores, quality values, and segmentation maps for each image frame processed by the local spatio-temporal processing module 110 as inputs to generate the final output score for each image. In some embodiments, after the completion of a medical procedure, the time series analysis module 130 can use the scores, quality values, and / or segmentation maps of the images generated by the local spatio-temporal processing module to generate a dashboard or other display having a summary of the quality scores of all the images. The components of the time series analysis module 130 can be used to generate different output scores and values presented as an aggregated summary of the images processed by the intelligent detector system 100. As illustrated in Figure 3C, the components of the time series analysis module 130 can include an event detector 331, a frame selector 332, an object descriptor 333, and a time segmenter 334 to assist in generating the final output score for the images of the input video.
[0096] The event detector 331 can determine the start and stop times in the input video of an event associated with the characteristics of the requested object of interest. In some embodiments, the event detector 331 determines the start and end images in the input video of an event associated with the characteristics of the requested object of interest. In some embodiments, the start and stop times or image frames of an event may overlap.
[0097] The start and stop times of the event may be the start and end of a portion of the input video in which some characteristics regarding the feature of interest are detected. The start and stop times of the video may be estimated values due to missing image frames from the analysis by the local spatio-temporal processing module 110. The event detector 331 can output a list of pairs (t, d), where t is a time instance and d is a description of the event detected at that point in time. Various events can be identified based on different features of interest processed by the intelligent detection system 100.
[0098] The portion of the input video identified from the event can include the portion of the organ scanned by a healthcare professional or other operator to generate the input video as part of a medical procedure. For example, a medical procedure such as a colonoscopy can include events configured for various portions of the colon such as the ascending colon, transverse colon, or descending colon.
[0099] The time series analysis module 130 can provide a summary report of events for different portions of the video representing different portions of the medical procedure. The summary report can include, for example, the length of the video, or the time taken to complete the scan of the portion of the organ associated with the event, which can be listed as the departure time. The event detector 331 can assist in generating a summary report of different portions of the medical procedure that include events regarding the feature of interest.
[0100] The time series analysis module 130 can present a summary report of different parts of a medical procedure video (e.g., a colonoscopy video) on a dashboard or other display that shows a pie chart having the amount of video for different actions required for parts of the video or parts of the organ represented by the video parts, such as a careful second review, performing a biopsy, or removing a lesion. In some embodiments, the dashboard can include a quality summary detail of the events identified by the event detector 331 in a color-coded manner. For example, the dashboard can include buttons or other icons colored red, orange, and green to identify the quality of the video of the part of the medical procedure representing the event. The dashboard can also include a summary detail of the entire video representing the medical procedure at the same level of information provided for the individual parts of the medical procedure.
[0101] In some embodiments, the summary report generated by the time series analysis module 130 can identify one or more frames to more carefully review a part and / or to deal with other tissues. The summary report can also indicate what percentage of the video is for performing additional actions such as a second review. The time series analysis module 130 can use the frame selector 332 to identify a particular frame of the video or the percentage of that video for performing additional actions.
[0102] The frame selector 332 can search for image frames in the input video based on the scores generated by the feature and local spatio-temporal processing module 110. In some embodiments, the frame selector 332 can also utilize quality values provided by the user to select frame images. The frame selector 332 can select image frames based on the relevance of the image frames to the features of the characteristics and / or objects of interest requested by the user of the intelligent detection system 100.
[0103] In some embodiments, the summary report generated by the time series analysis module 130 can include one or more image frames identified by the frame selector 332. The image frames presented in the report can be enhanced to display marks applied to one or more portions of the frames. In some embodiments, the marks can identify features of interest such as lesions or polyps in the image frame. For example, a colored bounding box can be used as a mark surrounding the feature of interest (see, for example, the green bounding box applied to the image frame shown in FIG. 10 and therein). In some embodiments, different marks (including different combinations of shapes and / or colors) can be used to indicate different features of interest. For example, the image frame can be enhanced to include one or more marks in the form of differently colored bounding boxes representing different features of interest identified by the intelligent detector system 100.
[0104] The object descriptor 333 can fuse image frames of the input video that contain matching characteristics from the requested features of interest. The object descriptor fuses the image frames based on the spatio-temporal consistency information provided by the local spatio-temporal processing module 110. The output of the object descriptor 333 can include a set of objects described using a set of characteristics. The set of characteristics can include the timestamps of the image frames relative to other image frames of the input video. In some embodiments, the set of characteristics can include statistical values for the estimated scores and positions of the detected characteristics or requested features of interest in the image frame.
[0105] The time segmenter 334 divides the input video into time intervals. The time segmenter 334 can be divided based on the consistency for the task of determining the characteristics of the requested object of interest. The time segmenter 334 can output labels for each image frame of the input video in the form of {L_i}. The output labels can indicate the presence and probability of the characteristics of the requested object of interest in the image frame, as well as the position within the image frame. In some embodiments, the time segmenter 334 can output separate labels for each characteristic of the object of interest in each image frame.
[0106] In some embodiments, the time series analysis module 130 can generate a dashboard or other display that includes quality scores for medical procedures performed by a physician, a healthcare professional, or other operator. To provide the quality scores, the time series analysis module 130 can include a machine learning model trained based on videos of medical procedures performed by other physicians and operators performing different inspection actions. In particular, the machine learning model can be trained to recognize portions of the video while the inspection actions of the healthcare professional indicate a need for additional scrutiny. For example, an endoscopy specialist who carefully explores the surface of the colon / small intestine, as opposed to the time spent cleaning the colon / small intestine surface, performing surgery, or navigating, can indicate a requirement for additional scrutiny of the small intestine surface. The machine learning model used by the time series analysis module 130 can learn based on the time spent, the number of photos taken, and / or the number of repeated scans of a section of the medical procedure representing a portion of the organ, for special actions of the healthcare professional such as careful exploration. In some embodiments, the machine learning model can learn based on the amount of marks in the form of notes or flags added to the video, or to a region of an image frame in the video, for the actions of the healthcare professional.
[0107] In some embodiments, the time series analysis module 130 can generate a summary report of the quality score of the actions of a healthcare professional using information about the time spent performing an action (e.g., careful exploration, navigation, cleaning, etc.). In some embodiments, the percentage of the total time of a medical procedure for an action can be used to calculate the quality score of the medical procedure or a part of the medical procedure. The time series analysis module 130 can be configured to generate a quality summary report of the actions of a healthcare professional based on the configuration of the intelligent detector system 100 to include the actions performed by the healthcare professional as features of interest.
[0108] To generate a dashboard with the above summary scores, the time series module can utilize a combination of an event detector 331, a frame selector 332, an object descriptor 333, and a time segmenter 334. The dashboard can include one or more frames from a medical procedure selected by the frame selector 332, and information regarding the total time spent on the medical procedure and the time spent examining portions with a medical condition or other features of interest. The quality score summary of the statistical values describing the actions of the healthcare professional can be calculated for the entire medical procedure (e.g., a full colon scan) and / or for portions identifying anatomical sites (e.g., parts of the colon such as the ascending colon, transverse colon, and descending colon, etc.).
[0109] The time series analysis module 130 can use the event detector 331, the frame selector 332, the object descriptor 333, and the time segmenter 334 to generate aggregated information about different features of interest, such as different parts of an organ imaged during a medical procedure, the presence of each medical condition, and / or the actions of a healthcare professional performing the medical procedure. For example, the aggregated information can be generated based on a listing of various medical conditions at different sites using the object descriptor 333, a frame indicating the medical condition selected by the frame selector 332, and the amount of identified time spent at the site of each medical condition and in the actions of the healthcare professional determined by the event detector 331.
[0110] In some embodiments, the time series analysis module 130 can generate a summary of the input video processed by the local spatio-temporal processing module 110 and the global spatio-temporal processing module 120. The summary of the input video can include a portion of the input video that is extracted and combined into the summary of the input video having features of interest. In some embodiments, the user can choose to view only the summary video or to expand each of the intervals of the video discarded by the module. The time segmenter 334 of the time series analysis module 130 can assist in extracting portions of the input video having features of interest. In some embodiments, the time series analysis module 130 can generate a video summary by selecting relevant frames to produce a variable frame rate video output. The frame selector 332 of the time series analysis module 130 can assist in the selection and removal of frames in the output video summary. In some embodiments, the time series analysis module 130 can provide additional metadata to the input video or video summary. For example, the time series analysis module 130 can color-code the timeline of the input video when features of interest are present in the input video. The time series analysis module 130 can use different colors to highlight the timelines of different features of interest. In some embodiments, portions of the output video summary having features of interest can include text and graphics overlaid on the output video summary.
[0111] In some embodiments, to maximize performance, modules 110-130 can be trained to select optimal parameter values for the neural networks in each of modules 110-130.
[0112] The components of the local spatio-temporal processing module 110 shown in FIG. 3A can include a neural network that has been trained prior to being used to process an image and detect characteristics regarding features of interest. The neural network in the local spatio-temporal processing module 110 can be trained based on a video data set that includes three subsets: a training set 415, a validation set 416, and a test set 417 (see FIG. 4A). The training subset for the intelligent detector system 100 can require a labeled video set. For example, the labeled video set can include a target score assigned to video processing by components of the intelligent detector system 100. The labeled video set can also include the location of characteristics detectable in each image frame of the video set and the values of each characteristic. Both the labels and the input video set can be used for training purposes. In some embodiments, the labels for a subset of the video set can be used to determine the labels of other subsets used for the training neural network in components 110-130 of the intelligent detector system 100.
[0113] During the training process, the intelligent detector system 100 can perform sampling from the training dataset images processed by the neural network in the components 110-130 of the intelligent detector system 100 or from the image buffer, and can update their parameters by error backpropagation. The intelligent detector system 100 can use the validation set 416 of the video set to control the convergence of the ground truth value y' of the desired characteristics and the encoder 312 output value y. The intelligent detector system 100 can use the test set 417 of the video set to evaluate the performance of the encoder 312 for determining the values of the characteristics in the image frames of the training subset of the video set. The intelligent detector system 100 can continue training until the ground truth value y' converges with the output value y. Upon reaching convergence, the intelligent detector system 100 can complete the training procedure and can remove the temporary fully connected network. The intelligent detector system 100 finalizes the encoder 312 for the latest values of the parameters.
[0114] Figure 4A is a flow diagram showing the training as an example of the encoder component of the local spatio-temporal processing module of FIG. 3A, which is consistent with an embodiment of the present disclosure. Encoder components such as encoder 312 for FIG. 3A take as input a single image frame or a small buffer of N image frames and generate as output an M-dimensional feature vector. As illustrated in FIG. 4A, the intelligent detector system 100 can train encoder 312 by adding a temporary network. The temporary network may be a fully connected network (FCN) 411 added as an additional layer at the end of encoder 312 to train encoder 312. FCN 411 can take as input the feature vector of each image frame of the input generated by encoder 312 and return a single floating point value or a one-hot vector y. The intelligent detector system 100 can use a loss function 413 to evaluate the convergence of the ground truth value y' and the output y of encoder 312. The loss function may be an additional layer added as the last layer of encoder 312. The loss function 413 can be expressed as L(y, y'), which represents the distance between the ground truth value y' for the characteristic and the output y generated by encoder 312 for the image frames of the input video. The intelligent detector system 100 can use mean squared error (MSE), cross entropy, or Huber loss as the loss function 413 for training encoder 312 using FCN 411.
[0115] In some embodiments, the temporary network may be a decoder network used by the intelligent detector system 100 to train the encoder 312. The decoder network 412 may be a convolutional neural network that maps each feature vector estimated by the encoder 312 to a large matrix (l_out) of the same dimension as the image frame (l_in) of the input video. The decoder network 412 can use L(l_out, l_in) as the loss function 413 to calculate the distance between two images (or a buffer of N images). The loss function 413 used with the decoder network 412 can include the mean squared error (MSE), structural similarity (SSMI), or L1 norm. The decoder network 412 used as the temporary network to train the encoder 312 does not require the determination of ground truth values for the training / validation / test subsets 415-417 of the video set. The intelligent detector system 100 that uses the decoder network 412 as the temporary network to train the encoder 312 can control convergence in the validation set 416 and use the test set 417 to evaluate the expected performance of the encoder 312. After completing the training of the encoder 312, the intelligent detector system 100 can remove the decoder network 412 or stop its activation.
[0116] In the training methods of both using the fully-connected network 411 and the decoder network 412, the encoder 312 and other components of the intelligent detector system 100 can use techniques such as data augmentation, learning rate adjustment, label smoothing, mosaic, MixUp, and CutMix data augmentation, and / or batch training to improve the training process of the encoder 312. In some embodiments, the neural network in the intelligent detector system 100 may be affected by class imbalance, and ad-hoc weighted loss functions and importance sampling can be used to avoid prediction bias for the majority classes.
[0117] Figure 4B is a flowchart showing an example of training of neural network components of the local spatio-temporal processing module as an example of FIG. 3A. The training as an example of FIG. 4B can be used, for example, to train the recurrent neural network (RNN) 313 and the temporal convolutional network (TCN) 314 of the local spatio-temporal processing module 110, which is consistent with an embodiment of the present disclosure. Training a deep neural network (DNN) in the intelligent detector system 100, such as the RNN 313 and the TCN 314, requires preparing a set of annotated videos or images and a loss function during the training procedure. During the training procedure, the intelligent detector system 100 can adjust the network parameters using an optimization algorithm based on gradient descent.
[0118] The intelligent detector system 100 can train the RNN 313 and the TCN 314 using the output of the previously trained encoder 312. The input to the RNN 313 and the TCN 314 may be an M-dimensional feature vector per time instance output of the encoder 312. The RNN 313 and the TCN 314 aggregate a plurality of feature vectors generated by the encoder 312 by buffering the feature vectors generated by the encoder 312. The intelligent detector system 100 can train the RNN 313 and the TCN 314 by supplying a sequence of consecutive image frames to the encoder 312 and passing the generated feature vectors to the RNN 313 and the TCN 314. For a sequence of B images (or a set of buffered images), the encoder 312 generates B vectors of M encoded features and sends them to the RNN 313 or the TCN 314 to generate B vectors of K features.
[0119] The intelligent detector system 100 can train the RNN 313 and the TCN 314 by including a temporary fully connected network (FCN) 411 at the end of the RNN 313 and the TCN 314. The FCN 411 converts the K-dimensional feature vector generated by the RNN 313 or the TCN 314 into a one-dimensional score and compares it to the ground truth in the loss function to correct the parameters until there is convergence between the output vector and the ground truth vector. In some embodiments, the intelligent detector system 100 improves the RNN 313 and the TCN 314 by using data augmentation, learning rate adjustment, label smoothing, batch training, weighted sampling, and / or importance subsampling as part of training the RNN 313 and the TCN 314.
[0120] FIG. 4C is a flow diagram showing an example of training of the quality network and the segmentation network components of the local spatio-temporal processing module of FIG. 3A. The example training of FIG. 4C can be used to train the quality network 315 and the segmentation network 316 of the local spatio-temporal processing module 110 as an example, which is consistent with the embodiments of the present disclosure. The intelligent detector system 100 can train the quality network 315 similar to the training of the encoder 312, but does not require a temporary network (FCN 411 or decoder network 412 in FIG. 4A). The quality network 315 outputs a scalar value q representing the quality of the image. The intelligent detector system 100 can train the quality network 315 by comparing its output quality score q with the ground truth quality score q' associated with each image frame of the training set 415 of the video set until the difference between the values is minimized. The intelligent detector system 100 can use a loss function 413 represented as L(q,q') to minimize the difference between the ground truth value q' and the output quality score q and adjust the parameters of the quality network 315. The intelligent detector system 100 can train the quality network 315 with MSE, L1 norm as the loss function 413. The intelligent detector system 100 can use data augmentation, learning rate adjustment, label smoothing, and / or batch training techniques to improve the training results of the quality network 315.
[0121] The intelligent detector system 100 can train the segmentation network 316 using one or more individual images or small buffers of size N. The buffer size N can be based on the number of images considered by the encoder 312 trained in FIG. 4A. The intelligent detector system 100 requires an annotated ground truth map as part of the training set 415 to train the segmentation network 316. The intelligent detector system 100 can use a loss function 413 represented as a loss L(m, m') that defines the distance between the map m estimated by the segmentation network 316 and the ground truth map m'. The intelligent detector system 100 can compare the predicted map m and the ground truth map m' by using, for example, MSE with respect to pixels, L1 loss function, dice score, and smooth dice score.
[0122] In some embodiments, the intelligent detector system 100 can use data augmentation such as ad-hoc morphological operations and affine transformations on each image frame in the input video and the masks generated for each image frame, learning rate adjustment, label smoothing, and / or batch training to improve the results of the segmentation network 316.
[0123] FIG. 4D is a flowchart showing an example of training as the global spatio-temporal processing module of FIG. 3B. The example training of FIG. 4D can be used to train the global spatio-temporal processing module 120 consistent with the embodiments of the present disclosure. As illustrated in FIG. 4D, the intelligent detector system 100 can train the global spatio-temporal processing module 120 by using the output of the local spatio-temporal processing module 110. As recognized from the present disclosure, the local spatio-temporal processing module 110 needs to be trained before being used to train the global spatio-temporal processing module 120. The local spatio-temporal processing module 110 is trained by training each of its components individually as described in the above description of FIGS. 4A-4C.
[0124] The temporal convolutional network (TCN) 321 of the global spatio-temporal processing module 120 can access the entire time series of T'×K features extracted by functioning on T image frames so that the local spatio-temporal processing module 110 generates a 1×K dimensional feature vector. The global spatio-temporal processing module 120 takes the entire matrix T'×K of features as input and returns a time series of scalar values of length T". The intelligent detector system 100 can train the global spatio-temporal processing module 120 by training the TCN 321.
[0125] When the intelligent detector system 100 trains the TCN 321, as a result, the global spatio-temporal processing module 120 can consider the number of processing layers of the TCN 321 and their architecture. The number of layers and connections can vary based on the task of determining the features of interest and need to be adjusted for each task.
[0126] The intelligent detector system 100 trains the global spatio-temporal processing module 120 by calculating the K-dimensional time series of scores for the image frames of each video in the training set 415. The intelligent detector system 100 calculates the time series scores by providing the training set 415 videos as input to the pre-trained local spatio-temporal processing module 110 and providing its output to the global spatio-temporal processing module 120. The intelligent detector system 100 can use optimization based on gradient descent to estimate the network parameters of the TCN 321 neural network. Optimization based on gradient descent can minimize the distance between the time series score s output by the global spatio-temporal processing module 120 and the ground truth time series score s'. The loss function 413 used to train the global spatio-temporal processing module 120 can be the mean squared error (MSE), cross entropy, or Huber loss.
[0127] In some embodiments, the intelligent detector system 100 can use data augmentation, learning rate adjustment, label smoothing, and / or batch training techniques to improve the results of the trained global spatio-temporal processing module 120.
[0128] Figures 5A and 5B are schematic representations of a pipeline constructed from the components of an intelligent detector system as an example for processing a set of input videos or images. By way of example, the pipeline of Figures 5A and 5B can be constructed from the components of the intelligent detector system 100 of Figure 1A for processing an input video. The pipeline for processing an input video using the modules of the intelligent detector system 100 can vary in structure based on the input video to be processed and the type of features of the object of interest requested. The pipeline can include all or some of the components of each module (e.g., the encoder 312 of Figure 3A).
[0129] As illustrated in Figure 5A, the pipeline 500 includes the components of the intelligent detector system 100 to process the input video 501 and determine the features of the object of interest that can be requested by a user or a physician (e.g., via the user device 170 or the physician device 180 of Figure 1A). The pipeline 500 includes the components of the local spatio-temporal processing module 110 and the global spatio-temporal processing module 120 to generate matrices 531 and 541 of the determined characteristics in each image frame and the scores of the features of the object of interest requested. The pipeline 500 also includes the time series analysis module 130 to determine the features of the object of interest using the spatio-temporal information of the characteristics present in matrices 531 and 541.
[0130] The local spatio-temporal processing module 110 can output a K×T’ matrix 531 of characteristic scores. T’ is the number of frames of the input video 501 repeatedly analyzed by the local spatio-temporal processing module 110. The local spatio-temporal processing module 110 generates a vector of size K of characteristic scores for each analyzed frame of the T’ frames. The size K can match the number of features of the object of interest requested by the user of the intelligent detector system 100. The local spatio-temporal processing module 110 can process the input video 501 using the sampler 311 to search for some or all of the T image frames. The T’ frames analyzed by the components of the local spatio-temporal processing module to generate the characteristic scores may be less than or equal to the total number of image frames T of the input video 501. The sampler 311 can select T’ frames for analysis by other components 312 and 315 - 317. In some embodiments, the RNN 313 and the TCN 314 can generate scores only for the T’ image frames of the sampled frames. The networks 313 and 314 can include T’ image frames based on the presence of at least one characteristic of the requested features of the object of interest. The local spatio-temporal processing module uses only one set of the networks 313 or 314 to process the image frames and generate the matrix 531 of characteristic scores.
[0131] The local spatio-temporal processing module generates the matrix 531 of characteristic scores for the T’ image frames by scrutinizing each image frame individually or in combination with a subset of the image frames of the input video 501 buffered and provided by the sampler 311.
[0132] The local spacetime processing module 110 can generate additional matrices 532-534 of scores using networks 3115-317. The quality network 315 can generate a quality score for each image frame considered by the sampler 311 to determine characteristics regarding features of the object of interest in each image frame. As illustrated in FIG. 5A, the quality network 315 can generate a matrix 532 of quality scores. The matrix 532 may be a 1×T” vector of quality scores for T” frames analyzed by the quality network 315. The quality network 315 can analyze the image frames extracted by the sampler 311 to generate quality scores for the T” image frames. In some embodiments, T” may be less than the total number of frames T of the input video 501. The quality network 315 can only process T” frames having quality scores exceeding a threshold value.
[0133] The segmentation network 316 can generate a matrix 533 of segmentation masks by processing the T” image frames of the input video 501. The dimensions of the matrix 533 are W’×H’×T”, and it includes T” masks of height H’ and width H’. In some embodiments, the width W’ and height H’ of the segmentation mask may be less than the dimensions of the processed image frames. The segmentation network 316 can analyze the image frames extracted by the sampler 311 to generate segmentation masks for the T” image frames. In some embodiments, T” may be less than the total number of frames T of the input video 501. The segmentation network 316 can only process them with a segmentation mask if the T” frames contain at least some of the characteristics or features of the object of interest required.
[0134] As illustrated in FIG. 5A, the global spatio-temporal processing module 120 processes the scores of the K×T’ matrix 531 of characteristics output by the local spatio-temporal processing module 110 to generate a corrected matrix 541 of characteristic scores. The global spatio-temporal processing module 120 scrutinizes the characteristic scores of all T’ analyzed image frames by processing the matrix 531 of characteristics together. The TCN 321 in the global spatio-temporal processing module 120 can process the matrix 531 of characteristic scores to generate a matrix 526 of scores of dimension 1×T’. The TCN 321 generates the matrix 526 of scores by combining the scores of T’ image frames represented by vectors of size K. The global spatio-temporal processing module 120 can use the post-processor 522 to remove any outliers in the matrix 526. The post-processor 522 can employ standard signal processing techniques such as low-pass filters and Gaussian smoothing to remove outliers. The global spatio-temporal processing module 120 outputs a matrix 541 of floating-point scores of dimension 1×U’. The dimension of U’ may be less than or equal to T’. The post-processor 522 could also remove some of the T’ image frame scores to generate a matrix 541 with improved scores. In some embodiments, the U’ dimension may be larger than T’ obtained by performing upsampling on the input video (e.g., video 501) to increase the number of image frames of the input video. The global spatio-temporal processing module 120 can include an upsampling module to increase the number of frames. The global spatio-temporal processing module 120 can perform upsampling on the video 501 if the number of image frames with quality scores is less than a threshold. The global spatio-temporal processing module 120 can perform upsampling based on the image frames with high quality scores as determined by the quality network 315 of the local spatio-temporal processing module 110.In some embodiments, upsampling can be performed on video 501 before processing by global spatio-temporal processing module 120. For example, sampler 311 can perform upsampling on input video 501 to create additional image frames.
[0135] Global spatio-temporal processing module 120 can use matrices 532-534 of additional scores used to determine features of interest required to generate matrices 542-544 and post-processors 523-525 to refine in detail.
[0136] As an example, post-processor 523 refines quality score matrix 532 using one or more standard signal processing techniques such as low-pass filtering and Gaussian smoothing. Post-processor 523 outputs matrix 542 of dimension 1×U” of the refined scores. In some embodiments, the value U” may be different from the value T”. For example, U” may be less than T” if an image frame with a low quality score is ignored by post-processor 523. Alternatively, U” may be greater than T” when upsampling is performed on video 501 to generate more image frames and higher resolution image frames.
[0137] Post-processor 524 can refine segmentation mask matrix 533 using a cascade of morphological operations that utilize prior information about the shape and distribution of each feature of interest. Post-processor 524 can output matrix 543 of dimension W’×H’×U’’’. In some embodiments, the dimension of U’’’ may be different from T’’’. For example, U’’’ may be less than T’’’ if an image frame with a low quality score is ignored by post-processor 524. Alternatively, U’’’ may be greater than T’’’ when video 501 is upsampled to generate more image frames and higher resolution image frames.
[0138] As illustrated in FIG. 5A, the time series module 130 of the pipeline 500 can take the output matrices 541-544 from the global spatio-temporal processing module 120 to generate numerical values indicating the location in each image frame of the input video of the location in the input video and the features of the requested object of interest. The time series module 130 can use the feature scores in the matrix 541 and the quality score matrix 542 to select the image frame that best represents the presence of the features of each object of interest. In some embodiments, the time series module 130 can utilize the spatio-temporal information of the features of the image frames in the matrix 541 to determine the intervals in the input video 501 that contain the features of the object of interest.
[0139] As illustrated in FIG. 5B, the pipeline 502 shows an alternative architecture that does not include additional components such as the networks 315-317 and the post-processors 523-525 (also shown in FIG. 5A as such). The pipeline 502 can still generate the same feature score matrices 531 and 541 as the output of the local spatio-temporal processing module 110 and the global spatio-temporal processing module 120. The time series module 130 can take the matrix 541 as input to generate values that identify the features of the object of interest in the input video 501.
[0140] Figures 6A and 6B illustrate a set-up of different pipelines for performing multiple tasks using an intelligent detector system such as system 100 as an example of FIG. 1A, which is consistent with embodiments of the present disclosure. The modules of the intelligent detector system 100 can be configured and managed as pipelines for processing image data for different tasks. The task manager 140 can maintain different pipeline architectures and manage the data flow across different modules in the pipeline. In some embodiments, the intelligent detector system 100 can be utilized as different tasks to determine various features of interest requested by different users from the same input video. In such scenarios, the neural network of the intelligent detector system 100 can be trained for different tasks to determine the features of interest associated with each task. The intelligent detector system 100 can also be trained for different types of input videos generated by different medical devices and / or other imaging devices to determine different features of interest.
[0141] The task manager 140 can maintain separate pipelines for each task and train them independently. As illustrated in FIG. 6A, the intelligent detector system 100 can generate two separate pipelines 610 and 620 and use modules to train them to function for separate tasks 602 and 603 to process the input video 601 and detect features of different objects of interest. In some embodiments, the pipelines 610 and 620 are pre-trained to handle different tasks. Additionally, the intelligent detector system 100 can instantiate a pipeline by searching for relevant pre-trained modules for processing the input video 601. For example, the task manager 140 can include a local spatio-temporal processing module 611, a global spatio-temporal processing module 612, and a time-series module 613 in pipeline 610 as part of task 602 to process the input video 601 and determine the features of the required object of interest. Similarly, pipeline 620 can be constructed using a local spatio-temporal processing module 621, a global spatio-temporal processing module 622, and a time-series module 623 as part of task 603 to process video 601 and determine the features of the object of interest. Maintaining multiple pipelines aids in easily scaling to multiple tasks, but can result in redundant processing of image data by certain components. An efficient alternative approach of a hybrid pipeline architecture with a partial set of shared components is described in the embodiment of FIG. 6B as an example below.
[0142] FIG. 6B shows an alternative pipeline 650 having shared modules of the intelligent detector system 100 between tasks. The intelligent detector system 100 shares modules between different tasks by sharing some or all of the components of each module. The intelligent detector system 100 can share those components in the pipeline that have a lower dependence on the required tasks for processing the image data.
[0143] The sampler 631 and the quality network 635 may depend on the input image data and can function on the image data in the same way regardless of the features of the requested object of interest. Therefore, in the pipeline 650, the sampler 631 and the quality network 635, which are components of the local spatio-temporal processing module 630 that depend on the input data and are irrelevant to the requested tasks, are shared between the tasks 602 and 603 that process the input video 601. The pipeline 650 can share their outputs among multiple tasks processed by downstream components in the pipeline 650.
[0144] The encoder 632 may depend on the requested task of identifying the correct annotations for the image frames of the input video 601, but may depend more on the input data and can also be shared between different tasks. Therefore, the pipeline 650 can share the encoder 632 between the tasks 602 and 603. Furthermore, sharing the encoder 632 between tasks can improve its training due to a larger number of available samples across multiple tasks.
[0145] The quality network 635 functions directly on the quality of the image without depending on the requested task. Therefore, using separate instances of the quality network 635, the quality scores of the image frames in the input video (e.g., input video 601) have no relation to the requested tasks (e.g., tasks 602 and 603), and the same operation will be applied multiple times to the input video 601, resulting in redundant per-task operations.
[0146] The segmentation network 636 depends on the tasks required rather than the components considered above. However, since it is easier to generate multiple outputs for different tasks (e.g., tasks 602 and 603), it can still be shared. As illustrated in FIG. 6B, the segmentation network 636 is a modified version of the segmentation network 316 that can return multiple segmentation masks for each task as a matrix 653 for each image frame.
[0147] The neural networks 633 and 634 can include an instance of either the RNN 313 or the TCN 314 that generates a matrix of characteristic scores specific to the features of the required object of interest for identification in different tasks. The local spatio-temporal processing module 630 of the pipeline 650 can be configured to generate multiple copies of the encoder outputs 637 and 638, and one can be provided for each task as input to the multiple neural networks 633 and 634.
[0148] FIGS. 6C and 6D show an example pipeline setup for performing multiple tasks with aggregated outputs using an example intelligent detector system consistent with embodiments of the present disclosure. As illustrated in FIG. 6C, the pipelines 610 and 620 generate outputs by simultaneously using multiple time series analysis modules 671 - 673. For example, the time series analysis module 673 takes as input the data generated by both the pipelines 610 and 620. The time series analysis modules 671 and 672 can generate outputs of the intelligent detector system similar to the outputs of the pipelines 610 and 620 in FIG. 6A described above. The additional time series analysis module 673 can aggregate the data generated by the local and global spatio-temporal processing modules 611 and 612 and 621 and 622.
[0149] FIG. 6D illustrates a pipeline that shares local and global spatio-temporal processing modules 630 and 640, in addition to sharing the output of the modules, to perform a time series analysis. As illustrated in FIG. 6D, the time series analysis module 683 uses a sampler 631 and an encoder 632 to preprocess an image and take both vectors 661 and 664 that include scores of the images generated without preprocessing. Similar to the time series analysis module 673 in FIG. 6C, the time series analysis module 683 aggregates data to generate an output.
[0150] FIG. 6E shows a dashboard as an example having an output summary for a plurality of tasks generated using an exemplary intelligent detector system consistent with embodiments of the present disclosure. The dashboard 690 can provide a summary of a medical procedure, e.g., a colonoscopy performed by a healthcare professional. The dashboard 690 can provide information about different parts of the medical procedure and can include scores and / or other information for summarizing features of identified objects of interest, such as inspection actions of the healthcare professional's actions and the number of identified polyps.
[0151] As illustrated in FIG. 6E, the dashboard 690 can include quality score summaries for different parts of the colon (quality score summary 691 for the right colon, quality score summary 692 for the transverse colon, and quality score summary 693 for the left colon), along with a quality score summary 694 for the overall colon quality. The quality score summaries 691 - 694 can include time statistics for different actions such as careful exploration, performing surgery, cleaning / washing the mucosa, and quickly moving or navigating the colon or other human organs. The system 100 can determine the departure time and the length and / or percentage of time identified as "careful exploration" based on, for example, characteristics or factors related to the actions of healthcare professionals. The intelligent detector system 100 can, for example, identify actions performed by a healthcare professional as "careful exploration" based on the time spent by the healthcare professional analyzing a scanned part of an organ relative to other parts. For example, an endoscopist analyzing the mucosa as opposed to other actions such as cleaning / partially excising a lesion can be considered a "careful exploration" action. The time statistics can include summaries of other actions such as performing surgery, cleaning / washing an anatomical site or part of an organ (e.g., the mucosa), and quickly moving / navigating an anatomical location or organ during a medical procedure. Different medical procedures (e.g., colonoscopy, video surgery, video capsule-based scans) can include different actions by a healthcare professional as "careful exploration". The intelligent detector system 100 can be configured to label the actions of a healthcare professional as "careful exploration". The quality score summary dashboard 690 can also include a color-coded representation of the quality of the examination for each part of the medical procedure. For example, as illustrated in FIG. 6E, the quality score summary dashboard 690 can include circles or icons (e.g., red, orange, and green) colored with traffic signal colors highlighted to indicate the quality level of the examination for each part of the procedure.
[0152] FIG. 7 is a flowchart showing the operation of a method as an example for detecting a medical condition in an input video of an image, consistent with an embodiment of the present disclosure. The steps of method 700 can be performed, for example, by the intelligent detector system 100 of FIG. 1A operating on or using the functionality of the computing device 200 of FIG. 2. It will be recognized that the illustrated method 700 can be changed to modify the order of steps or to include additional steps.
[0153] In step 710, the intelligent detector system 100 can receive an ordered set of input videos or images via the network 160. As disclosed herein, the images to be processed can be temporally ordered. The intelligent detector system 100 can request images directly from the image source 150. In some embodiments, other external devices, such as the medical device 180 and the user device 170, can instruct the intelligent detector system 100 to request images from the image source 150. In some embodiments, the user device 170 can issue a request to detect features of interest in the currently streaming image or the image received by the image source 150.
[0154] In step 720, the intelligent detector system 100 can individually analyze a subset of the images to determine characteristics regarding each requested feature of interest. The intelligent detector system 100 can use the sampler 311 (as shown in FIG. 1A) to select a subset of the images for analysis using other components of the local spatio-temporal processing module 110 (as shown in FIG. 1A). Further, as disclosed herein, the local spatio-temporal processing module 110 can have a limited subset of the images when determining characteristics in the image.
[0155] As disclosed herein, the intelligent detector system 100 can enable the configuration of the number of images included in a subset of images. The intelligent detector system 100 can automatically configure the size of the subset based on the characteristics of the requested object of interest or related thereto. In some embodiments, the user of the intelligent detector system 100 can configure the size of the subset based on input from the user or a physician (e.g., via the user device 170 or the physician device 180 in FIG. 1A). The subset of images can overlap the images among them and can share the images. The intelligent detector system 100 can enable the configuration of the number of overlapping images among the subsets of images processed by the local spatio-temporal processing module 110. The intelligent detector system 100 can select subsets of images simultaneously. In some embodiments, the intelligent detector system 100 can receive an image stream from the image source 150 and store those images in a buffer until the requested number of images forming the subset is achieved.
[0156] The intelligent detector system 100 can analyze a subset of images using the local spatio-temporal processing module 110 to determine the likelihood of characteristics in each image of the subset of images. The likelihood of characteristics for each object of interest feature can be represented by a range of continuous or discrete values. For example, the likelihood of characteristics can be represented using values in the range between 0 and 1.
[0157] The intelligent detector system 100 can detect characteristics by encoding each image of the subset of images using the encoder 312. As part of the analysis process, the intelligent detector system 100 can use a recurrent neural network (e.g., RNN 313 as shown in FIG. 3A) to aggregate spatio-temporal information of the determined characteristics. In some embodiments, the intelligent detector system 100 can use a causal temporal convolutional network (e.g., TCN 314 as shown in FIG. 3A) to extract spatio-temporal information of the determined characteristics in each image of the subset of images.
[0158] The intelligent detector system 100 can use a quality network 315 (as shown in FIG. 3A) to determine additional information about each image. The intelligent detector system 100 can use the quality network 315 to determine a vector of quality scores corresponding to each image in a subset of images. The quality scores can be used to rank each image against an ideal image having the features of interest requested. The quality network 315 can output the quality scores as ordinals. The ordinals can be within a range of numbers, and if outside that range, the image may be of such low quality that it needs to be ignored. For example, the quality network 315 can output quality scores between 0 and R.
[0159] In some embodiments, the intelligent detector system 100 can use a segmentation network 316 to generate additional information about the characteristics. The additional information can include information about parts in each image. The intelligent detector system 100 can use the segmentation network 316 to extract parts of an image having the features of interest requested by generating a segmentation mask for each image in a subset of images. The segmentation network 316 can use a deep convolutional neural network to extract the image.
[0160] In step 730, the intelligent detector system 100 can process the image and a vector of information about the determined characteristics of the image in step 720. The intelligent detector system 100 can use the global spatio-temporal processing module 120 to process the output generated by the local spatio-temporal processing module 110 in step 720. The intelligent detector system 100 can process a vector of information associated with all the images to improve on the vector of information containing the characteristics determined in each image. The global spatio-temporal processing module 120 can apply a non-causal temporal convolutional network (e.g., the temporal convolutional network 321 of FIG. 3B) to improve on the characteristic information generated by the components of the local spatio-temporal processing module 110.
[0161] The intelligent detector system 100 can also use a post-processor (e.g., post-processor 322 as shown in FIG. 3B) to improve the vectors with additional information about the image and characteristics such as quality scores and segmentation masks. The intelligent detector system 100 can use one or more signal processing techniques to improve the quality score of each image in the ordered set of images. By way of example, the intelligent detector system 100 can use one or more signal processing techniques such as a low-pass filter or Gaussian smoothing to improve the quality score.
[0162] As shown, for example, in FIG. 5A, the post-processor 523 can take the quality score matrix 532 of the quality scores to generate an improved score matrix 542.
[0163] In some embodiments, the intelligent detector system 100 can use a post-processor (e.g., post-processor 322 as shown in FIG. 3B) to improve the segmentation mask used for image segmentation to extract the portion of each image that contains the features of the requested object of interest. The intelligent detector system 100 can use morphological operations to improve the segmentation mask by leveraging prior information about the shape and distribution of the characteristics or features of interest across the ordered set of images. For example, as shown in FIG. 5A, the post-processor 524 can take the matrix of the segmentation mask 533 as an input to generate a matrix with an improved segmentation mask 543.
[0164] In step 740, the intelligent detector system 100 can associate a numerical value with each image based on the characteristics with improvements added to each image in the ordered set of images in step 730. The components of the intelligent detector system 100 can interpret the assigned numerical value of each image to determine the probability and identify the features of the object of interest within each image. The intelligent detector system 100 can present different numerical values to indicate different states of the features of each required object of interest. For example, the intelligent detector system 100 can output a first numerical value for each image when the features of the required object of interest are detected, and can output a second numerical value for each image when the features of the required object of interest are not detected.
[0165] In some embodiments, the intelligent detector system 100 can interpret the associated numerical value to determine the position in the image when the characteristics of the features of the required object of interest are present, or the number of images containing the characteristics. Following step 740, the intelligent detector system 100 can generate a report having information about the features of each object of interest based on the numerical value associated with each image (step 750). As disclosed above, the report can be electronically presented in different formats (e.g., file, display, data transmission, etc.), and can include information about the presence of the features of each required object of interest, along with additional information and / or recommendations based on medical guidelines. When step 750 is completed, the intelligent detector system 100 completes the process (step 799), and, for example, completes the execution of method 700 on the computing device 200.
[0166] FIG. 8 is a flowchart showing the operation of a method as an example for spatio-temporal analysis of video content, consistent with an embodiment of the present disclosure. The steps of method 800 can be executed, for example, by the intelligent detector system 100 of FIG. 1A operating on the computing device 200 of FIG. 2, or operating using the functions of the computing device 200. It will be recognized that the illustrated method 800 can be modified to change the order of steps or to include additional steps.
[0167] In step 810, the intelligent detector system 100 can access a temporally ordered set of images of video content from an image source 150 (as shown in FIG. 1A) via a network 160 (as shown in FIG. 1A). In some embodiments, the intelligent detector system 100 can access the images by extracting them from the input video. In some embodiments, the received images can be stored and accessed from memory.
[0168] In step 820, the intelligent detector system 100 can detect the occurrence of an event in the temporally ordered set of images using the spatio-temporal information of the characteristics in each image of the ordered set of images. The intelligent detector system 100 can detect the event using an event detector 331 (as shown in FIG. 3C). The intelligent detector system 100 can use a local spatio-temporal processing module 110 and a global spatio-temporal processing module 120 to determine the spatio-temporal information. The intelligent detector system 100 can determine the spatio-temporal information in a two-step method. First, the local spatio-temporal processing module can search for spatio-temporal information about the characteristics by examining each image in the accessed set of images. In some embodiments, the local spatio-temporal processing module 110 can use a subset of the images. Second, the global spatio-temporal processing module 120 can use the spatio-temporal information about the local characteristics for each image to generate the combined spatio-temporal information for all the images by examining the spatio-temporal information of all the images generated by the local spatio-temporal processing module 110.
[0169] When the intelligent detector system 100 detects an event, it can add color to the portion of the video content timeline that corresponds to the subset of the temporally ordered set of images of the video content where the event was found.
[0170] Color can vary with the level of relevance of an image of a subset of a temporally ordered set of images to a characteristic of interest, with respect to the characteristics of the feature of interest. Color can vary with the level of relevance of an image of a temporally ordered subset of images to one or more characteristics.
[0171] The intelligent detector system 100 can use the determined spatio-temporal information of the characteristics to determine in a temporally ordered set of images in which there are events representing the occurrence of features of interest.
[0172] In step 830, the intelligent detector system 100 can use a frame selector 332 (as shown based on FIG. 2) to select an image from a group of images based on the associated score and the quality score of the image indicating the presence of characteristics regarding at least one feature of interest. The intelligent detector system 100 can use a quality network 335 to evaluate the quality score of each image of the images accessed in step 810. The frame selector 332 can scrutinize the images and use the quality score generated by the quality network 335 and the characteristic score generated in step 820 to determine the images with information. The intelligent detector system 100 can select an image frame by adding a bookmark to the image in the temporally ordered set of images.
[0173] In step 840, the intelligent detector system 100 can use an object descriptor 333 to fuse a subset of images having matching characteristics based on spatio-temporal consistency. The intelligent detector system can determine the spatio-temporal consistency of the characteristics using the spatio-temporal information of the characteristics in each image determined in step 820.
[0174] In step 850, the intelligent detector system 100 can use a time segmenter 334 (as shown in FIG. 3C) to divide a temporally ordered set of images that satisfy the temporal consistency of the selected task. The intelligent detector system 100 can divide the set of images by identifying subsets in which features of one or more objects of interest are present. The intelligent detector system 100 can use the spatio-temporal information of the characteristics determined in step 820 to determine temporal consistency. The intelligent detector system 100 can consider that an image has temporal consistency if the image has a matching presence of features of one or more objects of interest.
[0175] The intelligent detector system 100 can extract a clip of video content that matches one of the divided subsets of the temporally ordered set of images of the video. The extracted clip can include features of at least one object of interest. When step 850 is completed, the intelligent detector system 100 completes the execution 800 on the computing device 200 (step 899).
[0176] FIG. 9 is a flowchart showing the operations of a method 900 as an example for a plurality of tasks for a set of input images, consistent with an embodiment of the present disclosure. The steps of method 900 can be performed, for example, by the intelligent detector system 100 of FIG. 1A operating on the computing device 200 of FIG. 2 or operating using the functions of the computing device 200. It will be recognized that the illustrated method 900 can be changed to modify the order of steps or to include additional steps.
[0177] In step 910, the intelligent detector system 100 can receive a plurality of tasks (e.g., tasks 602 and 603 of FIG. 6A) and an input video (e.g., input video 601 of FIG. 6A) including a set of images. Each of the received tasks 602 and 603 can include a request to identify features of an object of interest in the set of input images in the input video.
[0178] In step 920, the intelligent detector system 100 can analyze a subset of images using the local spatio-temporal processing module 110 (as shown in FIG. 1A) to identify the presence of characteristics regarding each required feature of interest in each image of the subset of images.
[0179] In some embodiments, the intelligent detector system 100 can use the global spatio-temporal processing module 120 to improve the characteristics identified by the local spatio-temporal processing module 110 by removing mis-identified characteristics. In some embodiments, the global spatio-temporal processing module 120 can highlight and flag some of the characteristics identified by the local spatio-temporal processing module 110. In some embodiments, the global spatio-temporal processing module 120 can perform filtering using additional components such as the quality network 315 and the segmentation network 316 that have been applied once to a set of images to generate additional information about the input images.
[0180] In step 930, the intelligent detector system 100 can repeatedly execute the time series analysis module 130 for each task in the required set of tasks for associating a numerical score with each image in the input set of images. In some embodiments, the intelligent detector system 100 can include multiple instances of the time series module 130 to process multiple tasks simultaneously. For example, the time series modules 671 and 672 (as shown in FIG. 6B) simultaneously identify different sets of characteristics in the same set of images for different tasks 602 and 603. When step 930 is completed, the intelligent detector system 100 completes the execution 900 on the arithmetic unit 200 (step 999).
[0181] The figures and components in the drawings described above illustrate the architecture, functions, and operations of possible realizations of systems, methods, and computer hardware or software products according to various example embodiments of the present disclosure. For example, each block in a flowchart or diagram can represent a module, segment, or portion of code that includes one or more executable instructions for implementing the specified logical function. In some alternative realizations, it should also be understood that the functions indicated in the block can occur in an order other than that described in the drawings. As an example, two blocks or steps shown consecutively can be executed or realized substantially simultaneously, or the two blocks or steps can be executed in the reverse order depending on the functions involved. Additionally, some blocks or steps can be omitted. It should also be understood that each block or step in the figure, and combinations of blocks or steps, can be realized by a system based on special-purpose hardware for performing the specified functions or operations, or by a combination of special-purpose hardware and computer instructions. A computer program product (e.g., software or program instructions) can also be realized based on the described embodiments and the examples shown.
[0182] It should be recognized that the systems and methods described above can be varied in many ways and that different features can be combined in different ways. In particular, not all of the features shown above in a particular embodiment or realization are required in all embodiments or realizations. Further combinations of the above features and realizations are also considered to be within the scope of the embodiments or realizations disclosed herein.
[0183] Certain embodiments and features have been described and illustrated herein, but modifications, substitutions, variations, and equivalents will be apparent to those skilled in the art. Accordingly, it is to be understood that the scope of the appended claims is intended to cover all such modifications and variations that fall within the scope of the disclosed embodiments and the features of the illustrated embodiments. The embodiments described herein are presented by way of example only and are not limiting, and it is also to be understood that various changes may be made in form and detail. Any part of the systems and / or methods described herein can be implemented in any combination other than mutually exclusive combinations. By way of example, the embodiments described herein can include various combinations and / or smaller combinations of the functions, components, and / or features of the different embodiments described.
[0184] Furthermore, although embodiments have been described herein by way of example, the scope of the present disclosure includes embodiments having equivalent elements, modifications, omissions, combinations (e.g., combinations of aspects across different embodiments), adaptations, or alterations based on the embodiments disclosed herein. Moreover, elements in the claims should be construed broadly based on the language employed in the claims and not limited to the examples described herein or the examples during the prosecution of this application. Rather, these examples are to be construed as non-exclusive. Accordingly, it is intended that the specification and examples herein be considered as exemplary only, and the true scope and spirit are indicated by the scope of the following claims and the full scope of their equivalents.
Claims
1. A system implemented by a computer, comprising: one or more memory devices storing instructions executable by a processor; one or more processors configured to execute instructions to cause the system to perform operations for spatio-temporal analysis of an image captured by an imaging device, the operations comprising: accessing a temporally ordered set of images from the captured image; using an event detector to detect the occurrence of an event in the temporally ordered set of images, wherein a start time and an end time are identified by a start image frame and an end image frame in the temporally ordered set of images; using a frame selector to select an image from a group of images whose range is defined by the start image frame and the end image frame in the temporally ordered set of images, based on the associated score and quality score of the image, wherein the associated score of the selected image indicates the presence of at least one feature of interest; using an object descriptor to fuse a subset of images identified based on spatio-temporal consistency using spatio-temporal information from the selected image, based on the matching presence of the at least one feature of interest; a time segmenter to divide the temporally ordered set of images into time intervals that satisfy the temporal consistency of a selected task.
2. The one or more processors are further configured to: use a local spatio-temporal processing module to determine spatio-temporal information about characteristics of the at least one feature of interest for a subset of images of video content; use a global spatio-temporal processing module to determine the spatio-temporal information for all images of the video content.
3. To divide the temporally ordered set of images into time intervals, the one or more processors are further configured to: identify a subset of the temporally ordered set of images having the presence of the at least one feature of interest, or identify a subset of the temporally ordered set of images having the presence of an event.
4. To identify a subset of a temporally ordered set of images having the presence of the features of the at least one object of interest, the one or more processors are further configured to add a bookmark to the images in the temporally ordered set of images, where the bookmarked images are part of a subset of the temporally ordered set of images, the system of claim 3.
5. To identify a subset of a temporally ordered set of images having the presence of the features of the at least one object of interest, the one or more processors are further configured to extract a set of images from the subset of the temporally ordered set of images, the system of claim 3 or 4.
6. The system of claim 5, wherein the extracted set of images includes characteristics regarding the features of the at least one object of interest.
7. To identify a subset of a temporally ordered set of images having the presence of the features of the at least one object of interest, the one or more processors are further configured to add color to a portion of the timeline of the captured images that matches the subset of the temporally ordered set of images, the system of any one of claims 3 to 6.
8. The system of claim 7, wherein the color varies with the level of relevance of the subset of the temporally ordered set of images to the features of the at least one object of interest.
9. The system of claim 7 or 8, wherein the color varies with the level of relevance of the subset of the temporally ordered set of images to the features regarding the features of the at least one object of interest.
10. The system of any one of claims 7 to 9, wherein the color is different for different features of the at least one object of interest.
11. The system of any one of claims 7 to 10, wherein the color is different for different events detected using the event detector.
12. The system of any one of claims 7 to 11, wherein the timeline is presented as part of a video summary, the video summary including overlaid text and graphics.
13. The video summary is generated by selecting relevant frames from the captured images and has a variable frame rate video output, the system according to claim 12.
14. The occurrence of the event represents a part of a medical procedure, the system according to any one of claims 1 to 13.
15. The operation further generating a dashboard having a summary of a temporally ordered set of the images, the summary is selected using the module of the frame selector and includes images extended with display marks, the system according to any one of claims 1 to 14.
16. The generated dashboard includes a quality score of a medical procedure performed while an image is captured using the imaging device, the system according to claim 15.
17. The generated dashboard includes a quality score of an operator of the imaging device who performs a medical procedure, the system according to claim 15 or 16.
18. The generated dashboard comprises aggregated information from one or more of the event detector, frame selector, object descriptor, and time segmenter, the system according to any one of claims 15 to 17.
19. The ordered set of images is received directly from the imaging device during a medical procedure, the system according to any one of claims 1 to 18.
20. The presence of at least one characteristic of interest is determined from a portion of the captured image, the system according to any one of claims 1 to 19.
21. The system is configured to perform operations for performing a plurality of tasks on a set of images, the operations are receiving a plurality of tasks, at least one task being associated with a request to identify at least one characteristic of interest in the set of images, analyzing a subset of the images of the set of images to identify the presence of the characteristics associated with the at least one characteristic of interest using a local spatio-temporal processing module, repeating the execution of a time series analysis module for each task of the plurality of tasks that associates a numerical score for each task with each image of the subset of the images, the system according to any one of claims 1 to 20.
22. The system according to claim 21, wherein the local spatio - temporal processing module outputs an analyzed subset of images of the set of images, and each subset is associated with a task of the plurality of tasks.
23. The system according to claim 21 or 22, wherein the local spatio - temporal processing module determines the presence of the characteristic by determining a vector of quality scores, and each quality score in the vector of quality scores corresponds to each image of the subset of images.
24. The system according to any one of claims 21 to 23, wherein the local spatio - temporal processing module generates a set of feature vectors for features of interest regarding the plurality of tasks.
25. The operation further comprises using a global spatio - temporal processing module to analyze a set of feature vectors for the subset of images analyzed by the local spatio - temporal processing module, according to any one of claims 21 to 24.
26. The operation further comprises using the time - series analysis module to aggregate the output of the local spatio - temporal processing module for each task of the plurality of tasks, according to any one of claims 21 to 25.
27. A computer - implemented method for spatio - temporal analysis of images captured by an imaging device, comprising the following, executed by at least one processor: accessing a temporally ordered set of images from the captured images; using an event detector to detect the occurrence of an event in the temporally ordered set of images, where the start time and end time are identified by a start image frame and an end image frame in the temporally ordered set of images; using a frame selector to select an image from a group of images whose range is determined by the start image frame and the end image frame in the temporally ordered set of images, based on the associated score and quality score of the image, where the associated score of the selected image indicates the presence of at least one feature of interest. Using the object descriptor, based on the matching presence of the characteristics of the at least one object of interest, fuse a subset of images identified based on spatial and temporal consistency using the spatio-temporal information generated by the local spatio-temporal processing module from the selected images. A method implemented by a computer, including an operation of using a time segmenter to divide the temporally ordered set of the images into time intervals that satisfy the temporal consistency of the selected task. Claim 28 A non-transitory computer-readable medium, including instructions that, when executed by at least one processor, cause the at least one processor to perform operations for spatio-temporal analysis of images captured by an imaging device, and the operations include: accessing a temporally ordered set of images from the captured images; using an event detector to detect the occurrence of an event in the temporally ordered set of the images, where the start time and the end time are identified by a start image frame and an end image frame in the temporally ordered set of the images; using a frame selector to select an image from a group of images whose range is determined by the start image frame and the end image frame in the temporally ordered set of the images, based on the associated score and quality score of the image, where the associated score of the image indicates the presence of the characteristics of at least one object of interest; using the object descriptor, based on the matching presence of the characteristics of the at least one object of interest, fuse a subset of images identified based on spatial and temporal consistency using the spatio-temporal information generated by the local spatio-temporal processing module from the selected images; A non-transitory computer-readable medium, including using a time segmenter to divide the temporally ordered set of the images into time intervals that satisfy the temporal consistency of the selected task.
Citation Information
Patent Citations
Systems and method for selecting for display images captured in vivo
US20210137368A1