COMPUTER-IMPLEMENTED SYSTEM AND METHOD FOR OBJECT DETECTION - Patent application

JP2024526751A5Pending Publication Date: 2025-06-13COSMO ARTIFICIAL INTELLIGENCE AI LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024501862
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-12
Filing Date
2022-07-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Existing image analysis systems lack the ability to efficiently detect and characterize objects in real-time video, often resulting in false positives and limited response times, and do not provide comprehensive object information, particularly in medical applications such as polyp detection during endoscopic procedures.

Method used

A computer-implemented system utilizing trained neural networks to detect and characterize objects in real-time video, including classification, location, and size, with the capability to generate medical guidelines for immediate display during procedures, employing multiple neural networks in parallel to enhance processing efficiency.

Benefits of technology

The system provides accurate, real-time object detection and characterization with reduced latency, enabling informed medical decision-making by presenting essential guidelines, thus improving diagnostic accuracy and procedural efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented system is provided for receiving real-time video captured from a medical imaging device during a medical procedure. The real-time video may include a plurality of frames. The system may be adapted to apply one or more neural networks configured to detect objects of interest in the plurality of frames and identify a plurality of characteristics of the detected objects of interest, such as classification, size and / or location. In some embodiments, the system is adapted to identify a medical guideline based on one or more of the plurality of characteristics and display information of the medical guideline on a display device in real-time during the medical procedure.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 220,585, filed July 12, 2021, and European Priority Application No. 21185179.5, filed July 12, 2021, the entireties of each of the above-referenced applications being incorporated herein by reference.

[0002] The present disclosure relates generally to the field of imaging systems and computer-implemented systems and methods for processing real-time video. More specifically, but not by way of limitation, the present disclosure relates to systems, methods, and computer-readable media for processing frames of real-time video and performing object detection and characterization. The systems and methods disclosed herein may be used in a variety of applications, such as medical image analysis for polyp detection and characterization, including determining polyp classification, size, and location. The systems and methods disclosed herein may be implemented to provide real-time image processing functions, such as identifying medical guidelines based on one or more object characteristics and presenting medical guideline information in real time on a display device. [Background technology]

[0003] Modern vision and image analysis systems require the ability to detect and characterize objects of interest in a scene. The objects of interest may be people, places, features or things. In some applications, such as medical image analysis systems, the accuracy of object detection and characterization is important to ensure proper diagnosis and / or treatment. Exemplary objects of interest in medical applications include lesions, polyps and / or other anomalies in human tissue. Summary of the Invention [Problem to be solved by the invention]

[0004] Although various object detectors and classifiers have been developed, many of them suffer from shortcomings. For example, existing systems may lack the ability to detect changes in object type and / or produce false positives. Some may suffer from limited response time or an inability to efficiently process real-time video signals. Furthermore, existing systems may not provide object characterization capabilities or may provide only limited object information.

[0005] Thus, there is a need for improved image analysis systems, including medical image analysis. There is also a need for improved object detection and characterization solutions, including systems that can efficiently process real-time video and provide image analysis and object characterization information. Additionally, there is a need for computer-implemented systems and methods that can aggregate data for objects of interest and provide information according to a given context, including information related to the object's location, size, and / or classification. Such systems and methods are useful for applications such as polyp detection and characterization, including during endoscopy or other medical procedures. [Means for solving the problem]

[0006] A system, method and computer readable medium for processing real-time video including processing frames of real-time video and performing object detection and characterization according to some disclosed embodiments are provided. The disclosed embodiments also relate to systems and methods for object detection and characterization using real-time video from a medical imaging device. The disclosed embodiments have a trained neural network for detecting objects and determining a characterization, such as a classification, location and / or size, of the identified objects. In some embodiments, the trained neural networks are arranged to operate in parallel to more efficiently determine the characterization of each object during a medical procedure and optionally provide information related to medical guidelines. By way of example, a characterization network may be provided having a plurality of trained neural networks, each of which is configured to detect a characterization of an identified object, such as a classification, location or size. The trained neural networks of the characterization network may be applied and operated in parallel simultaneously to determine an object characterization for each identified object. As further disclosed herein, object detection and characterization may include detection and characterization of polyps as well as detection and characterization of other anomalies. These and other embodiments, features and implementations are described herein.

[0007] In accordance with the present disclosure, one or more computer systems may be configured to perform certain operations by installing software, firmware, hardware, or combinations thereof into the system that cause or initiate an operation during operation. One or more computer programs may be configured to perform operations by including instructions that, when executed by a data processing device (such as one or more processors), cause the device to perform such operations.

[0008] One general aspect includes a computer-implemented system for processing real-time video. The computer-implemented system may have at least one processor configured to receive real-time video captured from a medical imaging device during a medical procedure, the real-time video including a plurality of frames. The at least one processor may be configured to detect objects of interest in the plurality of frames and apply one or more neural networks implementing a trained classification network configured to determine a classification of the object of interest, a trained location network configured to determine a location associated with the object of interest, and a trained size network configured to determine a size associated with the object of interest. Additionally, the at least one processor may be configured to identify a medical guideline based on one or more of the classification, location, and size of the object of interest. The at least one processor may be further configured to present information of the identified medical guideline on a display device in real-time during the medical procedure. Other embodiments include corresponding computer methods, computer devices, and computer programs stored on one or more computer storage devices configured to perform the operations or features described above.

[0009] The implementation may have one or more of the following features: The medical procedure may include at least one of an endoscopy, a gastroscopy, a colonoscopy, or an intestinal endoscopy. The object of interest may include at least one of a formation of human tissue, a change in human tissue from one type of cell to another type of cell, an absence or lesion of human tissue from a location where human tissue is expected. As an example, the object of interest may be a polyp. The identified medical guideline information may include an indication to leave or resect the object of interest. The identified medical guideline information may include a type of resection. The at least one processor may be further configured to generate a confidence value associated with the identified medical guideline. The determined classification may be based on at least one of a histological classification, a morphological classification, a structural classification, or a malignancy classification. The determined location associated with the object of interest may be a location of a human body. The location of the human body may be one of a location of a rectum, a sigmoid colon, a descending colon, a transverse colon, an ascending colon, or a cecum. The determined size associated with the object of interest may be a numerical value or a size classification.

[0010] The at least one processor may be further configured to apply one or more neural networks implementing a trained quality network configured to determine a frame quality associated with at least one of the plurality of frames and generate a confidence value associated with the determined frame quality. The at least one processor may be further configured to aggregate data associated with the determined classification, location and size when at least one of the determined frame quality or confidence value exceeds a predefined threshold and present at least a portion of the aggregated data on a display device. Furthermore, the at least one processor may be further configured to detect a plurality of objects of interest in the plurality of frames and determine a plurality of classifications and a plurality of sizes associated with the plurality of objects of interest, wherein one classification and one size of the determined plurality of classifications and sizes is associated with a detected object of interest of the detected plurality of objects of interest. The at least one processor may be further configured to present information associated with one or more classifications and sizes of the plurality of classifications and the plurality of sizes on the display device. Implementations of these and other above-mentioned operations and techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0011] Another general aspect includes a computer-implemented system for processing real-time video. The computer-implemented system may have at least one processor configured to receive real-time video, the real-time video including a plurality of frames collected during a medical procedure. The at least one processor may be configured to apply one or more neural networks implementing a trained characterization network configured to detect objects of interest in the plurality of frames, determine a plurality of features associated with the objects of interest, and determine a confidence value associated with the plurality of features. The at least one processor may be further configured to identify medical guidelines based on one or more of the plurality of features and the confidence value, and present information regarding the identified medical guidelines on a display device in real-time during the medical procedure. Other embodiments include corresponding computer methods, computer devices, and computer programs stored on one or more computer storage devices configured to perform the operations or features described above.

[0012] The computer-implemented system implementation described above may have one or more of the following features: The medical procedure may include at least one of an endoscopy, a gastroscopy, a colonoscopy, or an intestinal endoscopy. The object of interest may include at least one of a formation of human tissue, a change in human tissue from one type of cell to another type of cell, an absence or lesion of human tissue from a location where human tissue is expected. As an example, the object of interest may be a polyp. The medical guideline information may include an indication to leave or resect the object of interest. The identified medical guideline information may include a type of resection. The at least one processor may be further configured to generate a confidence value associated with the identified medical guideline. The trained characterization network may include a trained classification network configured to determine a classification associated with the object of interest and generate a classification confidence value associated with the determined classification, a trained location network configured to determine a location associated with the object of interest and generate a location confidence value associated with the determined location, and a trained size network configured to determine a size associated with the object of interest and generate a size confidence value associated with the determined size. The at least one processor may be further configured to present information associated with at least one of the classification, location, or size on a display device.

[0013] The at least one processor may be further configured to apply one or more neural networks implementing a trained quality network configured to determine a frame quality associated with at least one of the plurality of frames and generate a confidence value associated with the determined frame quality. The at least one processor may be further configured to aggregate data associated with the plurality of features and present at least a portion of the aggregated data on a display device when at least one of the determined frame quality or confidence value is above a predefined threshold. The at least one processor may be further configured to detect a plurality of objects of interest in the plurality of frames, determine a set of a plurality of features associated with the plurality of objects of interest, one set of features of the plurality of features including characterization and size information associated with the detected one of the detected objects of interest, and present information associated with one or more sets of features of the plurality of sets of features on a display device. Implementations of these and other above-mentioned operations and techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0014] Another general aspect includes a computer-implemented system for processing real-time video. The computer-implemented system may have at least one processor, which may be configured to detect an object of interest in a plurality of frames received from a medical imaging device and perform a characterization of the object of interest, the characterization including determining a plurality of features associated with the object of interest. The plurality of features may include a location and a size of the object of interest. The at least one processor may be further configured to aggregate information associated with the determined location and size of the object of interest if the object of interest persists across two or more of the plurality of frames. The at least one processor may be further configured to present the aggregated information of the object of interest on a display device when the determined location is in a first body region and the determined size is within a first range, and to present information on a display device indicative of a status of the characterization of the object of interest when the determined location is in a second body region and the determined size is within a second range. Other embodiments include corresponding computer methods, computer devices, and computer programs stored on one or more computer storage devices configured to perform the operations or features described above.

[0015] The implementation may have one or more of the following features: The at least one processor may be further configured to identify a medical guideline based on the determined location and size of the object of interest and present information related to the identified medical guideline on the display device. The plurality of features may further include a classification of the object of interest based on at least one of a histological classification, a morphological classification, a structural classification, or a malignancy classification. The determined location associated with the object of interest may be at least one of a rectum, a sigmoid colon, a descending colon, a transverse colon, an ascending colon, or a cecum. The determined size associated with the object of interest may be a numerical value or a size classification.

[0016] The at least one processor may be further configured to detect a plurality of objects of interest in a plurality of frames, perform a characterization of the plurality of objects of interest, the characterization comprising determining a plurality of sets of features associated with the plurality of objects of interest, a feature set of the plurality of sets of features including the characterization and size information associated with the detected objects of interest of the plurality of objects of interest, and present information associated with one or more of the feature sets of the plurality of sets of features on a display device. Implementations of these and other above-mentioned operations and techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0017] Another general aspect includes a computer-implemented method for processing real-time video. The computer-implemented method may include receiving real-time video captured from a medical imaging device during a medical procedure, the real-time video including a plurality of frames. The method further includes detecting an object of interest in the plurality of frames and applying one or more neural networks implementing a trained classification network configured to determine a classification of the object of interest, a trained location network configured to determine a location associated with the object of interest, and a trained size network configured to determine a size associated with the object of interest. The method further includes identifying a medical guideline based on two or more of the classification, location, and size, and presenting the identified medical guideline information on a display device in real time during the medical procedure.

[0018] Systems and methods according to the present disclosure may be implemented using any suitable combination of software, firmware, and hardware. Implementations of the present disclosure may include programs or instructions that are specifically mechanically constructed and / or programmed to perform the functions associated with the disclosed operations. Additionally, non-transitory computer-readable storage media may be used that store program instructions executable by at least one processor to perform the steps and / or methods described herein.

[0019] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory but are not restrictive of the disclosed embodiments. [Brief description of the drawings]

[0020] The following drawings, which constitute a part of this specification, illustrate several embodiments and, together with the description, serve to explain the principles and features of the disclosed embodiments.

[0021] [Figure 1] FIG. 1 is a schematic diagram of an example computer-implemented system for processing real-time video in accordance with an embodiment of the present disclosure.

[0022] [Diagram 2] FIG. 2 is a block diagram of an example computing device that may be used in conjunction with the example system of FIG. 1 and other embodiments of the present disclosure.

[0023] [Figure 3A] FIG. 3A shows an exemplary frame image of a polyp according to an embodiment of the present disclosure. [Figure 3B] FIG. 3B shows an exemplary frame image of a polyp according to an embodiment of the present disclosure.

[0024] [Figure 4A] FIG. 4A illustrates an example system for processing real-time video according to an embodiment of the present disclosure.

[0025] [Figure 4B] FIG. 4B illustrates another exemplary system for processing real-time video according to an embodiment of the present disclosure. [Figure 4C] FIG. 4C illustrates another exemplary system for processing real-time video according to an embodiment of the present disclosure.

[0026] [Figure 5A]FIG. 5A illustrates a frame image of an exemplary augmented frame including characterization and medical guideline information in accordance with an embodiment of the present disclosure. [Figure 5B] FIG. 5B illustrates a frame image of an exemplary augmented frame including characterization and medical guideline information in accordance with an embodiment of the present disclosure. [Figure 5C] FIG. 5C illustrates a frame image of an exemplary augmented frame including characterization and medical guideline information in accordance with an embodiment of the present disclosure. [Figure 5D] FIG. 5D illustrates a frame image of an exemplary augmented frame including characterization and medical guideline information in accordance with an embodiment of the present disclosure. [Figure 5E] FIG. 5E illustrates a frame image of an exemplary augmented frame including characterization and medical guideline information in accordance with an embodiment of the present disclosure.

[0027] [Figure 6] FIG. 6 illustrates another example of a system for processing real-time video according to an embodiment of the present disclosure.

[0028] [Figure 7] FIG. 7 illustrates a block diagram of an exemplary method for aggregating information for presentation on a display device according to an embodiment of the present disclosure.

[0029] [Figure 8] FIG. 8 illustrates a block diagram of an exemplary method for processing real-time video according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0030] Exemplary embodiments will now be described with reference to the accompanying drawings, which are not necessarily drawn to scale. Although examples and features of the disclosed principles are described herein, modifications, adaptations and other implementations are possible without departing from the spirit and scope of the disclosed embodiments. In addition, the words "comprises," "has," "contains," and "includes," and other similar forms, have equivalent meanings and are open-ended in that they do not imply an exhaustive listing of such items or a limitation to only the listed items following any of these phrases. It should also be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.

[0031] In the following description, various examples are provided for purposes of explanation, but it will be understood that the present disclosure may be practiced without one or more of these details.

[0032] Throughout this disclosure, references are made to "disclosed embodiments," which refer to examples of the ideas, concepts, and / or implementations of the invention described herein. Many related and unrelated embodiments are described throughout this disclosure. The fact that some "disclosed embodiments" are described as exhibiting a feature or characteristic does not imply that other disclosed embodiments necessarily share that feature or characteristic.

[0033] The present disclosure is provided for the convenience of providing the reader with a basic understanding of some exemplary embodiments, and is not intended to completely define the scope of the disclosure. The present disclosure is not an extensive overview of all contemplated embodiments, and is not intended to identify key or critical elements of all embodiments, nor to delineate the scope of some or all aspects. Its purpose is to present some features of one or more embodiments in a simplified form as a prelude to the more detailed description presented later. For convenience, the terms "specific embodiment" or "exemplary embodiment" may be used herein to refer to a single embodiment or multiple embodiments of the present disclosure.

[0034] The embodiments described herein may refer to a non-transitory computer readable medium that includes instructions that, when executed by at least one processor, cause at least one processor to perform a method or sequence of operations. The non-transitory computer readable medium may be any medium that can store data in any memory in a manner that can be read by any computing device with a processor to execute the method or any other instructions stored in the memory. The non-transitory computer readable medium may be implemented as software, firmware, hardware, or any combination thereof. Preferably, the software is implemented as an application program embodied in a program storage unit or computer readable medium consisting of components, specific devices, and / or combinations of devices. The application program may be uploaded to and executed by a machine having any suitable architecture. Preferably, the machine may be implemented on a computer platform having hardware such as one or more central processing units (CPUs), memory, and input / output interfaces. The computer platform may include an operating system and microinstruction code. Various processes and functions described in this disclosure may be part of the microinstruction code or part of the application program or any combination thereof, and may be executed by the CPU, whether or not such a computer or processor is explicitly shown. In addition, various other peripheral units may be connected to the computer platform such as an additional data storage unit and a printing unit. Further, the non-transitory computer readable medium may be any computer readable medium except for a transitory propagating signal.

[0035] Memory may include any mechanism for storing electronic data or instructions, including random access memory (RAM), read only memory (ROM), hard disk, optical disk, magnetic media, flash memory, other permanent, fixed, volatile, or non-volatile memory. Memory may include one or more separate storage devices, collocated or distributed, capable of storing data structures, instructions, or any other data. Memory may further include memory portions that contain instructions for the processor to execute. Memory may be used as a working memory device for the processor or as a temporary storage device.

[0036] Some embodiments may include at least one processor. A processor is any physical device or group of devices with electrical circuitry that performs logical operations on inputs. For example, the at least one processor may include one or more integrated circuits (ICs) including an application specific integrated circuit (ASIC), a microchip, a microcontroller, a microprocessor, all or part of a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), a server, a virtual server, or other circuitry suitable for executing instructions or performing logical operations. The instructions executed by the at least one processor may be preloaded into a memory integrated or embedded in the controller, for example, or may be stored in a separate memory.

[0037] In some embodiments, the at least one processor may include multiple processors. Each processor may have a similar configuration, or the processors may be of different configurations that are electrically connected or disconnected from one another. For example, the processors may be separate circuits or integrated into a single circuit. When multiple processors are used, the processors may be configured to operate independently or in concert. The processors may be coupled electrically, magnetically, optically, acoustically, mechanically, or by other means that allow them to interact with each other.

[0038] According to the present disclosure, the disclosed embodiments may include a network. The network may comprise any type of physical or wireless computer networking configuration used for data exchange. For example, the network may be the Internet, a private data network, a virtual private network using a public network, a Wi-Fi network, a LAN or WAN network, and / or other suitable connections that enable information exchange between various components of the system. In some embodiments, the network may include one or more physical links used to exchange data, such as Ethernet, coaxial cable, twisted pair cable, optical fiber, or any other suitable physical medium for exchanging data. The network may include a public switched telephone network (PSTN) and / or a wireless cellular network. The network may be a secure network or an unsecure network. In other embodiments, one or more components of the system may communicate directly over a dedicated communications network. The direct communication may use any suitable technology, including, for example, BLUETOOTH, BLUETOOTH LE (BLE), Wi-Fi, Near Field Communication (NFC), or other suitable communications method that provides a medium for exchanging data and / or information between separate entities.

[0039] In some embodiments, a machine learning network or algorithm may be trained using training examples, for example as described below. Non-limiting examples of such machine learning algorithms may include classification algorithms, data regression algorithms, image segmentation algorithms, visual detection algorithms (such as object detectors, face detectors, people detectors, motion detectors, edge detectors), visual recognition algorithms (such as face recognition, people recognition, object recognition), speech recognition algorithms, mathematical embedding algorithms, natural language processing algorithms, support vector machines, random forests, nearest neighbor algorithms, deep learning algorithms, artificial neural network algorithms, convolutional neural network algorithms, recurrent neural network algorithms, linear machine learning models, nonlinear machine learning models, ensemble algorithms, etc. For example, the trained machine learning network or algorithm may include predictive models, classification models, regression models, clustering models, segmentation models, artificial neural networks (such as deep neural networks, convolutional neural networks, recurrent neural networks), random forests, support vector machines, etc. In some examples, the training examples may include input examples and desired outputs corresponding to the input examples. Additionally, in some examples, training machine learning algorithms using training examples may generate trained machine learning algorithms, which may be used to estimate outputs for inputs not included in the training examples. Training may be supervised or unsupervised, or a combination thereof. In some examples, engineers, scientists, processes, and machines that train machine learning algorithms may further use validation examples and / or test examples.For example, the validation examples and / or test examples may include input examples and desired outputs corresponding to the input examples, the trained machine learning algorithm and / or the intermediately trained machine learning algorithm may be used to estimate outputs of the input examples of the validation examples and / or the test examples, the estimated outputs may be compared with the corresponding desired outputs, and the trained machine learning algorithm and / or the intermediately trained machine learning algorithm may be evaluated based on the results of the comparison. In some examples, the machine learning algorithm may have parameters and hyperparameters, where the hyperparameters are set manually by a human or automatically by a process external to the machine learning algorithm (such as a hyperparameter search algorithm), and the machine learning algorithm is set by the machine learning algorithm according to the training examples. In some implementations, the hyperparameters are set according to the training examples and the validation examples, and the parameters are set according to the training examples and the selected hyperparameters. The machine learning network or algorithm may be further retrained based on the outputs.

[0040] Certain embodiments disclosed herein may include a computer-implemented system for performing operations or methods that include a series of steps. The computer-implemented systems and methods may be implemented by one or more computing devices that may include one or more processors described herein configured to process real-time video. The computing devices may be one or more computers or any other device capable of processing data. Such computing devices may have displays such as LED displays, augmented reality (AR), or virtual reality (VR) displays. However, the computing devices may also be implemented in a computing system that includes back-end components (e.g., as a data server), or includes middleware components (e.g., an application server), or includes front-end components (e.g., a user device with a graphical user interface or web browser through which a user can interact with an implementation of the systems and techniques described herein), or includes any combination of such back-end, middleware, or front-end components. The components of the system and / or computing devices may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet. The computing devices may include clients and servers. Generally, a client and server are remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0041] FIG. 1 illustrates an example of a computer-implemented system 100 for processing real-time video according to an embodiment of the present disclosure. As shown in FIG. 1, the system 100 includes an imaging device 140 and an operator 120 that operates and controls the imaging device 140 through control signals transmitted from the operator 120 to the imaging device 140. As an example, in an embodiment in which the video feed includes medical video, the operator 120 may be a doctor or other medical professional. The imaging device 140 may comprise a medical imaging device, such as an endoscopic imaging device, an X-ray device, a computed tomography (CT) device, a magnetic resonance imaging (MRI) device, or other medical imaging device that generates video or one or more images of a human body or a portion thereof. The operator 120 may control the imaging device 140, among other things, by controlling the capture rate of the imaging device 140 and / or the movement or navigation of the imaging device 140, for example, through or relative to the patient's or individual's body.

[0042] In the example of FIG. 1, imaging device 140 may transmit the captured video as a number of image frames to computing device 160. Computing device 160 may include one or more processors for processing the video as described herein (see, e.g., FIG. 2). In some embodiments, the one or more processors may be implemented as separate components (not shown) that are not part of computing device 160 but are in network communication therewith. In some embodiments, the one or more processors of computing device 160 may implement one or more networks, such as a trained neural network. Examples of neural networks include object detection networks, classification detection networks, position detection networks, size detection networks, or frame quality detection networks, as described further herein. Computing device 160 may receive and process the number of image frames from imaging device 140. In some embodiments, control or information signals may be exchanged between computing device 160 and operator 120 to control or direct the creation of one or more augmented videos. These control or information signals may be transmitted and received as data via imaging device 140 or directly from operator 120 to computing device 160. Examples of control and information signals include signals for controlling components of computing device 160, such as an object detection network, a classification detection network, a position detection network, a size detection network, or a frame quality detection network as described herein.

[0043] In the example of FIG. 1, the computing device 160 may process and enhance the video received from the imaging device 140 and then transmit the enhanced video to the display device 180. In some embodiments, enhancing or modifying the video may include providing one or more overlays, alphanumeric characters, shapes, diagrams, images, animated images, or any other suitable graphical representation within or along with the video frames. The video enhancement may provide information related to the object of interest, such as classification, size, and / or location information. Additionally or alternatively, the video enhancement may provide information related to medical guidelines. The information related to medical guidelines may be displayed separately and / or adjacent to the object of interest in the video, and simultaneously with other information related to the object of interest, such as classification, size, and / or location. As shown in FIG. 1, the computing device 160 may be configured to directly relay the original, unenhanced video from the imaging device 140 to the display device 180. For example, the computing device 160 may perform the direct relay under certain conditions, such as when there are no overlays or other enhancements to be generated. In some embodiments, the computing device 160 may perform the direct relay when the operator 120 sends a command to the computing device 160 as part of the control signal to perform the direct relay. Commands from operator 120 may be generated by operation of buttons and / or keys included on an operator device and / or input device (not shown), such as a mouse click, cursor hover, mouse over, button press, keyboard entry, voice command, interaction performed in virtual or augmented reality, or other input.

[0044] To enhance the video, computing device 160 may process the video from imaging device 140 and create a modified video stream for transmission to display device 180. The modified video may include the original image frames with the augmented information that is displayed to the operator via display device 180. Display device 180 may comprise any suitable display or similar hardware for displaying the video or modified video, such as an LCD, LED or OLED display, an augmented reality display or a virtual reality display.

[0045] 2 is a block diagram of an example computing device 200 for processing real-time video according to an embodiment of the present disclosure. Computing device 200 may be used in connection with implementing the example system of FIG. 1 (e.g., with computing device 160). It should be understood that in some embodiments, the computing device may include multiple subsystems, such as a cloud computing system, a server, or any other suitable components for receiving and processing real-time video.

[0046] 2, computing device 200 may include one or more processors 230, which may include, for example, one or more integrated circuits (ICs) including application specific integrated circuits (ASICs), microchips, microcontrollers, microprocessors, some or all of a central processing unit (CPU), graphics processing unit (GPU), digital signal processor (DSP), field programmable gate array (FPGA), server, virtual server, or other circuitry suitable for executing instructions or performing logical operations, as described above. In some embodiments, processor 230 may include or be a component of a larger processing unit implemented with one or more processors. One or more processors 230 may be implemented with a combination of general purpose microprocessors, microcontrollers, digital signal processors (DSPs), field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gate logic, discrete hardware components, dedicated hardware finite state machines, or other suitable entities capable of performing calculations or other manipulations of information.

[0047] As further shown in FIG. 2, the processor(s) 230 may be communicatively coupled to memory 240 via a bus or network 250. The bus or network 250 may be adapted for communication of data and other types of information. The memory 240 may include a memory portion 245 that includes instructions that, when executed by the processor(s) 230, perform operations and methods as described in more detail herein. The memory 240 may also serve as a working memory, temporary storage, and other memory or storage for the processor(s) 230, as the case may be. By way of example, but not limited to, the memory 240 may be a volatile memory, such as a random access memory (RAM), or a non-volatile memory (NVM), such as a flash memory.

[0048] The processor(s) 230 may be communicatively coupled to one or more I / O devices 210 via a bus or network 250. The I / O devices 210 may include any type of input and / or output device or peripheral device. The I / O devices 210 may include one or more network interface cards, APIs, data ports, and / or other components to support connectivity with the processor(s) 230 via the network 250.

[0049] As further shown in FIG. 2, the processor(s) 230 and other components (210, 240) of the computing device 200 may be communicatively coupled to a database or storage device 220. The storage device 220 may electronically store data in an organized format, structure, or set of files. The storage device 220 may include a database management system that facilitates storage and retrieval of data. Although the storage device 220 is depicted in FIG. 2 as a single device, it should be understood that the storage device 220 may include multiple devices, either collocated or distributed. In some embodiments, the storage device 220 may be implemented in a remote network, such as cloud storage.

[0050] The processor(s) 230 and / or memory 240 may include a machine-readable medium for storing software or instruction sets. As used herein, "software" refers broadly to any type of instructions, such as software, firmware, middleware, microcode, hardware description languages, etc. The instructions may include code (e.g., in source code format, binary code format, executable code format, or other suitable format of code). The instructions, when executed by the one or more processors 230, may cause the processors to perform various operations and functions, as described in more detail herein.

[0051] Implementations of the computing device 200 are not limited to the exemplary embodiment shown in FIG. 2. The number and arrangement of the components (210, 220, 230, 240) may be changed and rearranged. Additionally, although not shown in FIG. 2, the computing device 200 may be in electronic communication with other network(s), including the Internet, local area networks, wide area networks, metro area networks, and other networks that may enable communication between elements of a computing architecture. The computing device 200 may also retrieve data or other information described herein from any source, including the storage device 220, and from network(s) or other database(s). Additionally, the computing device 200 may include one or more machine learning models used to implement the neural networks described herein, and may retrieve or receive weights or parameters of the machine learning models, training information or training feedback, medical guidelines and / or guideline rules, and / or other data and information described herein.

[0052] According to embodiments of the present disclosure, a system, method, and computer readable medium for processing real-time video are provided. The systems and methods described herein may be implemented with the aid of at least one processor or CPU, FPGA, ASIC, or any other processing structure(s) of a computing device, or a non-transitory computer readable medium, such as a storage medium. As used herein, "real-time video" may refer to video received by at least one processor, computing device, and / or system without perceptible delay from a video source (e.g., an imaging device). For example, at least one processor may be configured to receive real-time video captured from a medical imaging device during a medical procedure according to the disclosed embodiments. The medical imaging device may be any device capable of generating a video or one or more images of a human body or a portion thereof, such as an endoscopy device, an X-ray device, a CT device, or an MRI device, as discussed above. A medical procedure is any act performed to determine, detect, measure, or diagnose a condition of a patient, such as an endoscopy, a gastroscopy, a colonoscopy, or an intestinal endoscopy. In embodiments where the medical procedure is an endoscopic procedure, the medical procedure may be used to identify an object of interest (e.g., a lesion or a polyp) at a location of the human body. The body location may be the rectum, sigmoid colon, descending colon, transverse colon, ascending colon or cecum, however, it should be understood that the disclosed systems and methods may be used in other contexts and applications.

[0053] A real-time video may include multiple frames, according to disclosed embodiments. A "frame" as used herein may refer to any digital representation, such as a collection of pixels, that represents a scene or field of view in a real-time video. In such embodiments, a pixel may represent a discrete element that is characterized by a value or intensity in a color space (e.g., based on an RGB, RYB, CMY, CMYK, or YUV color model). A frame may be encoded in any suitable format, such as Joint Photographic Experts Group (JPEG), Graphics Interchange Format (GIF), bitmap, Scalable Vector Graphics (SVG), Encapsulated PostScript (EPS), etc. The term "video" refers to a digital representation of a scene or area of ​​interest that is composed of multiple consecutive frames. A video may be encoded in any suitable format, such as Moving Picture Experts Group (MPEG), Flash Video, Audio Video Interleave (AVI), or other formats. However, a video need not be encoded, and more generally, a video may include multiple frames. The frames may be in any order, including random order. In some embodiments, the video or multiple frames may be paired with the audio.

[0054] The plurality of frames may include a representation of an object of interest. As used herein, "object of interest" may refer to any visual item or feature in the plurality of frames that is desired to be detected or characterized. For example, the object of interest may be a person, a place, an entity, a feature, an area, or other identifiable visual item or thing. In an embodiment in which the plurality of frames includes images captured from a medical imaging device, for example, the object of interest may include at least one of a formation of human tissue, a change in human tissue from one type of cell to another type of cell, an absence or lesion of human tissue from a location where human tissue is expected. Examples of objects of interest in a video captured by an imaging device may include a polyp (a growth protruding from the gastrointestinal mucosa), a tumor (a swelling in a part of the body), a bruise (a change from healthy cells to discolored cells), a depression (an absence of human tissue), or an ulcer or abscess (damaged tissue or lesion). Other examples of objects of interest will be apparent from this disclosure.

[0055] Although some embodiments are described herein with reference to objects of interest that are polyps, the disclosed systems and methods are not limited to polyps and may be utilized in other contexts and applications, including non-medical applications. As used herein, "polyp" may refer to a growth or lesion of the gastrointestinal mucosa, and may be used more generally herein to refer to candidate tissues whose detection and characterization may be objects of interest. Polyps may be characterized based on classifications such as histological, morphological, structural, or malignant tumor classifications. For example, polyps may be classified histologically using the Narrowband Imaging International Colon and Rectal Endoscopy (NICE) or Vienna classifications. According to the NICE classification system, polyps fall into one of three types: (Type 1) atypia serrated polyp or hyperplastic polyp; (Type 2) conventional adenoma; and (Type 3) carcinoma with deep submucosal invasion. According to the Vienna classification, polyps fall into one of five types: (Category 1) Neoplasia / dysplasia negative; (Category 2) Indefinite for neoplasia / dysplasia. (Category 3) Non-invasive low-grade neoplasms (low-grade adenoma / dysplasia). (Category 4) High-grade neoplasms of the mucosa, such as high-grade adenoma / dysplasia, non-invasive carcinoma (carcinoma in situ) or suspected invasive carcinoma. (Category 5) Invasive tumors, intramucosal carcinoma, submucosal carcinoma, etc.

[0056] Polyps may be classified morphologically using the Paris classification. According to the Paris classification system, polyps fall into one of three general types: (Type I) elevated or polypoid, such as peduncular, sessile, and broad-based; (Type II) flat or superficial, such as flat elevated, completely flat, and superficially depressed; (Type III) excavated, including excavated and ulcerated. Type I is often referred to as polypoid, while Types II and III are often referred to as non-polypoid. Polyps may be classified structurally based on their shape or appearance. For example, if the surface of the polyp is smooth or has a round appearance, the polyp is classified as benign, and if the surface of the polyp contains abnormal growths or is irregular in appearance, the polyp is classified as non-benign or malignant. Polyps may be classified based on the grade of malignancy. Malignancy may be based on the degree of invasiveness of the disease, such as the invasiveness of cancer. For example, if there is little or no cancer infiltration in or around the polyp, the polyp may be classified as benign, and if there is cancer infiltration in or around the polyp, the polyp may be classified as malignant or cancerous. Other classifications will be apparent based on this disclosure, and other classifications may be selected according to a particular application. Thus, the present disclosure is not limited to any particular classification or type of object of interest.

[0057] Polyps may be characterized based on their size. The size of a polyp may be expressed as a numerical value or a size classification. The size of a polyp may be expressed using any suitable measurement, such as, for example, millimeters (mm), but any other measurement (e.g., inches) may be used. Thus, a polyp may have a size of 1 mm, 5 mm, 10 mm, etc. The size of a polyp may be expressed as a classification based on one or more suitable size categories, such as: (1) "small" or "small" for polyps having a size of 5 mm or less, (2) "not small" or "large" for polyps having a size between 6 mm and 9 mm, and (3) "very large" for polyps having a size of 10 mm or more. As will be appreciated, other values, categories, or labels may be used. Other size representations may be used depending on the particular application and object of interest.

[0058] 3A and 3B are diagrams illustrating framed images of an example polyp as part of an augmented video display according to an embodiment of the present disclosure. As illustrated, the augmented display (e.g., presented on display device 180) includes a rectangular bounding box surrounding the identified object of interest (i.e., polyp) and information of determined properties of the object, such as classification, location, and / or size. In the examples of FIGS. 3A and 3B, the illustrated polyp is characterized based on classification and size according to the above description. For example, as shown in FIG. 3A, the illustrated polyp is classified as non-adenomatous (i.e., not adenomatous) having a small size (i.e., having a size of 5 mm or less). Meanwhile, in FIG. 3B, the illustrated polyp is classified as adenoma (i.e., pre-cancerous) having a non-small size (i.e., having a size between 6 mm and 9 mm). As described herein, other characterizations, classifications, and size categories may be used. As described further below, information related to medical guidelines may be provided as part of the displayed augmented video (see, e.g., FIGS. 5A-5E).

[0059] At least one processor of the computing device 160 (FIG. 1) may be configured to detect an object of interest in a plurality of frames. The object of interest may be detected, for example, based on determining the presence or absence of an object of interest in a plurality of frames of a video. The object detection may be performed, for example, using one or more machine learning detection networks or algorithms, traditional detection algorithms, or a combination of both. For example, the plurality of frames may be fed into one or more neural networks (e.g., deep neural networks, convolutional neural networks, recurrent neural networks, etc.), random forests, support vector machines, or any other suitable models trained to detect the object of interest as described above. The machine learning detection algorithms, models, or weights may be stored in the computing device and / or system, or may be fetched from a network or database prior to detection. In some embodiments, the machine learning detection network or algorithm may be retrained based on one or more outputs, such as true / false positive detection or true / false negative detection. Feedback for retraining may be generated automatically by the system or computing device, or may be manually input by an operator or another user (e.g., via a mouse, keyboard, or other input device). Weights or other parameters of the machine learning detection network or algorithm may be adjusted based on the feedback. Additionally, traditional non-machine learning detection algorithms may be used alone or in combination with the machine learning detection network or algorithm.

[0060] For example, the presence of a polyp may be detected using one or more machine learning detection networks or algorithms, traditional detection algorithms, or a combination of both according to the disclosed embodiments. The detection of a polyp may include a determined location of the polyp in the frame, and may be indicated using any suitable graphical representation overlaid on the detected frame. For example, in Figures 3A and 3B, each illustrated polyp is surrounded by a rectangular bounding box indicating the detection of the illustrated polyp. The detection may be represented using other graphical representations (e.g., different shapes, colors, patterns, images, videos, and / or alphanumeric characters) or may not be represented at all as part of the video display.

[0061] At least one processor of the computing device 160 may be configured to apply one or more neural networks to implement a trained characterization network configured to determine a plurality of features associated with an object of interest from a plurality of frames according to the disclosed embodiments. The trained characterization network may include one or more suitable machine learning networks or algorithms for determining a plurality of features associated with an object of interest, including one or more neural networks (e.g., deep neural networks, convolutional neural networks, recurrent neural networks, etc.), random forests, support vector machines, or any other suitable models as described above trained to determine an object of interest feature of a plurality of objects of interest. The characterization network may be trained using a plurality of training frames, or portions thereof, labeled based on a desired feature (e.g., classification, size, or location). For example, a first set of training frames (or portions of frames) that may or may not include the object of interest are labeled as having the feature, and a second set of training frames (or portions of frames) that may or may not include the object of interest are labeled as not having the feature. The weights or other parameters of the characterization network may be adjusted based on the output for a third set of unlabeled training frames (or portions of frames) until convergence or other metric is achieved, and the process may be repeated using additional training frames (or portions thereof) or live video or frame data as further described herein.

[0062] In some embodiments, the characterization network of the computing device 160 may be implemented using multiple trained neural networks, each of which is adapted to determine a particular property or characteristic of an object identified from a real-time video or frame. For example, the multiple trained neural networks may include at least one trained neural network for determining a classification of an object, at least one trained neural network for determining a size of an object, and at least one trained neural network for determining a location of an object. As part of the computing device 160, each of the trained neural networks may be configured to operate simultaneously or in parallel with each other to more efficiently determine object properties and other information based on real-time video or frames from the imaging device 140. For example, for each identified object, a trained neural network of the characterization network is applied and operates simultaneously in parallel with each other to more efficiently determine the object's property. Optimizing the placement of the trained neural network also enables computing device 160 to generate an augmented video display including all determined information related to detected objects with little or no delay or latency perceived by a clinician or operator viewing the output of display device 180 while performing a medical procedure using imaging device 140. In some embodiments, the trained neural network of computing device 160 may be implemented to determine other information, such as confidence values ​​or medical guidelines, as further disclosed herein.

[0063] In some embodiments, the trained characterization network of the computing device 160 may be configured to determine a confidence value associated with a plurality of features. The confidence value of an identified feature may refer to an indication of a level of certainty associated with the identified feature. For example, a confidence value of 0.9 or 90% indicates that there is a 90 percent certainty that the identified feature is present in the object of interest, while a confidence value of 0.4 or 40% indicates that there is a 40 percent certainty that the identified feature is present in the object of interest, and so on. However, other values, metrics, or representations may be used to represent the confidence values, such as alphabetical letters (e.g., "A" for high confidence values ​​and "F" for low confidence values), colors (e.g., green for high confidence values ​​and red for low confidence values), shapes (e.g., a check for high confidence values ​​and a cross for low confidence values), or other suitable representations.

[0064] At least one processor of the computing device 160 may be configured to apply one or more neural networks to determine the confidence values. In some embodiments, the confidence values ​​associated with the outputs may be implicitly defined in a one-hot encoding formula, where a score is output by the network for each possible class. During training, neural network calibration methods such as mix-up and label smoothing may be used to control the range and distribution of confidence values ​​or scores. Alternatively, in other embodiments, a dedicated output node may be added to each of the neural networks that perform confidence estimation, and an abstraction term may be included in the loss function to train the neural network appropriately. By adding an additional term to the loss function, the neural network can predict low confidence values ​​when the estimation error is high, for example due to low quality or cluttered images. In yet another embodiment, the neural network is trained to predict both the output and the confidence score or label.

[0065] In some embodiments, each neural network is trained to obtain a correspondence between the confidence score threshold output by the network and the performance of a particular task, such as characterizing an identified object in an input image frame, in a validation step. The validation phase may use an independent labeled data set to obtain a correspondence between the confidence score threshold from the neural network and each of the specific properties or features of the object. The computing device 160 may determine a confidence score threshold to achieve an expected performance level of the neural network performing the specific task of characterizing an identified object. For example, it may be found that when the confidence score threshold is set to 0.6, the trained neural network achieves a performance sensitivity of 99% when performing the specific task. The computing device 160 may then execute the neural network to select image frames that have an expected performance level of the specific task based on the image frames that meet the determined confidence score threshold. Additionally, the computing device 160 may use the desired performance level to achieve medical guidelines when using frames from the imaging device 140. One or more metrics such as accuracy, precision, recall, sensitivity, npv, specificity, etc. may be used for the measurement. This approach allows one to obtain a correlation between a predefined threshold of the confidence score output by the neural network and the expected accuracy of the characteristics or features determined for the identified object.

[0066] In some embodiments, at least one processor of the computing device 160 may be configured to apply one or more specific machine learning networks or algorithms trained to detect one or more specific features. For example, at least one processor may be configured to apply one or more neural networks implementing a trained classification network configured to determine a classification of the object of interest. In some embodiments, the trained classification network may be configured to generate a classification confidence value associated with the determined classification. The classification may be the same or similar to those described above (e.g., based on at least one of a histological classification, a morphological classification, a structural classification, or a malignancy classification). The classification network may be trained using a plurality of training frames or portions thereof labeled based on one or more classifications. For example, a first set of training frames (or portions of frames) that may or may not include the object of interest may be labeled as "adenoma" and a second set of training frames (or portions of frames) that may or may not include the object of interest may be labeled as "non-adenoma" or another classification (e.g., "serrated"). Other labeling rules, both binary (e.g., "hyperplasia" vs. "non-hyperplasia") and multi-class (e.g., "adenoma" vs. "adenoid" vs. "hyperplasia") can be used. Weights or other parameters of the classification network may be adjusted based on the output on a third unlabeled set of training frames (or portions of frames) until convergence or other metric is achieved, and the process may be repeated using additional training frames (or portions thereof) or live data as described herein.

[0067] At least one processor of the computing device 160 may also be configured to apply one or more neural networks implementing a trained location network configured to determine an associated tagged location associated with the object of interest. The trained location network may be configured to generate a location confidence value associated with the determined location. The location may be the same or similar to those described above (e.g., a location of the human body such as the rectum, sigmoid colon, descending colon, transverse colon, ascending colon, or cecum). The location network may be trained using a plurality of training frames or portions thereof labeled based on one or more locations (e.g., body locations). For example, a first set of training frames (or portions of frames) that may or may not include the object of interest may be labeled as "rectus sigma" and a second set of training frames (or portions of frames) that may or may not include the object of interest may be labeled as "not rectus sigma" or another location in the body (e.g., "ascending colon"). The weights or other parameters of the location network may be adjusted based on the output for a third unlabeled set of training frames (or portions of frames) until convergence or other metric is achieved, and the process may be repeated using additional training frames (or portions thereof) or live data as described herein.

[0068] Additionally, at least one processor of the computing device 160 may be configured to apply one or more neural networks implementing a trained size network configured to determine a size associated with the object of interest, and in some embodiments, a trained classification network may be configured to generate a size confidence value associated with the determined size. The size may be the same or similar to those described above (e.g., a numerical value or a size classification). The location network may be trained using a plurality of training frames or portions thereof labeled based on size. For example, a first set of training frames (or portions of frames) with or without the object of interest may be labeled as "small" or "compact" and a second set of training frames (or portions of frames) with or without the object of interest may be labeled as "not small", "not compact" or another size value (e.g., "large" or "10 mm"). Weights or other parameters of the location network may be adjusted based on the output for a third unlabeled set of training frames (or portions of frames) until convergence or other metrics are achieved, and the process may be repeated using additional training frames (or portions thereof) or live data as described herein.

[0069] According to the above description, the trained classification, location and / or size network may be stored in the computer device and / or system, or may be fetched from the network or database before characterization. In some embodiments, the machine learning detection network or algorithm may be retrained based on one or more outputs, such as true / false positive detection or true / false negative detection. Feedback for retraining may be generated automatically by the system or computer device, or may be manually input by an operator or another user (e.g., via a mouse, keyboard, or other input device). Weights or other parameters of the machine learning detection network or algorithm may be adjusted based on the feedback. Additionally, conventional non-machine learning detection algorithms may be used alone or in combination with the machine learning detection network or algorithm.

[0070] At least one processor of the computing device 160 may be configured to identify a medical guideline based on one or more of the plurality of features and / or confidence values. As used herein, a "medical guideline" may refer to any information provided for the purpose of assisting in the determination, diagnosis, or treatment of a patient's condition. For example, in an embodiment where the object of interest is a polyp or other formation of human tissue, the identified medical guideline may include instructions to leave or remove the object of interest. In some embodiments, when the medical guideline includes instructions to remove the object of interest, the medical guideline may include a specification or description of a particular type of resection. For example, the medical guideline may include instructions to perform an endoscopic mucosal resection (EMR) to remove a small polyp or precancerous growth, or an endoscopic submucosal dissection (ESD) to remove a large polyp or a potentially cancerous growth, or other types of instructions. However, other medical guidelines may be identified based on a particular application or situation, such as a detailed examination of the object of interest, performing other medical tests, performing surgery, prescribing a medication, or performing other treatment. The at least one processor may be configured to present information regarding the identified medical guideline on the display device in real time during capture (e.g., during the medical procedure), as described above. The information displayed about the identified medical guideline may be any suitable display, such as one or more alphanumeric characters (e.g., the word "save" or "resection"), or an abbreviation or text that more specifically indicates the type of resection (e.g., "EMR" or "EDS"), a shape (e.g., a check mark or cross), a color (e.g., green or red), an image (e.g., an image of a hand or medical instrument), a video (e.g., a video of the recommended procedure), or a combination thereof. As another example, in an augmented video display, the medical guideline may be displayed separately from and / or near the object of interest (see, e.g., FIGS. 5A-5E). Additionally, the medical guideline information may be displayed simultaneously with other information related to the object of interest, such as classification, size, and / or location.When there are multiple objects of interest identified in a frame, the medical guidelines and other information presented in the augmented video display may be color adjusted to use unique colors for the bounding box, classification information, and / or display of the medical guideline for each identified object. Similar techniques may be used to allow a clinician or operator to more quickly identify or differentiate the displayed information, such as using unique colors for the presented information depending on the type and / or urgency of the classification of the recommended medical guideline (e.g., green for "non-adenoma" and / or "leave alone" but red for "adenoma" and / or "resection").

[0071] The at least one processor may be configured to generate a confidence value associated with the identified medical guideline. The confidence value of the identified medical guideline may refer to an indication of a level of confidence associated with the identified medical guideline. For example, a confidence value of 0.9 or 90% indicates that there is a 90 percent certainty that the identified medical guideline is accurate, while a confidence value of 0.4 or 40% indicates that there is a 40 percent certainty that the identified medical guideline is accurate, and so on. However, other values, metrics, or representations may be used to represent the confidence value. For example, the confidence value may be classified as "high confidence" when the confidence value is above a predefined threshold (e.g., above 66% confidence), as "low confidence" when the confidence value is below a predefined threshold (e.g., below 33% confidence), or as "undetermined" when the confidence value is between two predefined thresholds (e.g., confidence is between 33% and 66% confidence). In some embodiments, the at least one processor may present a confidence score associated with the medical guideline. For example, the confidence score may be expressed as an alphanumeric representation aligned with medical guidelines, such as "Keep - High Confidence" or "Remove - 30% Confidence", although other suitable representations (e.g., shape, color, image, video, or combinations thereof) may be used.

[0072] In some embodiments, at least one processor of the computing device 160 may be configured to apply one or more neural networks to determine confidence values ​​associated with medical guidelines. For example, as part of or after training of the neural network, a validation phase may be used to associate confidence values ​​with medical guidelines. This phase may use an independent labeled dataset to obtain a correspondence between confidence score thresholds from the neural network and performance of the task of interest measured according to multiple metrics such as accuracy, precision, recall, sensitivity, npv, specificity, etc. This approach allows obtaining a correlation between predefined thresholds of confidence scores output by the neural network and predicted performance in real case scenarios.

[0073] In embodiments in which the characterization network determines a classification, a body location, and / or a size of the object of interest, the at least one processor may be configured to identify a medical guideline based on one or more of the classification, location, and size. For example, in embodiments in which the object of interest is a polyp, the medical guideline may be to leave the polyp when the polyp is classified as hyperplastic and to remove the polyp when the polyp is classified as dysplastic or neoplastic. Similarly, the medical guideline may be to leave the polyp when the size of the polyp is determined to be 5 mm or less or "small" and to remove the polyp when the size of the polyp is determined to be greater than 5 mm or "small" or "large". Similarly, the medical guideline may be to leave the polyp when the polyp is determined to be harmless in its body location and to remove the polyp when the polyp is determined to be in a harmful body location. In some embodiments, the medical guideline may be determined based on a combination of the classification, location, and size. For example, the medical guideline may be to leave the polyp when the polyp is determined to be a hyperplastic polyp 5 mm or less in size located in the rectum sigma. As another example, the medical guideline may be to leave the polyp when it is determined to be a hyperplastic polyp with a "small" size located in the rectum. On the other hand, the medical guideline may be to remove the polyp when it is determined to be an adenoma with a "small" size located in the cecum. Similarly, the medical guideline may be to remove the polyp when it is determined to be an adenoma with a "non-small" size located in the ascending colon. In some embodiments, a confidence value associated with the associated characteristic evaluation may be used to determine the medical guideline. For example, the medical guideline may be to remove the polyp only when there is a 90% confidence value that the polyp is neoplastic, 10 mm or "large" in size, and / or in a harmful location in the body, and to leave the polyp otherwise. The confidence value may be expressed as "high confidence" or "low confidence."However, as noted above, other confidence values ​​and characteristic assessments may be used. The medical guideline determinations disclosed herein are provided for illustrative purposes only and are not intended to be exhaustive, and other medical guidelines will be apparent to those of skill in the art.

[0074] FIG. 4A illustrates an exemplary system for processing real-time video according to an embodiment of the present disclosure. As shown in FIG. 4A, the real-time processing system 400 may include an imaging device 410, an object detector 420, a characterization network 430, and a display device 470. The imaging device 410 may be the same as or similar to the imaging device 140 described above in connection with FIG. 1 (e.g., an endoscopic device, an X-ray device, a CT device, an MRI device, or other medical imaging device), and the display device 470 may be the same as or similar to the display device 180 described above in connection with FIG. 1 (e.g., an LCD, LED or OLED display, an augmented reality display, a virtual reality display, etc.). The imaging device 410 may be configured to capture real-time video. The imaging device 410 may be configured to capture real-time video, which may be captured during a medical procedure (e.g., an endoscopic procedure) in some embodiments as described above. The imaging device 410 may be configured to provide the captured real-time video to the object detector 420.

[0075] The object detector 420 may comprise one or more machine learning detection networks or algorithms, traditional detection algorithms, or a combination of both, as described above in connection with the embodiment of FIG. 1 and computing device 160. The object detector 420 may be configured to detect an object of interest, such as a polyp, in a frame of real-time video captured by the imaging device 410. The object detector 420 may be configured to output multiple detections when more than one object of interest is present in a frame of real-time video. The object detector 420 may be configured to output any detections to the characterization network 430.

[0076] The characterization network 430 may include one or more trained machine learning algorithms (e.g., one or more neural networks) configured to determine a plurality of features for each of the objects of interest detected by the object detector 420, as described above in connection with the embodiment of FIG. 1 and the computing device 160. As shown in FIG. 4A, the characterization network 430 may include trained machine learning networks or algorithms configured to detect specific features, such as a classification network 440 configured to determine a classification of the detected objects of interest, a location network 450 configured to determine a body location of the detected objects of interest, and a size network 460 configured to determine a size of the detected objects of interest. However, it should be understood that the characterization network 430 may be modified to include all, some, or more of the machine learning networks shown in FIG. 4A, or to include none of the machine learning networks shown in FIG. 4A. For example, in some embodiments, the characterization network 430 may be a single network configured to determine all features of interest, or the characterization network 430 may include machine learning networks for determining features other than classification, location, or size, depending on the particular application or situation. Characterization network 430 may include a trained neural network for determining confidence values ​​and / or medical guidelines as disclosed herein. Additionally, characterization network 430 may include one or more processors (similar to computing device 160) for generating an augmented video stream for display to a clinician or operator.

[0077] In some embodiments, characterization network 430 may be implemented with multiple trained neural networks (e.g., networks 440, 450, and 460 and / or networks for determining confidence values ​​and medical guidelines) arranged to operate in parallel with one another to more efficiently determine characteristics and other information associated with objects identified from real-time video or frames. For example, the multiple trained neural networks may include at least one trained neural network for determining a classification of an object (i.e., classification network 440), at least one trained neural network for determining a location of an object (i.e., location network 450), at least one trained neural network for determining a size of an object (i.e., size network 460), and a trained neural network for determining a confidence value and / or medical guidelines. By optimizing the arrangement of neural networks (e.g., networks 440, 450 and 460 and / or networks for determining confidence values ​​and medical guidelines) trained to operate simultaneously in parallel with one another, it becomes possible to efficiently determine all information associated with each of the identified objects, enabling the characterization network 430 to generate an augmented video display including the determined information associated with the detected objects with little or no perceived delay to a clinician or operator viewing the augmented video display on the display device 470 while performing a medical procedure using the imaging device 410.

[0078] In some embodiments, the classification network 440, location network 450, and size network 460 may be implemented to simultaneously provide multiple output values ​​for each identified object. For example, the classification network 440 may be configured to provide as outputs an optical characterization prediction (e.g., adenoma, hyperplasia, SSL, etc.) as well as a morphology estimate (e.g., sessile, pedunculated, etc.) and a pit pattern description (Type I, Type II, Type III, etc.). This may be achieved through branching in the final layer of the neural network of the classification network 440, where each branch infers a particular classification. Additionally or alternatively, multiple instances of the networks 440-460 may be created to operate simultaneously in parallel to determine multiple characteristics or features for one or more detected objects.

[0079] For each identified object, a classification network 440, a location network 450, and a size network 460 may be implemented to process the entire frame and / or image patches around the object of interest. The output of each of the networks 440, 450, and 460 may be one or more predictions for a given class or characteristic or estimates (regressions) and may have a confidence score, as disclosed above in connection with the embodiment of FIG. 1 and the computing device 160. Additionally, each of the networks 440, 450, and 460 may receive as input the estimated data in the current frame or a buffer of N items from past frames. Optionally, each network may also build an internal representation that stores information from past frames (e.g., a recurrent neural network (RNN) implementation). The training of these neural networks may be based on annotated data with ground truth labels. In some embodiments, the ground truth values ​​may have confidence or uncertainty values ​​that can be used during training of the neural network.

[0080] In the example of FIG. 4A, the characterization network 430 may be configured to identify and output information related to the determined features and / or associated confidence values ​​to the display device 470. For example, the characterization network 430 may generate an augmented video stream (as described herein and with reference to other figures herein) for display on the display device 470. In some embodiments, the characterization network 430 may be configured to identify and output medical guidelines based on the determined features and / or associated confidence values ​​of each of the identified objects of interest, as described above. The display device 470 may then present information related to the features and / or medical guidelines in real-time during capture (e.g., during a medical procedure). For example, the characterization network 430 may determine classification information (e.g., "adenoma" or "non-adenoma" or "hyperplasia" or "non-hyperplasia"), location information (e.g., "rectum" or "appendicum"), size information (e.g., "small" or "not small" or "small" or "large") and / or medical guidelines (e.g., "leave" or "resection" or "biopsy") of the detected objects of interest, which the display element location 470 may display. In situations where multiple objects of interest are present in multiple frames, the characterization network 430 may be configured to generate and output characterization and / or medical guideline information for each individual object of interest (e.g., as a feature vector or another suitable data model), and the display 470 may consequently display the characterization and / or medical guideline information for each individual object of interest (e.g., on or near each object of interest).

[0081] The real-time processing system 400 may receive frames of video from the imaging device 410, process them, and provide an augmented video display with relevant information about objects of interest identified in the frames to an operator of the imaging device 410 in real-time (i.e., simultaneously or nearly simultaneously as the physician or operator performs the medical procedure). As disclosed herein, the characterization network 430 of the real-time processing system 400 may be optimized by deploying neural networks (e.g., networks 440, 450, and 460 and / or networks for determining confidence values ​​and medical guidelines) that are trained to process frames in parallel (or nearly simultaneously) and provide output efficiently. With such a configuration, an augmented video display with all determined information may be presented with little or no delay to a clinician or operator performing an endoscopy or other medical procedure.

[0082] In some embodiments, the object detector 420 and / or the characterization network 430 may run multiple instances of the machine learning detection network in parallel on multiple frames of real-time video from the image device 410. Additionally or alternatively, frames may be buffered for processing by the trained neural network. For example, the real-time processing system 400 may buffer frames from the image device 410 and provide them as input to the neural network. At each iteration, the networks 440-460 may receive as input the N image frames buffered by the system 400. In some embodiments, the networks 440-460 process the current frame along with the past N-1 image frames, and the output of each of the networks for the current frame also depends on the past N-1 frames. This buffering implementation can provide real-time processing using one of three options: (i) there is no output for the first N-1 frames, or (ii) the output of the first N-1 frames depends only on the current frame (i.e., no buffering in the initial stage), or (iii) the output of the first N-1 frames depends only on the last frame and all previous frames available. Additionally, there may be other intervals during which one or more of the networks 440-460 provide no output. During intervals during which there is no output, the real-time processing system 400 may communicate the status of the system by displaying an appropriate message via the display device 470 (e.g., a status message such as "Processing," "Buffering," or "Analyzing").

[0083] In some embodiments, the real-time processing system 400 may determine the number of instances of trained neural networks to run in parallel based on operator input and / or relevant processing parameters (e.g., frame rate and / or frame buffer size of the video generated by the image device 410). The real-time processing system 400 may process only certain frames or regions of interest identified by the object detector 420 as containing objects. Additionally, the real-time processing system 400 may selectively process frames and regions of frames based on available system resources and / or performance requirements. Alternatively or additionally, in other embodiments, the real-time processing system 400 may adjust the size of the input object identified by the object detector 420 based on the neural network used to determine the characterization. The real-time processing system 400 may adjust the size of the identified object by adjusting the resolution of the image frame containing the identified object or the buffer size of the image frame.

[0084] In some embodiments, the real-time processing system 400 may control processing based on the number of frames and / or objects detected in the frames. For example, the real-time processing system 400 may adjust the processing of frames to keep up with the frame rate of the image device 410. In some embodiments, the real-time processing system 400 may adjust the size or length of a buffer to account for the frame rate of the image device 410. In some embodiments, the real-time processing system 400 may keep up with the frame rate by processing multiple frames in parallel. The real-time processing system 400 may determine the number of neural network instances to run in parallel based on the frame rate of the image device 410. The real-time processing system 400 may determine the number of neural network instances to run in parallel based on other real-time processing requirements or factors such as (one or more) processing time delays or limitations due to available system resources (e.g., available hardware and software resources) and accuracy requirements in detecting objects and features of detected objects of interest. The real-time processing system 400 may achieve real-time processing requirements by adjusting the sampling rate to select a subset of frames from the image device 410. Additionally or alternatively, the real-time processing system 400 may sample frames and / or persistent objects detected in received frames to meet real-time processing requirements.

[0085] In some embodiments, the real-time processing system 400 may skip execution of one or more trained neural networks in response to operator input or settings (such as a command to exclude confidence values ​​and / or medical guidelines and / or a command to select object characterization features to include in processing). Skipping frames for processing may occur when an identified object is missing across one or more frames. Additionally or alternatively, the real-time processing system 400 may skip or pause execution of one or more trained neural networks in response to the mode of operation of the imaging device 410 (e.g., cleaning vs. navigation) and / or the location of the imaging device 410 in the patient's body or organ during medical processing or its relative location. For example, the real-time processing system 400 may stop operation of one or more neural networks of the system when it determines that the endoscopic device is outside the patient's colon. Additionally or alternatively, the real-time processing system 400 may skip execution of one or more of the networks 440-460 based on an action on an object or the status of the object detector 420. For example, the real-time processing system 400 may disable operation of the neural networks 440-460 during resection / surgery of an object identified by the object detector 420 or while the operator performs other tasks on the object of interest (e.g., injecting a lesion). While the operation of one or more networks 440-460 is disabled, the real-time processing system 400 may continue to receive information regarding detected objects and / or features of interest from the object detector 420 and / or other system components (e.g., frame quality network, object tracker, aggregator, etc.; see FIG. 6 and other embodiments disclosed herein).

[0086] 4A , one or more computing devices (such as computing device 160) may be used to implement object detector 420 and characterization network 430. Such computing device(s) may have one or more processors and may be configured to modify video from the imaging device using augmented information, including the information described above, determined by object detector 420 and characterization network 430. The augmented video may then be provided to display 470 for viewing by an operator of imaging device 410 and other users.

[0087] As disclosed herein, the trained neural networks of the characterization network 430 may be configured to operate in parallel to more efficiently determine the characterization (e.g., classification, location, and size) of objects identified during a medical procedure. With reference to Figures 4B and 4C, another configuration and feature for optimizing real-time processing according to an embodiment of the present disclosure is disclosed. It should be understood that Figures 4B and 4C are non-limiting examples and may be modified to include other components and features disclosed herein (e.g., a frame quality network (see network 620 in Figure 6) in combination with the characterization network 430). Furthermore, the teachings of Figures 4B and 4C may be compatible with and / or combined with other teachings herein, such as the teachings of the embodiments described with reference to Figure 6 and / or other figures provided herein.

[0088] FIG. 4B illustrates another exemplary system for real-time processing according to an embodiment of the present disclosure. As shown in FIG. 4B, real-time processing system 480 includes an image device 410, an object detector 420, and a display device 470. These components may be implemented and configured similarly to the corresponding components described above for the embodiment of FIG. 4A. Additionally, characterization network 430 may have similar features as those described above with reference to FIG. 4A, but in the embodiment of FIG. 4B, characterization network 430 has several additional components and optimized features. For example, as shown, an encoder network 485 and latent representations 486 are provided in combination with trained neural networks including classification network 440, location network 405, and size network 460. Object detector 420 may send objects identified in a frame received from image device 410 to encoder network 485. Encoder network 485 may process entire frames including identified objects and / or patch regions around each of the objects of interest in the frame determined by object detector 420. Additionally or alternatively, in some embodiments, the encoder network 485 may receive as input for processing information related to a bounding box or region surrounding a detected object, which may be represented as a list of numbers or a mask. The encoder network 485 may encode the input information (e.g., one or more image frames or surrounding regions representing the detected object of interest) and generate a feature vector representing the input information. The encoder network 485 may include one or more recurrent neural networks and / or long-short-term memory neural networks. The encoder network 485 may receive input information related to the detected object of interest and encode the information by generating a feature vector representation of the detected object(s). After encoding by the encoder network 485, the latent representation 486 may receive the feature vector representation as input and provide an embedding representation in the latent space as output.The latent representation 486 reduces data from high dimensionality (i.e., feature vector representation) to low dimensionality (i.e., latent space representation) while providing storage savings and is useful for computational efficiency in processing object data. As an example, the latent space representation may include an array of N floats, where N represents the number of frames or objects processed by the encoder network 485.

[0089] The encoder network 485 and the latent representation 486 may be realized with one or more neural networks trained using a combination of unsupervised reconstruction losses and supervised losses based on classification, location and size tasks. Additionally or alternatively, the encoder network 485 may be trained with losses from contrasting loss families, such as triplet or quartet losses, that reinforce the structured organization of the latent space. In this way, the latent space can assign similar representations to image frames belonging to the same object, and a more robust distance metric can be defined between the latent representations. As described above, the encoder network 485 can embed the inherent structure of the detected objects by projecting them into a latent space, e.g., the latent representation 486. The encoder network 485 may process the image frame(s) containing each of the detected objects or the surrounding region(s) to project them into the latent space by encoding layers of the network, together with the latent representation 486, resulting in a latent vector of lower dimensionality than the detected object(s) in the processed frame or its surrounding region. This provides several advantages, including reduced storage requirements and increased processing efficiency for object data, as described above.

[0090] In some embodiments, a tracking module (not shown in FIG. 4B, but see tracker 497 in FIG. 4C and further description thereof) may exploit and benefit from the improved latent representation in associating image frames belonging to the same object. To learn the most effective latent space, multiple losses may be combined and their relative magnitudes adjusted during training of the tracking module's neural network(s). For example, in the first part of learning, reconstruction and contrast losses may be large and act as regularization terms to prevent overfitting, and then the weights of specific losses for the task of interest may be increased at later stages during training to maximize their performance.

[0091] For each of the objects, the embedding representation in the latent space (i.e., the output of the latent representation 486) may be fed in parallel to three characterization networks (i.e., classification network 440, location network 405, and size network 460) to determine the features or characteristics of the object. Advantageously, in this implementation, the trained neural networks 440-460 are small (i.e., only a few fully connected layers) since the encoding part is shared and performed in the encoder network 485. This reduces the overall computational cost and the efficiency of the characterization network 430. As a result, the real-time processing system 480 benefits from a reduction in the time required to process and characterize the object of interest and provide an output to the display device 470.

[0092] FIG. 4C illustrates another example of a system for real-time processing according to an embodiment of the present disclosure. The real-time processing system 490 may be similarly constructed using the components and features described above for the real-time processing system 480 (FIG. 4B), but further includes a tracker 497 implemented with one or more neural networks that utilize temporal information. More specifically, in the embodiment of FIG. 4C, the neural networks 440-460 of the real-time processing system 490 may be networks that receive as input a buffer of consecutive frames of the same object identified by the object detector 420. Advantageously, the real-time processing system 490 may utilize the temporal information by using a recurrent neural network (RNN) and / or a long short-term memory (LSTM) neural network that maintains a memory of the same object detected in past frames by the object detector 420. The real-time processing system 490 may continue to track objects previously identified in past frames using the tracker 497. In some embodiments, the tracker 497 may perform a tracking operation by associating currently detected objects of a current set of frames with objects detected in previous frames. The tracker 497 may perform object tracking operations after computing a latent representation 496 of the object's encoding information such that the box information of each detected object is associated with its past history. The tracker 497 may exploit the similarity of the object's latent representation 496 to track the object's history across the current frame and past frames. The tracker 497 may determine the similarity of the object's latent representation 496 between the current image frame and past image frames based on one or more metrics such as mean squared error (MSE), mean absolute error (MAE), etc. As a result of the operations performed by the tracker 497, the characterization may be performed by the network 440-460, which may exploit the temporal information of each of the objects to accurately associate the objects with the past history of the frames, thereby more accurately determining the object's characteristics (i.e., classification, location, and size) and providing an improved confidence value.

[0093] The characterization network 430 may help determine features of an object of interest in frames accessed directly from the image device 410 or a storage device (e.g., storage device 220 or a buffer device with memory) that contains previously generated image frames. The characterization network 430 allows for the inclusion of various networks for simultaneously determining various features of an object and optimizing the processing of the image frames when determining the features.

[0094] The feature network 430 may be composed of multiple networks and may be configured to select and deselect various configurations of the network to identify objects of interest in an image frame and determine its properties in an optimized manner. The feature network 430 may be configured to have multiple copies of the same network to process image frames or patches of image frames in parallel to determine properties. In some embodiments, the feature network 430 optimizes frame processing by configuring the order of networks to preprocess image frames by selecting the required networks and performing common operations across the networks. For example, the feature estimation network 430 may utilize the encoder network 485 and latent representation 486 to preprocess the image and may have smaller networks to efficiently determine properties while using fewer computational resources.

[0095] 5A-5E show example augmented video frames containing characterization and medical guideline information according to embodiments of the present disclosure. The augmented frames shown in FIGS. 5A-5E may be displayed in real-time during video capture (e.g., during a medical examination), as described above. Although a polyp is shown in the example frames of FIGS. 5A-5E, it should be understood that the disclosed embodiments may be used with any other object of interest. As shown in FIGS. 5A-5E, the frames may include characterization information related to one or more characteristics of the polyp, such as its classification, size, and / or location. For example, as shown in FIG. 5A, the polyp may be classified as a non-adenomas and a corresponding representation such as "non-adenomas" may be displayed, or as shown in FIGS. 5B and 5C, the polyp may be classified as an adenoma and a corresponding representation such as "adenomas" may be displayed. Similarly, the polyp may be determined to have a small size and a corresponding representation such as "small" may be displayed, as shown in FIG. 5A and FIG. 5C, or the polyp may be determined to have a non-small size and a corresponding representation such as "not small" may be displayed, as shown in FIG. 5B. Similarly, the body location of the polyp may be determined and a representation such as "rectum" (FIG. 5A), "ascending" (FIG. 5B), "cecum" (FIG. 5C), "cecum" (FIG. 5C) and "sigma-rectum" (FIG. 5D) may be displayed, as shown in FIG. 5A-FIG. 5D. If the location cannot be determined, no location information may be displayed in the expansion frame. Other representations such as one or more images, icons, videos, shapes or numbers may be used according to the present disclosure.

[0096] Additionally, as described above, medical guidelines may also be displayed. As shown in FIG. 5A, for example, a medical guideline based on classification, size, and / or location may be to not remove the polyp and may display a corresponding expression such as "leave". However, as shown in FIG. 5B and FIG. 5C, a medical guideline may be to remove the polyp (or perform other procedures or treatments) and may display a corresponding expression such as "remove". Furthermore, characterization and medical guideline information may be displayed (and determined) independently for each of the multiple polyps in the frame. As shown in FIG. 5C, for example, separate classification and size information may be displayed adjacent to each detected polyp, but this may vary depending on the situation. Additionally, other information besides characterization and medical guideline information may be displayed. For example, as shown in FIG. 5D and FIG. 5E, the status of the characterization network and / or computer device may be displayed, such as whether the characterization network and / or computer device is currently "analyzing" the frame (FIG. 5D), "no prediction" or there is no characterization yet (shown as a blank field in the detected object in the center of the frame in FIG. 5E). Examples of other information displayed include the characterization network and / or computer device performing partial characterization of the polyp, confidence values ​​that are too low, errors that have occurred, system restarts, user actions or inputs, or other relevant information related to the processing of the real-time video.

[0097] In some embodiments, additional modules or processes may be provided and executed before, after, or simultaneously with the characterization network 430. For example, in some embodiments, at least one processor may be configured to apply one or more neural networks implementing a trained quality network configured to determine a frame quality associated with at least one of the plurality of frames. As used herein, "frame quality" may refer to the visual clarity of one or more frames to perform the operations described herein. The frame quality may be based on visual characteristics such as blur, sharpness, brightness, lighting, exposure, contrast, motion, visibility, or other characteristics of one or more frames. The trained frame quality network may be trained to generate a numeric value associated with the frame quality and / or a quality classification. For example, the trained frame quality network may be configured to output a numeric value associated with the frame quality (e.g., 0.7) or may be configured to assign a quality class (e.g., "sufficient quality" or "inadequate quality") to the frame.

[0098] The trained frame quality network may include one or more suitable machine learning networks or algorithms for determining a quality value associated with one or more frames of a real-time video, including one or more neural networks (e.g., deep neural networks, convolutional neural networks, recurrent neural networks, etc.), random forests, support vector machines, or any other suitable models trained above for determining frame quality. The frame quality network may be trained using a plurality of training frames or portions thereof that are labeled based on one or more quality values ​​or classifications. For example, a first set of training frames (or portions of frames) may be labeled as "sufficient quality" and a second set of training frames (or portions of frames) may be labeled as "not sufficient quality." Weights and other parameters of the frame quality network may be adjusted based on output for a third set of unlabeled training frames (or portions of frames) until convergence or other metrics are achieved, and the process may be repeated using additional training frames (or portions thereof) or using live data as described herein. The trained frame quality network may be stored in the computing device and / or system or may be fetched from a network or database prior to determining frame quality. In some embodiments, the trained frame quality network may be retrained based on one or more of its outputs, such as accurate frame quality detection or inaccurate frame quality detection. Feedback for retraining may be generated automatically by the system or computing device or may be manually entered by an operator or another user (e.g., via a mouse, keyboard, or other input device). Weights or other parameters of the trained frame quality network may be adjusted based on the feedback. In some embodiments, conventional non-machine learning frame quality detection networks or algorithms may be used alone or in combination with the trained frame quality network.

[0099] In some embodiments, the trained frame quality network may be configured to generate a confidence value associated with the determined frame quality. The confidence value for the determined frame quality may refer to an indication of a level of confidence associated with the determined frame quality. For example, a confidence value of 0.9 or 90% indicates that there is a 90 percent certainty that the determined frame quality is correct, while a confidence value of 0.4 or 40% indicates that there is a 40 percent certainty that the determined frame quality is correct, and so on. However, other values, metrics, or representations may be used to represent the confidence value. For example, the confidence value may be classified as "high confidence" when the confidence value is above a predetermined threshold (e.g., above 66% confidence), as "low confidence" when the confidence value is below a predetermined threshold (e.g., below 33% confidence), or as "undecided" when the confidence value is between two predetermined thresholds (e.g., the confidence is between 33% and 66% confidence). In some embodiments, the at least one processor may present information regarding the determined frame quality in real time on a display device in any suitable format (e.g., a frame quality value and / or classification, a thumbs up or thumbs down, a check mark or cross, or a color).

[0100] At least one processor of the computing device or system may be configured to aggregate data related to the determined characterization (e.g., classification, location, and / or size) when at least one of the determined frame quality or confidence value exceeds a predefined threshold, according to an embodiment of the present disclosure. Aggregation in this context may refer to any operation for combining, collecting, or receiving a plurality of data. For example, in an embodiment in which the classification, location, and size of an object of interest of a frame are determined, the at least one processor may be configured to collect the classification, location, and size determination from the characterization network only when the frame quality from the frame is determined to be above a predefined threshold (e.g., the frame quality is greater than 0.4 or classified as "sufficient quality"). In some embodiments, the at least one processor may be configured to present at least a portion of the aggregated data on a display device. The aggregated data may be displayed in the same or similar manner as described above (e.g., using an LED display, a virtual reality display, or an augmented reality display).

[0101] Other information, such as other determined features, may be aggregated, and other metrics may be used to determine whether to aggregate data depending on the particular application or situation. In some embodiments, for example, the at least one processor may be configured to aggregate information associated with the determined features (e.g., location and size) of the object of interest when the object of interest persists across two or more of the multiple frames. As used herein, "persistence" or variations thereof may refer to the object of interest continuing to be present at a location in one or more frames. Persistence may be determined using any process for comparing the presence of the object of interest in one or more frames. For example, an Intersection over Union (loU) value of the location of the object of interest in two or more image frames may be calculated, and the loU value may be compared to a threshold value to determine whether the object of interest persists across two or more of the frames. The loU value may be estimated using the following formula:

[0102]

number

[0103] In the above formula, Area of ​​Overlap is the area where the object of interest is present in two or more frames, and Area of ​​Union is the total area where the object of interest is present in two or more frames. As a non-limiting example, an loU value of more than 0.5 (e.g., about 0.6 or 0.7 or more, such as 0.8 or 0.9) between two consecutive frames may be used to determine that the object of interest persists in two consecutive frames. In contrast, an loU value of less than 0.5 (e.g., about 0.4 or less) between two consecutive frames may be used to determine that the object of interest does not persist. However, other methods of determining persistence may be used depending on the application and circumstances. When an object is determined to persist beyond multiple frames, information related to the determined characteristics (e.g., location and size) of the object of interest may be aggregated in the same or similar manner as described above. The at least one processor may be configured to present the aggregated data or a portion thereof on a display device. In this manner, only information about the object of interest that is sufficiently present in two or more frames may be displayed to avoid displaying unnecessary or distracting information during capture (e.g., during a medical procedure).

[0104] In some embodiments, the aggregated information may be displayed based on one or more criteria. For example, the at least one processor may be configured to present on the display device aggregated information (e.g., location and size) of the object of interest when the determined location is in a first body region and the determined size is within a first range. Additionally, the at least one processor may be configured to present on the display device information indicative of a status of a characterization of the object of interest (i.e., non-aggregated information) when the determined location is in a second body region and the determined size is within a second range. As a non-limiting example, in an embodiment where the object of interest is a polyp and the aggregated information is location and size, the classification information (i.e., non-aggregated information) may be displayed when the polyp is determined to be small and located in a body location other than the sigma rectum. Conversely, the location and size (i.e., aggregated information) may be displayed when the polyp is determined to be not small or located in the sigma rectum. In this manner, only relevant information may be displayed to an operator based on predefined aggregation and / or display criteria to provide only important information during capture (eg, during a medical procedure).

[0105] Figure 6 illustrates another exemplary system 600 for processing real-time video in accordance with disclosed embodiments. As shown in Figure 6, the real-time processing system 600 may include an object detector 610, a frame quality network 620, a characterization network 630, a tracker 640, an aggregator 650, and a display device 660. The components (610, 620, 630, 640, 650) of Figure 6 may be implemented using a computing device(s) or one or more processors. The object detector 610 may be the same as or similar to the object detector 420 described above in connection with FIG. 4A (e.g., a machine learning detection network of algorithms or a conventional detection algorithm), the characterization network 630 may be the same as or similar to the characterization network 430 described above in connection with FIG. 4A (the characterization network 630 may include more or fewer networks), and the display device 470 may be the same as or similar to the display device 180 shown in FIG. 4A (e.g., an LCD, LED or OLED display, an augmented reality display, or a virtual reality display). Comparing FIG. 4A with FIG. 6, it can be seen that additional operations may be performed before, after, or simultaneously with the characterization network 630. For example, in FIG. 6, the frame quality network 620, the tracker 640, and the aggregator 650 may perform operations at any time in conjunction with the operation of the characterization network 630. Although arrows are used in FIG. 6 and in other figures to indicate the general flow of information in some embodiments, it should be understood that operations may be performed in a different order or simultaneously with one or more other operations and that other steps may be added or skipped entirely.

[0106] 6, object detector 610 may detect one or more objects of interest in multiple frames and send its output to frame quality network 620 and / or characterization network 630. Although frame quality network 620 is shown as two separate components in some embodiments, frame quality network 620 may be part of characterization network 630. Frame quality network 620 may be configured to determine frame quality and / or a confidence value associated with the frame quality of frames that include the detected object of interest and send its output to characterization network 630 and / or tracker 640, as described above.

[0107] In some embodiments, the frame quality network 620 may classify frames from the imaging device and / or patches of frames containing objects of interest detected by the object detector 610 based on a set of classes learned in training. The frame quality network 620 may output its output confidence values ​​by learning implicitly during training, for example, using a one-hot encoding formula. Alternatively, a dedicated output node may be added to the frame quality network 620 to provide confidence estimates and abstraction terms included in the loss function for training the network. The frame quality network 620 may use abstraction terms to predict low confidence values, for example, when the estimation error is high due to low quality or cluttered image frames. In some embodiments, the frame quality network 620 may learn to generate its output confidence values ​​explicitly when ground truth confidence values ​​are available for the training data used to train the frame quality network 620. In some embodiments, the frame quality network 620 may use neural network calibration methods such as mix-up and label smoothing during training to control the range and distribution of confidence values. One or more neural networks may be used to implement the quality network 620 and, where possible, may be trained to predict both the output and the confidence values ​​or label uncertainties.

[0108] The characteristic evaluation network 630 may identify one or more features and / or confidence values ​​associated with one or more features of the detected object of interest, as described above. In embodiments in which the characteristic evaluation network 630 receives frame quality and / or a confidence value associated with the frame quality from the frame quality network 620, the characteristic evaluation network 630 may be configured to detect (or provide an output for) features of the object of interest where the frame quality and / or a confidence value associated with the frame quality is above a predefined threshold (e.g., greater than 0.4 or classified as "sufficient quality"). The tracker 640 may be configured to determine persistence of the detected object of interest, as described above. In some embodiments, the tracker 640 may receive feature detections from the characteristic evaluation network 630 and may be configured to provide the feature detections as an output only upon determining that the detected object of interest persists across multiple frames or a predefined number of frames. In some embodiments, tracker 640 may be configured to receive frame quality and / or a confidence value associated with the frame quality, and tracker 640 may use that persistence determination to determine whether to provide an output (e.g., tracker 640 may provide an output only when a detected object of interest persists across multiple frames or a predetermined number of frames and the frame quality and / or confidence value is above a predetermined threshold).

[0109] The aggregator 650 may receive output from any of the components described above, including the object detector 610, the frame quality network 620, the characterization network 630, and / or the tracker 640, and may aggregate or combine any of the received information based on one or more criteria for presentation to the display 660, as described above. For example, the aggregator 650 may receive information related to features detected by the characterization network 630 to determine which features to aggregate according to predefined criteria. For example, the aggregator 650 may output aggregated position and size information of the object of interest to the display 660 when the determined position is within a first body region and the determined size is within a first range. Additionally or alternatively, the aggregator 650 may instead output information indicative of a status of characterization of the object of interest to the display 660 when the determined position is within a second body region and the determined size is within a second range. As an example, the status of the characterization of the object of interest provided by aggregator 650 may include an aggregation status to inform the operator or user whether there is aggregated information or whether there is only non-aggregated information. Aggregator 650 may use other rules and criteria for outputting information for presentation on display device 660, as described above.

[0110] 7 illustrates a block diagram of an exemplary method 700 for aggregating information for presentation on a display according to an embodiment of the present disclosure. The exemplary method 700 may be performed using one or more processors. In one embodiment, the method 700 may be performed using a computing device or system such as the system 600 of FIG. 6. It is understood that this is a non-limiting example.

[0111] As shown in Figure 7, in step 710, a frame quality network may be applied to frames with or without an object of interest to determine frame quality and a confidence value associated with the frame quality, as described above. In step 720, the frame quality and confidence value may be compared to a threshold to determine whether the frame has sufficient frame quality. If the frame has sufficient frame quality, the method proceeds to aggregate the data. If the frame does not have sufficient frame quality, the method returns to step 710 so that another frame can be examined.

[0112] In steps 730, 740, and 750, the classification network, location network, and size network may be applied to determine the classification, location, and size, respectively, of the detected object of interest, as described above. In step 760, if the frame has sufficient frame quality, at least a portion or set of the classification, location, and size information is aggregated. In step 770, one or more criteria may be applied to determine which portion of the aggregated data to present for display (e.g., whether the determined location is in a first body region and the determined size is within a first range, or whether the determined location is in a second body region and the determined size is within a second range), as described above. If a first set of criteria is met, then in step 780, a first portion of the aggregated information may be displayed. If a second set of criteria is met, then in step 790, a second portion of the aggregated information may be displayed. Although not shown in FIG. 7, additional criteria or combinations of criteria may be applied in step 770 to determine whether to display information as part of an extended display.

[0113] 8 illustrates a block diagram of an exemplary method 800 for processing real-time video according to an embodiment of the present disclosure. The exemplary method 800 may be performed using a computing device(s) or one or more processors. In an embodiment, the method 800 may be performed using a computing device or system such as the system 100 of FIG. 1 or the system 400 of FIG. 4A. It is understood that these are non-limiting examples.

[0114] FIG. 8 includes steps 801-813. In step 801, the at least one processor may receive real-time video captured from a medical imaging device during a medical procedure, the real-time video including a plurality of frames, as described above. In step 803, the at least one processor may detect an object of interest in the plurality of frames, as described above. In block 805, the at least one processor may apply one or more neural networks implementing a trained classification network configured to determine a classification of the object of interest, as described above. In block 807, the at least one processor may apply one or more neural networks implementing a trained location network configured to determine a location associated with the object of interest, as described above. In block 809, the at least one processor may apply one or more neural networks implementing a trained size network configured to determine a size associated with the object of interest, as described above. In block 811, the at least one processor may identify a medical guideline based on one or more of the classification, location, and size, as described above. In block 813, the at least one processor may present information of the identified medical guideline on a display device in real time during a medical procedure, as described above.

[0115] The above-mentioned figures and elements of the figures illustrate the architecture, functionality and operation of possible implementations of systems, methods and computer hardware or software products according to various exemplary embodiments of the present disclosure. For example, each block of the flowchart or figures may represent a module, segment or part of code including one or more executable instructions for implementing a specified logical function. It should also be understood that in some alternative implementations, the functions shown in the blocks may be executed in a different order than that shown. For example, two blocks shown in succession may be executed or implemented substantially simultaneously, or the two blocks may be executed in reverse order depending on the functions involved. Some blocks may also be omitted. It should also be understood that each block and combination of blocks in the figures may be implemented by a dedicated hardware-based system that executes the specified functions or operations, or by a combination of dedicated hardware and computer instructions. A computer program product (e.g., software or program instructions) may be implemented based on the described embodiments and illustrated examples.

[0116] It should be understood that the above-described systems and methods can be modified in many ways, including omitting or adding steps, changing the order of steps and the types of functions and / or components used. It should also be understood that different features can be combined in different ways. In particular, not all features described above in a particular embodiment or implementation are required in all embodiments or implementations. Other combinations of the above-described features and implementations are considered to be within the scope of the embodiments or implementations disclosed herein.

[0117] While certain embodiments and features of implementations have been described and illustrated herein, modifications, substitutions, variations, and equivalents will be apparent to those skilled in the art. It should therefore be understood that the appended claims are intended to cover all such modifications and variations that fall within the scope of the disclosed embodiments and features of the illustrated implementations. It should also be understood that the embodiments described herein are presented by way of example only and not by way of limitation, and that various changes in form and details are possible. Any part of the systems and / or methods described herein may be implemented in any combination, except in mutually exclusive combinations. By way of example, the implementations described herein may include various combinations and / or subcombinations of the functions, components and / or features of the various embodiments described.

[0118] Furthermore, although exemplary embodiments have been described herein, the scope of the disclosure includes any embodiment having equivalent elements, modifications, omissions, combinations, adaptations, or variations based on the embodiments disclosed herein (e.g., aspects across various embodiments). Furthermore, the elements of the claims should be interpreted broadly based on the language used in the claims, and not limited to the examples described in the specification or during the practice of this application. Instead, these examples should be interpreted as non-limiting. Furthermore, the steps of the disclosed methods can be modified in any manner, including changing the order of steps or inserting or deleting steps.

[0119] As another example, the systems and methods according to the present disclosure include the following implementations and aspects.

[0120] 1. A computer-implemented system for processing real-time video, comprising at least one processor, the at least one processor configured to receive real-time video captured from a medical imaging device during a medical procedure, the real-time video comprising a plurality of frames, detect objects of interest in the plurality of frames, encode the objects of interest in the plurality of frames to generate embedded representations using an encoder network that processes regions surrounding the objects of interest, generate latent representations of the encoded objects of interest, apply one or more neural networks implementing a trained property estimation network to the latent representations to determine one or more properties of the objects of interest, modify the real-time video using augmented information of the detected objects of interest and the one or more properties of the objects of interest, and present the modified video on a display device during the medical procedure.

[0121] In the above-mentioned system, the medical procedure may include at least one of an endoscopy, a gastroscopy, a colonoscopy, or an intestinal endoscopy. Further, the object of interest may include at least one of a formation of human tissue, a change in human tissue from one type of cell to another type of cell, an absence or lesion of human tissue from a location where human tissue is expected.

[0122] In the above-described system, the at least one processor may be further configured to identify a medical guideline based on one or more characteristics of the object of interest. The at least one processor may be configured to present information related to the identified medical guideline on the display device as part of the modified video. As one example, the information related to the identified medical guideline may include instructions to leave or resect the object of interest. As another example, the identified medical guideline information may include a type of resection.

[0123] In the above-described system, the at least one processor may be further configured to generate a confidence value associated with the identified medical guideline.

[0124] In the above-mentioned system, the determined one or more characteristics of the object of interest may include a classification of the object of interest based on at least one of a histological classification, a morphological classification, a structural classification, or a malignancy classification. As another example, the determined one or more characteristics of the object of interest may include a location associated with the object of interest. In some embodiments, the determined location is a location of a human body or a location associated with a human organ. Examples of determined locations may include the rectum, the sigmoid colon, the descending colon, the transverse colon, the ascending colon, or the cecum. As yet another example, the determined one or more characteristics of the object of interest may be a size of the object of interest. The determined size of the object of interest may be expressed as a numerical value or a size classification.

[0125] In the above system, the at least one processor may be further configured to apply one or more neural networks implementing a trained quality network, the trained quality network may be configured to determine a frame quality associated with one or more of the plurality of frames and to generate a confidence value associated with the determined frame quality.

[0126] In the above-described system, the at least one processor may be further configured to aggregate data associated with the determined one or more characteristics when at least one of the determined frame quality or confidence is above a predefined threshold, and present at least a portion of the aggregated data on a display device.

[0127] In the above-described system, the at least one processor may be further configured to detect a plurality of objects of interest in a plurality of frames, determine a plurality of classifications and sizes associated with the detected plurality of objects of interest, and present information associated with one or more of the determined classifications and sizes on a display device.

[0128] In the above-described system, the at least one processor may be further configured to track an object of interest in the multiple frames to determine temporal information associated with the object of interest.

[0129] A computer-implemented method for processing real-time video, comprising: receiving real-time video captured from a medical imaging device during a medical procedure, the real-time video including a plurality of frames; detecting an object of interest in the plurality of frames; encoding the object of interest in the plurality of frames to generate an embedded representation using an encoder network that processes an area surrounding the object of interest; generating a latent representation of the encoded object of interest; applying one or more neural networks implementing a trained property estimation network to the latent representation to determine one or more properties of the object of interest; modifying the real-time video using augmented information of the detected object of interest and the one or more properties of the object of interest; and presenting the modified video on a display device during the medical procedure.

[0130] 1. A computer-implemented system for processing real-time video, comprising at least one processor, the at least one processor being configured to receive real-time video captured from a medical imaging device during a medical procedure, the real-time video including a plurality of frames, detect an object of interest in the plurality of frames, encode the object of interest in the plurality of frames to generate an embedded representation using an encoder network that processes a region surrounding the object of interest, generate a latent representation of the encoded object of interest, track the object of interest in the plurality of frames based on the latent representation to determine temporal information of the object of interest, and apply one or more neural networks implementing a trained characteristic estimation network to determine one or more properties of the object of interest, the one or more properties being determined based on at least one of the latent representation and the temporal information.

[0131] In the above system, the at least one processor may be further configured to modify the real-time video using the detected object of interest and the augmented information of one or more properties of the object of interest and present the modified video on a display device during a medical procedure. The medical procedure may include at least one of an endoscopy, a gastroscopy, a colonoscopy, or an intestinal endoscopy. Further, the object of interest may include at least one of a formation of human tissue, a change in human tissue from one type of cell to another type of cell, an absence or lesion of human tissue from a location where human tissue is expected.

[0132] In the above-described system, the at least one processor may be further configured to identify a medical guideline based on one or more characteristics of the object of interest. The at least one processor may be configured to present information related to the identified medical guideline on the display device as part of the modified video. As one example, the information related to the identified medical guideline may include instructions to leave or resect the object of interest. As another example, the identified medical guideline information may include a type of resection.

[0133] In the above-described system, the at least one processor may be further configured to generate a confidence value associated with the identified medical guideline.

[0134] In the above-mentioned system, the determined one or more characteristics of the object of interest may include a classification of the object of interest based on at least one of a histological classification, a morphological classification, a structural classification, or a malignancy classification. As another example, the determined one or more characteristics of the object of interest may include a location associated with the object of interest. In some embodiments, the determined location is a location of a human body or a location associated with a human organ. Examples of determined locations may include the rectum, the sigmoid colon, the descending colon, the transverse colon, the ascending colon, or the cecum. As yet another example, the determined one or more characteristics of the object of interest may be a size of the object of interest. The determined size of the object of interest may be expressed as a numerical value or a size classification.

[0135] In the above system, the at least one processor may be further configured to apply one or more neural networks implementing a trained quality network, the trained quality network may be configured to determine a frame quality associated with one or more of the plurality of frames and to generate a confidence value associated with the determined frame quality.

[0136] In the above-described system, the at least one processor may be further configured to aggregate data associated with the determined one or more characteristics when at least one of the determined frame quality or confidence is above a predefined threshold, and present at least a portion of the aggregated data on a display device.

[0137] In the above-described system, the at least one processor may be further configured to detect a plurality of objects of interest in a plurality of frames, determine a plurality of classifications and sizes associated with the detected plurality of objects of interest, and present information associated with one or more of the determined classifications and sizes on a display device.

[0138] In the above system, the at least one processor may be further configured to track the object of interest in the plurality of frames to determine temporal information associated with the object of interest. To track the object of interest, the at least one processor may be configured to track the object of interest in the plurality of frames based on similarity of the latent representation of the object of interest in the plurality of frames. Additionally or alternatively, to track the object of interest, the at least one processor may be configured to determine a number of frames for which the object of interest persists. The at least one processor may be configured to track the object of interest based on frame quality information for each of the frames.

[0139] 1. A computer-implemented method for processing real-time video, comprising: receiving real-time video captured from a medical imaging device during a medical procedure, the real-time video including a plurality of frames; detecting an object of interest in the plurality of frames; encoding the object of interest in the plurality of frames to generate an embedded representation using an encoder network that processes an area surrounding the object of interest; generating a latent representation of the encoded object of interest; tracking the object of interest in the plurality of frames based on the latent representation to determine temporal information of the object of interest; and applying one or more neural networks implementing a trained property estimation network to determine one or more properties of the object of interest, the one or more properties being determined based on at least one of the latent representation and the temporal information.

[0140] In the above method, the method may further comprise modifying the real-time video using the detected object of interest and the augmented information of one or more properties of the object of interest, and presenting the modified video on a display device during a medical procedure. The medical procedure may include at least one of an endoscopy, a gastroscopy, a colonoscopy, or an intestinal endoscopy. Further, the object of interest may include at least one of a formation of human tissue, a change in human tissue from one type of cell to another type of cell, an absence or lesion of human tissue from a location where human tissue is expected.

[0141] It is therefore intended that the specification and examples be considered as exemplary only, with a true scope and spirit being indicated by the following claims and their full scope of equivalents.

Claims

1. A computer-executable system for processing real-time video, comprising at least one processor, wherein the at least one processor receives real-time video captured from a medical imaging device during a medical procedure, the real-time video including a plurality of frames, detects an object of interest in the plurality of frames, applies one or more neural networks that implement a trained classification network configured to determine a classification of the object of interest, a trained position network configured to determine a position associated with the object of interest, and a trained size network configured to determine a size associated with the object of interest, identifies medical guidelines based on two or more of the classification, the position, and the size, and is configured to present information of the identified medical guidelines in real time on a display device during the medical procedure.

2. The system according to claim 1, wherein the medical procedure includes at least one of an endoscopy, a gastroscopy, a colonoscopy, or an enteroscopy.

3. The system according to claim 1, wherein the object of interest includes at least one of a formation of human tissue, a change of human tissue from one type of cell to another type of cell, an absence of human tissue from a location where human tissue is expected, or a lesion.

4. The system according to claim 3, wherein the information of the identified medical guidelines includes an instruction to leave or excise the object of interest.

5. The system according to claim 4, wherein the information of the identified medical guidelines includes a type of excision.

6. The system according to claim 1, wherein the at least one processor is further configured to generate a confidence value associated with the identified medical guidelines.

7. The system according to claim 1, wherein the determined classification is based on at least one of a histological classification, a morphological classification, a structural classification, or a malignancy classification.

8. The system according to claim 1, wherein the determined position associated with the object of interest is a position in the human body.

9. The system according to claim 8, wherein the position in the human body is one of a position of the rectum, sigmoid colon, descending colon, transverse colon, ascending colon, or cecum.

10. The system according to claim 1, wherein the determined size related to the object of interest is a numerical value or a size classification. **Claim 11** The at least one processor applies one or more neural networks configured to implement a trained quality network that determines a frame quality related to at least one of the plurality of frames and generates a confidence value related to the determined frame quality. The system according to claim 1, further configured. **Claim 12** The at least one processor aggregates data related to the determined classification, the position, and the size when at least one of the determined frame quality or the confidence value exceeds a predetermined threshold, The system according to claim 11, further configured to present at least a portion of the aggregated data to the display device. **Claim 13** The at least one processor detects a plurality of objects of interest in the plurality of frames, determines a plurality of classifications and a plurality of sizes related to the plurality of objects of interest, and one classification and one size of the determined plurality of classifications and the plurality of sizes are related to one detected object of interest among the detected plurality of objects of interest, The system according to claim 1, further configured to present information related to one or more determined classifications and sizes to the display device. **Claim 14** The at least one processor applies one or more neural networks configured to implement an encoder network configured to encode the object of interest in the plurality of frames by processing an area surrounding the object of interest. The system according to any one of claims 1 to 13. **Claim 15** The system according to claim 14, wherein the at least one processor is further configured to generate a latent representation of the object of interest encoded by the encoder network. **Claim 16** The system according to claim 15, wherein the at least one processor is further configured to provide the latent representation of the object of interest using the classification network, the position network, and the size network. **Claim 17** The at least one processor The system according to claim 14, further configured to track the object of interest in the plurality of frames to determine time information of the object of interest.

18. A computer-executable system for processing real-time video, comprising a system having at least one processor, the at least one processor receives real-time video including a plurality of frames collected during a medical procedure, detects an object of interest in the plurality of frames, determines a plurality of features associated with the object of interest, applies one or more neural networks that implement a trained characteristic evaluation network configured to determine a confidence value associated with the plurality of features, identifies medical guidelines based on one or more of the plurality of features and the confidence value, A system configured to present information regarding the identified medical guidelines in real time to a display device during the medical procedure.

19. The system according to claim 18, wherein the medical procedure includes at least one of an endoscopy, a gastroscopy, a colonoscopy, or an enteroscopy.

20. The system according to claim 18, wherein the object of interest includes at least one of the formation of human tissue, the change of human tissue from one type of cell to another type of cell, the absence of human tissue from a location where human tissue is expected, or a lesion.

21. The system according to claim 20, wherein the information of the medical guidelines includes an instruction to leave or excise the object of interest.

22. The system according to claim 21, wherein the information of the identified medical guidelines includes the type of excision.

23. The system according to claim 18, wherein the at least one processor is further configured to generate a confidence value associated with the identified medical guidelines.

24. The trained characteristic evaluation network a trained classification network configured to determine a classification associated with the object of interest and generate a classification confidence value associated with the determined classification; a trained position network configured to determine a position associated with the object of interest and generate a position confidence value associated with the determined position; A trained size network configured to determine a size related to the object of interest and generate a size confidence value related to the determined size, The system according to claim 18, comprising:

25. The system according to claim 24, wherein the at least one processor is further configured to present information related to at least one of the classification, the position, or the size to the display device.

26. The at least one processor is Determine the frame quality related to at least one of the plurality of frames, The system according to claim 18, further configured to apply one or more neural networks that implement a trained quality network configured to generate a confidence value related to the determined frame quality.

27. The at least one processor is When at least one of the determined frame quality or the confidence value is above a predetermined threshold, aggregate data related to the plurality of features, The system according to claim 26, further configured to present at least a portion of the aggregated data to the display device.

28. The at least one processor is Detect a plurality of objects of interest in the plurality of frames, Determine a set of a plurality of features related to the plurality of objects of interest, wherein one set of features of the set of the plurality of features includes a characteristic evaluation and size information related to one detected object of interest among the plurality of detected objects of interest, The system according to any one of claims 18 to 27, further configured to present information related to one or more sets of features of the set of the plurality of features to the display device.

29. A computer-executable method for processing real-time video, comprising: Receiving real-time video captured from a medical imaging device during a medical procedure, the real-time video including a plurality of frames, Detecting an object of interest in the plurality of frames, A trained classification network configured to determine the classification of the object of interest, A trained position network configured to determine a position related to the object of interest, and Applying one or more neural networks configured to implement a trained size network configured to determine a size associated with the object of interest; Identifying medical guidelines based on two or more of the classification, the position, and the size; Presenting information of the identified medical guidelines in real time to a display device during the medical procedure; A method comprising the above.

30. A computer-executable system for processing real-time video, comprising at least one processor, the at least one processor: Detecting an object of interest in a plurality of frames received from a medical imaging device; Performing a characteristic evaluation of the object of interest, the characteristic evaluation comprising determining a plurality of characteristics associated with the object of interest, the plurality of characteristics including a position and a size of the object of interest; When the object of interest persists across two or more of the plurality of frames, aggregating information related to the determined position and size of the object of interest; Evaluating the determined position and size of the object of interest based on the aggregated information; When the determined position is in a first body region and the determined size is within a first range, presenting the aggregated information of the object of interest to a display device; A system configured to present information indicating a state of the characteristic evaluation of the object of interest to the display device when the determined position is in a second body region and the determined size is within a second range.

31. The at least one processor: Is further configured to identify medical guidelines based on the determined position and size of the object of interest; The system according to claim 30, further configured to present information related to the identified medical guidelines to the display device.

32. The system according to claim 30, wherein the plurality of characteristics further includes a classification of the object of interest based on at least one of a histological classification, a morphological classification, a structural classification, or a malignancy classification.

33. The system according to claim 30, wherein the determined position associated with the object of interest is at least one position of the rectum, the sigmoid colon, the descending colon, the transverse colon, the ascending colon, or the cecum.

34. The system according to claim 30, wherein the determined size related to the object of interest is a numerical value or a size classification.

35. The at least one processor detects a plurality of objects of interest in the plurality of frames, performs a characteristic evaluation of the plurality of objects of interest, the characteristic evaluation having determining a set of a plurality of features related to the plurality of objects of interest, and one set of features of the set of the plurality of features includes a characteristic evaluation and size information related to the detected object of interest among the plurality of objects of interest, The system according to claim 30, further configured to present information related to one or more sets of features of the set of the plurality of features to the display device.

36. The status of the characteristic evaluation of the object of interest includes identifying that the object of interest is not small, generating aggregate information of the object of interest, the aggregate information including the position and the size of the object of interest in the plurality of frames, The system according to claim 30, further comprising.

37. The status of the characteristic evaluation of the object of interest includes identifying that the object of interest is small, generating information indicating the status of the characteristic evaluation of the object of interest, the status of the characteristic evaluation including non-aggregate information of the classification of the object of interest, The system according to claim 30, further comprising.

38. A computer-executable method for processing real-time video, comprising: detecting an object of interest in a plurality of frames received from a medical imaging device; performing a characteristic evaluation of the object of interest, the characteristic evaluation having determining a plurality of features related to the object of interest, the plurality of features having the position and size of the object of interest; when the object of interest persists over two or more of the plurality of frames, aggregating information related to the determined position and the determined size of the object of interest; evaluating the determined position and the determined size of the object of interest based on the aggregated information; presenting the aggregated information of the object of interest to a display device when the determined position is in a first body region and the determined size is within a first range. When the determined position is in the second body region and the determined size is within the second range, presenting information indicating the state of the characteristic evaluation of the object of interest to the display device; A method comprising.

39. Identifying medical guidelines based on the determined position and size of the object of interest; Presenting information related to the identified medical guidelines to the display device; The method according to claim 38, further comprising.

40. The method according to claim 38, wherein the plurality of features further includes a classification of the object of interest based on at least one of histological classification, morphological classification, structural classification, or malignancy classification.

41. The method according to claim 38, wherein the determined position related to the object of interest is at least one position of the rectum, sigmoid colon, descending colon, transverse colon, ascending colon, or cecum.

42. The method according to claim 38, wherein the determined size related to the object of interest is a numerical value or a size classification.

43. Detecting a plurality of objects of interest in the plurality of frames; Performing a characteristic evaluation of the plurality of objects of interest, the characteristic evaluation having determining a set of a plurality of features related to the plurality of objects of interest, and one set of features of the set of the plurality of features including characteristic evaluation and size information related to the detected object of interest of the plurality of objects of interest; Presenting information related to one or more sets of features of the set of the plurality of features to the display device; The method according to claim 38, further comprising.

44. The status of the characteristic evaluation of the object of interest is Identifying that the object of interest is not small; Generating aggregate information of the object of interest, the aggregate information including the position and size of the object of interest in the plurality of frames; The method according to claim 38, further comprising.

45. The status of the characteristic evaluation of the object of interest is Identifying that the object of interest is small; Generating information indicating the status of the characteristic evaluation of the object of interest, the status of the characteristic evaluation including non-aggregate information of the classification of the object of interest; The method according to any one of claims 38 to 44, further comprising.

46. A non-transitory computer-readable medium including instructions that, when executed by at least one processor, cause the at least one processor to perform operations for performing real-time video processing, the operations comprising: Detecting regions of interest in a plurality of frames received from a medical imaging device; Performing a characteristic evaluation of the regions of interest, the characteristic evaluation comprising determining a plurality of characteristics associated with the regions of interest, the plurality of characteristics including the position and size of the regions of interest; Aggregating information related to the determined position and size of the regions of interest when the regions of interest persist across two or more of the plurality of frames; Evaluating the determined position and size of the regions of interest based on the aggregated information; Presenting the aggregated information of the regions of interest to a display device when the determined position is in a first body region and the determined size is within a first range; Presenting information indicating the state of the characteristic evaluation of the regions of interest to the display device when the determined position is in a second body region and the determined size is within a second range; A computer-readable medium comprising the above.

47. The operations further comprise: Identifying medical guidelines based on the determined position and size of the regions of interest; Presenting information related to the identified medical guidelines to the display device. The computer-readable medium according to claim 46, further comprising the above.

48. The plurality of characteristics further includes a classification of the regions of interest based on at least one of histological classification, morphological classification, structural classification, or malignancy classification. The computer-readable medium according to claim 46.

49. The determined position associated with the regions of interest is at least one position of rectum, sigmoid colon, descending colon, transverse colon, ascending colon, or cecum. The computer-readable medium according to claim 46. Cecum.

50. The determined size associated with the regions of interest is a numerical value or a size classification. The computer-readable medium according to claim 46.

51. The operations further comprise: Detecting a plurality of regions of interest in the plurality of frames; Performing characteristic evaluation of the plurality of objects of interest, the characteristic evaluation having determining a set of a plurality of features related to the plurality of objects of interest, wherein one set of features of the set of the plurality of features includes characteristic evaluation and size information related to the detected object of interest among the plurality of objects of interest, Presenting information related to one or more sets of features of the set of the plurality of features to the display device; The computer-readable medium according to claim 46, further comprising.

52. The status of the characteristic evaluation of the object of interest is Identifying that the object of interest is not small; Generating aggregate information of the object of interest, the aggregate information including the position and the size of the object of interest in the plurality of frames; The computer-readable medium according to claim 46, further comprising.

53. The status of the characteristic evaluation of the object of interest is Identifying that the object of interest is small; Generating information indicating the status of the characteristic evaluation of the object of interest, the status of the characteristic evaluation including non-aggregate information of the classification of the object of interest; The computer-readable medium according to any one of claims 46 to 52, further comprising.